WO2010094288A1 - Expressivity tag and use thereof - Google Patents

Expressivity tag and use thereof Download PDF

Info

Publication number
WO2010094288A1
WO2010094288A1 PCT/DK2010/050043 DK2010050043W WO2010094288A1 WO 2010094288 A1 WO2010094288 A1 WO 2010094288A1 DK 2010050043 W DK2010050043 W DK 2010050043W WO 2010094288 A1 WO2010094288 A1 WO 2010094288A1
Authority
WO
WIPO (PCT)
Prior art keywords
seq
tag
acid sequence
nucleic acid
infb
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/DK2010/050043
Other languages
French (fr)
Inventor
Kim Kusk Mortensen
Jon Gade Hansted
Hans Uffe Sperling-Petersen
Friederike HÖG
Laura PIETIKÄINEN
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Aarhus Universitet
Original Assignee
Aarhus Universitet
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Aarhus Universitet filed Critical Aarhus Universitet
Publication of WO2010094288A1 publication Critical patent/WO2010094288A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/70Vectors or expression systems specially adapted for E. coli
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/67General methods for enhancing the expression

Definitions

  • the present invention relates to tags which increase the yield of protein expression.
  • the present invention relates to hybrid nucleic acid encoding a hybrid polypeptide comprising an expressivity tag, which increases the yield of the hybrid polypeptide during the manufacturing of the hybrid polypeptide.
  • the gram-negative bacterium Escherichia coli is one of biotechnologies most commonly used organisms for heterologous protein expression. It has many advantages justifying its use. It can grow on inexpensive substrates and still rapidly reach high cell densities. In the bacterial cell, production of the protein of interest can reach sixty to seventy percent of total protein synthesized, giving a much wanted high product to cost ratio.
  • the genome is well characterized and numerous different strains with specific properties are commercially available. Many biotechnological tools are also compatible with the organism, making it easy to manipulate for individual use of interest. In conclusion it is an ideal protein factory for small to medium scaled productions. Being such a widely used expression host several issues also arise in using it for routine applications.
  • Protein synthesis is enormous complicated and a vast amount of variables are to be taken into account when evaluating expression profiles.
  • Expressivity tags are known to be able to enhance the yield of heterologous proteins expressed in E. coli.
  • the tag often placed 5' of the gene of interest, is usually a natural E. coli gene which expresses well, and for unknown reasons promotes a higher level of translation of its 3' carrier gene (Wolfgang Peti a and Rebecca Page, 2007).
  • the first 471 nucleotides of the InfB gene termed InfB(l-471) encodes the domain I of E. coli Initiation Factor II (IF2).
  • IF2 E. coli Initiation Factor II
  • domain I of IF2 is known to enhance its solubility (S ⁇ rensen et. al, 2003).
  • An object of the present invention relates to manufacturing an expression system providing a high yield of protein.
  • Expressivity tags are generally long (above 150 nucleotides) which is generally not favourable when you want to have a specific polypeptide expressed in high amounts. 1) Long expressivity tags have a higher probability of interfering with the polypeptide of interest. 2) Long expressivity tags consume a larger proportion of the overall amino acids available.
  • one aspect of the invention relates to a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide, said nucleic acid comprising a first nucleic acid sequence of maximum 99 nt encoding a polypeptide operably linked to the a second heterologous nucleic acid sequence encoding a polypeptide, wherein said first nucleic acid sequence comprises a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b).
  • the above mentioned hybrid nucleic acid sequence encoding a hybrid polypeptide sequence may be transcribed and translated into said hybrid polypeptide in a suitable expression system.
  • the present invention relates to a recombinant hybrid polypeptide, said polypeptide comprising a first polypeptide amino acid sequence and a second heterologous amino acid sequence, wherein said first amino acid sequence of maximum 33 amino acids is selected from the group consisting of a) an amino acid sequence selected from the group consisting of SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11 and SEQ ID NO: 12, and b) an amino acid sequence having at least 75% sequence identity to a), and c) an amino acid sequence which is a sub-sequence of a) and b).
  • a vector encoding the nucleic acid sequence of the invention incorporated in a vector In particular such vector could be an expression vector.
  • a third aspect of the present invention is to provide a vector comprising a first nucleic acid sequence of maximum 99 nt, said first sequence comprising a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b).
  • An optimized vector can be introduced into a suitable host cell, which can be used for expression the hybrid polypeptide.
  • a fourth aspect of the present invention is to provide a host cell comprising a vector according to the invention.
  • the expressivity tag be used as an affinity tag for an antibody binding.
  • a fifth aspect of the present invention is to provide an isolated antibody or antigen-binding fragment thereof which selectively binds to the first polypeptide amino acid sequence of the invention.
  • the incorporation of the expressivity tag of the invention into and operably linking the tag sequences to the nucleic acid sequences encoding a protein of interest may increase the yield of the polypeptide in a manufacturing process.
  • the present invention relates to the use of a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide according to the invention for the expression of a said polypeptide.
  • An expression system capable of expressing a polypeptide in high amounts by the use of an expressivity tag would be advantageous.
  • the present invention relates to a method for expressing a recombinant polypeptide comprising a) introducing into an expression system the hybrid nucleic acid sequence according to the invention encoding a recombinant hybrid polypeptide, b) expressing said nucleic acid sequence in the expression system, and optionally c) purifying the expressed recombinant hybrid polypeptide .
  • a final aspect relates to a kit comprising a vector according of the invention.
  • Figure 1 is a graphical representation of the protein secondary structure of domain I of IF2 and deletions in the InfB(l-471) gene. Deletions were placed in loop or random coil regions. Structure was obtained from PDB entry 1ND9 and tools at expasy.org.
  • Figure 2 shows InfB ⁇ l-x)-InfA constructs. Numbering refers to nucleotide position in coding region of the genes.
  • InfA gene encodes the Initiation Factor I protein from E.coli.
  • InfB(l-471) encodes the first domain of Initiation Factor II from E.coli.
  • the six consecutive CAC codons encode a polyhistidine tag.
  • the Sad site codes for Glutamic acid-Leucine amino acids. GenBank. InfB GenelD: 5586604. InfA GenelD: 945500.
  • Figure 3 shows InfB ⁇ l-x)-gfp deletion constructs. Numbering refers to nucleotide position in coding region of the genes.
  • the gfp gene encodes the Green Fluorescence Protein (GFP) from Aequorea victoria. The six consecutive CAC codons encode a polyhistidine tag. The Sad site codes for Glutamic acid-Leucine amino acids. GenBank. InfB GenelD: 5586604. gfp GenelD 193239069.
  • Figure 4 shows InfB ⁇ l-x)-gfp deletion constructs. Numbering refers to nucleotide position in coding region of the genes.
  • the gfp gene encodes the Green Fluorescence Protein (GFP) from Aequorea victoria.
  • the six consecutive CAC codons encode a polyhistidine tag.
  • the Sad site codes for Glutamic acid-Leucine amino acids. GenBank. InfB GenelD: 5586604. gfp GenelD 19
  • Figure 4 is a representation of SDS-PAGE of expression of InfB(l-x)-InfA by BL21(DE3) pLysS pET15b cells. Samples were taken after 3.5 hours of 37° C growth with 0.8 mM IPTG. Commassie stained. Lanes: 1. InfA, 2. In f B(IS)-In fA, 3. InfB ⁇ l-12)-InfA, 4. InfB ⁇ l-15)-InfA, 5. InfB ⁇ l-18)-InfA, 6. InfB ⁇ l-21)-InfA, 7. InfB ⁇ l-90)-InfA, 8. InfB ⁇ l-120)-InfA, 9. InfB ⁇ l-177)-InfA, 10. InfB ⁇ l-282)-InfA, and 11. In fB ⁇ l-471)-In fA. 12. Proten molecular mass marker.
  • Figure 5 is a representation of a Western blot of expression of InfB(l-x)-InfA by BL21(DE3) pLysS pET15b cells. Primary antibody used was targeted against IFl.
  • InfB(l-18)-InfA 6. InfB ⁇ l-21)-InfA, 7. InfB(l-90)-InfA, 8. InfB(l-120)-InfA, 9.
  • Figure 6 is a representation of SDS-PAGE of expression of InfB(l-x)-gfp by (A) BL21 pBAD and (B) BL21(DE3) pLysS pET24d. Samples were taken after 3.5 hours of 37° C growth with (A) 0.02 % L-arabinose and (B) 0.8 mM IPTG. Commassie stained. Lanes: 1. Un-induced cells, 2. gfp, 3. InfB(l-9)-gfp, 4.
  • InfB(l-12)-gfp f 5. InfB(l-15)-gfp, 6. InfB(l-18)-gfp, 7. InfB(l-21)-gfp f 8. InfB(l- 90)-gfp, 9. InfB ⁇ l-120)-gfp, 10. InfB ⁇ l-177)-gfp, 11. InfB ⁇ l-282)-gfp, and 12. InfB ⁇ 1-471) -gfp.
  • Figure 7 shows fluorescence from GFP moiety of fusion proteins expressed from InfB(l-x)-gfp in (A) BL21 pBAD and (B) BL21(DE3) pLysS pET24d. " ⁇ " represents background. The fluorescence count of the gfp construct for each expression system is set to 1. A and B are therefore not directly comparable.
  • Figure 8 shows fluorescence from GFP moiety of fusion proteins expressed from InfB(l-x)-gfp in (A) BL21 pBAD and (B) BL21(DE3) pLysS pET24d. " ⁇ " represents background. The fluorescence count of the gfp construct for each expression system is set to 1. A and B are therefore not directly comparable.
  • Figure 8 shows fluorescence from GFP moiety of fusion proteins expressed from InfB(l-x)-gfp in (A) BL21 pBAD and (B) BL21(DE3) pLysS
  • Figure 8 shows insoluble protein fractions of expression of InfB(l-x)-gfp from BL21 pBAD grown with 0.02 % L-arabinose at (A) 37° C and (B) 30° C. The bands observed in all lanes are chromosomal gene products. Insoluble proteins of interest are marked with arrows. Lanes: 1. gfp, 2. InfB ⁇ l-9)-gfp, 3. InfB ⁇ l-12)- gfp, 4. InfB ⁇ l-15)-gfp, 5. InfB ⁇ l-18)-gfp, 6. InfB ⁇ l-21)-gfp, 7. InfB ⁇ l-90)-gfp, 8. InfB ⁇ l-120)-gfp, 9. InfB ⁇ l-177)-gfp, 10. InfB ⁇ l-282)-gfp, 11. InfB ⁇ l-471)-gfp.
  • Figure 9 shows insoluble protein fractions from expression of InfB(l-x)-gfp by BL21(DE3) pET24d pLysS with 0.8 mM IPTG at (A) 37° C, (B) 30° C and (C) 20° C.
  • the bands observed in all lanes are chromosomal gene products.
  • Figure 10 is a representation of SDS-PAGE of expression of InfB(l-x)-gfp by (A) BL21 pBAD and (B) BL21(DE3) pLysS pET24d. Samples were taken after 3.5 hours of growth with (A) 0.02 % L-arabinose at 30° C and (B) 0.8 mM IPTG at 20° C. Commassie stained. Lanes: 1. gfp, 2. InfB ⁇ l-9)-gfp, 3. InfB ⁇ l-12)-gfp, 4. InfB ⁇ l-15)-gfp, 5. InfB ⁇ l-18)-gfp, 6. InfB ⁇ l-21)-gfp, 7.
  • InfB ⁇ l-90)-gfp 8. InfB ⁇ l-120)-gfp, 9. InfB ⁇ l-177)-gfp, 10. InfB ⁇ l-282)-gfp, 11. InfB ⁇ l-471)-gfp.
  • Figure 11 shows fluorescence from GFP moiety of fusion proteins expressed from InfB(l-x)-gfp in (A) BL21 pBAD at 30 ° C and (B) BL21(DE3) pLysS pET24d at 20 0 C. " ⁇ " represents the background fluorescence. The fluorescence count of the gfp construct for each expression system is set to 1. A and B are therefore not directly comparable.
  • Figure 12 shows the purification of protein expressed from (1) InfB(l-9)-gfp (2) InfB(l-12)-gfp (3) InfB(l-15)-gfp (4) InfB(l-18)-gfp (5) InfB(l-21)-gfp constructs by IMAC on a 8 ml_ column of chelating sepharose, charged with Ni2+. Relevant fractions were pooled and are shown in the SDS-PAGE.
  • Figure 13 is a representation of the MALDI-TOF spectra of purified protein from (A) InfB(l-9)-gfp (B) InfB(l-12)-gfp (C) InfB(l-15)-gfp (D) InfB(l-18)-gfp (E) InfB(l-21)-gfp constructs.
  • Mass spectra were acquired by a Voyager-DE PRO instrument. Final spectra obtained by averaging 150-200 single shot spectra which were externally calibrated with Bovine Insulin (5734,6 Da) and E.coli Thioredoxin (11674,5 Da).
  • Figure 14 is a representation of the values obtained from all spectra. The theoretical mass is calculated from the assumption that the N-formyl-methionine was cleaved off.
  • Figure 15 is a representation of the MALDI-TOF zoomed picture of [M + H + ] ion of protein expressed from (A) InfB ⁇ l-9)-gfp and (B) InfB ⁇ l-21)-gfp.
  • the peak with the lowest molecular weight (1) is the analyzed protein without an N-formyl- methionine group.
  • the middle peak (2), seen only clearly in A, is the protein with the methionine residue still attached.
  • the highest molecular weight peak (3) is the protein with the formyl-methionine group attached.
  • Figure 16 is a graphical representation of the folded mRNA sequences of 5'-end of pET15b vector transcript (87 nt) followed by the first 50 nucleotides of /nf ⁇ (l-x)- InfA.
  • the online available MFOLD program was used.
  • SD Shine-Dalgarno sequence.
  • Start Initiation start codon.
  • dG is in units of kcal/mol. The table summarizes the dG values.
  • Figure 17 is a graphical representation of the folded mRNA sequences of 5'-end of pBAD vector followed by 50 first nucleotides of InfB(l-x)-gfp. Folding was done by the MFOLD program. dG is in units of kcal/mol. The table summarizes the dG values.
  • Figure 18 shows the coding sequences of Trx and InfB having expressivity tag properties. Matching bases are marked in bold.
  • Figure 19 shows the fluorescence of GFP moiety of fusion proteins expressed from InfB(l-21)-gfp with point mutations at position (A) 7 (B) 13 and (C) 16. Fluorescence is plotted with the frequency of the codon which the mutation caused. Wildtype bases are indicated with an asterix and set to a fluorescence count of 1. The frequency of the codon plotted at the 2nd y-axis is obtained from a codon table of 8087 coding sequences in the E.coli genome. The fluorescence of gfp is represented with two dotted lines, each being the limits of the standard deviation on the dataset. Constructs were expressed in BL21 pBAD for 3.5 hours at 30° C with 0.02 % L-arabinose.
  • Figure 20 shows the numbers of A/U residues and CA repeat bases have been reported to be of importance for the efficiency of translation, and are shown for sequences of InfB, InfA and gfp.
  • Figure 21 shows the numbers of A/U residues and CA repeat bases have been reported to be of importance for the efficiency of translation, and are shown for sequences of InfB, InfA and gfp.
  • Figure 21 shows Thermus thermophilus 70S ribosome, fMet-tRNA and mRNA in complex, pdb entry: 2HGR .
  • a sphere of rRNA bases is shown around the mRNA at a 16 Anstrom radius in the downstream tunnel.
  • Figure 22 shows the possible interaction between rRNA position C1397 and mRNA position +7.
  • Figure 23 shows the mRNA sequences of InfB(l-21), InfA(l-21) and g/p(l-21). Possible Watson-Crick base pair interactions between mRNA position +7 to rRNA C1397 are marked in bold.
  • Figure 24 shows the expression of LacZ gene as amount of ⁇ -galactosidase fusion protein in percentage of total protein synthesized.
  • the codon producing the most protein is defined as 1, in both studies this was AAA (Lys).
  • Figure 25 shows primers used for making InfB(l-x)-InfA and InfB(l-x)-gfp and InfB(l-21)-gfp point mutants, constructs are listed in table I and II. Annealing temperatures for each PCR reaction is listed in the c°-column.
  • Figure 26 shows primers used for making InfB(l-21)-gfp (Table III). Screening primers are listed in table IV. Annealing temperatures for each PCR reaction is listed in the c°-column.
  • Figure 27 shows primers used for making InfB(l-21)-gfp (Table III). Screening primers are listed in table IV. Annealing temperatures for each PCR reaction is listed in the c°-column.
  • Figure 27 shows primers used for making InfB(l-21)-gfp (Table III). Screening primers are listed in table IV. Annealing temperatures for each PCR reaction is listed in the c°-column.
  • Figure 27 shows primers used for making InfB(l-21)-gfp (Table III). Screening primers are listed in table IV. Annealing temperatures for each PCR reaction is listed in the c°-column.
  • Figure 27 shows primers used for making InfB(l-21)-gfp (Table III). Screening primers are listed in
  • Figure 27 shows the relative expression of the Thio6-GFP system compared to the infB(l-21)-gfp system, measured by the fluorescence of the expressed reporter protein.
  • Figure 28 shows the expression levels of point mutations of infB(l-21)-GFP relative to infB(l-21)-GFP.
  • the mutated nucleic acids and the resulting new codons are indicated below each graph.
  • Figure 29 shows an increased expression of proteins when these proteins comprise the infB(l-21)-tag compared to control proteins.
  • A Streptavidin +/- infB(l-21)-tag.
  • B L32 +/- infB(l-21)-tag.
  • Figure 30 shows sequences and characteristics of the used primers for amplification of infB(l-21)-l32, InfB(l-21)-streptavidin, 132 and streptavidin. These constructs are used in example 9.
  • expressivity tag refers to nucleic acid sequences or amino acid sequences which increase the expression and thus the yield of a polypeptide to which the tag is operably linked.
  • operably linked refers to the connection of elements being a part of a functional unit such as a gene or an open reading frame. Accordingly, by operably linking a promoter to a nucleic acid sequence encoding a polypeptide the two elements becomes part of the functional unit - a gene. The linking of the promoter to the nucleic acid sequence enables the transcription of the nucleic acid sequence directed by the promoter. By operably linking two heterologous nucleic acid sequences encoding a polypeptide the sequences becomes part of the functional unit - an open reading frame encoding a fusion protein comprising the amino acid sequences encoding by the by the heterologous nucleic acid sequences. By operably linking two amino acids sequences, the sequences become part of the same functional unit - a polypeptide. Operably linking two heterologous amino acid sequences generates a hybrid (fusion) polypeptide.
  • Recombinant gene, promoter, nucleic acid sequence (DNA), amino acid sequences (polypeptide) refers to the products generated by genetic engineering such as the combination or insertion of one or more nucleic acid sequences, thereby combining nucleic acid sequences that would not normally occur together (recombinant nucleic acid sequence, recombinant polynucleotide). Accordingly, recombinant expression refers to the expression of a RNA transcript and/or polypeptide from a coding region directed by a heterologous promoter or a promoter comprising heterologous promoter/enhancer elements.
  • heterologous sequences refer to sequence elements of different sequences species.
  • the fusion of nucleic acid sequences of one species of nucleic acids sequence to nucleic acid sequences of another species of nucleic acids generates a recombinant fusion nucleic acid sequence comprising heterologous sequences.
  • the fusion of amino acid sequences of one species of polypeptide to amino acid sequences of another species of amino acid sequences generates a recombinant fusion amino acid sequences comprising heterologous sequences. Accordingly, elements originating from the same species of nucleic acids sequence or polypeptide are not considered heterologous according to the invention.
  • nucleic acid sequence originating from InfB encoding an expressivity tag and a 5' subsequence of InfB encoding a amino acid sequences are not heterologous sequences according to the invention.
  • the molecule formed by operably linking genetic/protein elements of heterologous genetic/protein species is referred to as a "chimera”, “chimeric”, fusion, or hybrid molecule.
  • Hybrid nucleic acid sequence refers to a nucleic acid sequence (polynucleotide) comprising at least two heterologous nucleic acid sequences.
  • the molecule is also referred to as "chimeric" nucleic acid sequence.
  • Hybrid polypeptide or “hybrid amino acid sequences” refer to an amino acid sequence (polypeptide) comprising at least two heterologous amino acid sequences.
  • the molecule is also referred to as "chimeric" polypeptide.
  • sequence identity indicates a quantitative measure of the degree of homology between two amino acid sequences or between two nucleic acid sequences of equal length. If the two sequences to be compared are not of equal length, they must be aligned to give the best possible fit, allowing the insertion of gaps or, alternatively, truncation at the ends of the polypeptide sequences or
  • N ref -N d , f (N ref -N d , f )l00 nucleotide sequences.
  • the sequence identity can be calculated as Nref , wherein Ndif is the total number of non-identical residues in the two sequences when aligned and wherein Nref is the number of residues in one of the sequences.
  • the percentage of sequence identity between one or more sequences may also be based on alignments using the clustalW software
  • nucleotide sequences may be analysed using programme DNASIS Max and the comparison of the sequences may be done at http://www.paraliqn.org/. This service is based on the two comparison algorithms called Smith-Waterman (SW) and ParAlign. The first algorithm was published by Smith and Waterman (1981) and is a well established method that finds the optimal local alignment of two sequences. The other algorithm, ParAlign, is a heuristic method for sequence alignment. Default settings for score matrix and Gap penalties as well as E-values were used.
  • vector refers to a DNA molecule used as a vehicle to transfer recombinant genetic material into a host cell.
  • the four major types of vectors are plasmids, bacteriophages and other viruses, cosmids, and artifical chromosomes.
  • the vector itself is generally a DNA sequence that consists of an insert (a heterologous nucleic acid sequence, transgene) and a larger sequence that serves as the "backbone" of the vector.
  • the purpose of a vector which transfers genetic information to the host is typically to isolate, multiply, or express the insert in the target cell.
  • Vectors called expression vectors are specifically adapted for the expression of the heterologous sequences in the target cell, and generally have a promoter sequence that drives expression of the heterologous sequences.
  • Simpler vectors called transcription vectors are only capable of being transcribed but not translated : they can be replicated in a target cell but not expressed, unlike expression vectors. Transcription vectors are used to amplify the inserted heterologous sequences. The transcripts may subsequently be isolated and used in as templates suitable in vitro translations systems. Functional equivalent
  • the term "functional equivalent” relates in the present context to a nucleic acid sequence or polypeptide sequence which pertains substantially the same activity as nucleic acid sequence or polypeptide sequence to which it is compared.
  • the same functional activity may be an activity of at least 25%, such as at least 40%, such as at least 60 %, such as at least 80% of the sequence to which it is compared.
  • a functional equivalent may have a higher activity such as at the most 400%, such as at the most 350%, such as at the most 300%, such as at the most 200%, such as at the most 175%, such as at the most 150%, or such as at the most 125% when compared to a sequence.
  • the same functional activity as an expressivity tag can be measured by comparing the expression levels of the protein or peptides to which the tags are operably linked. An example of such measurement can be seen in Figure 28 and the corresponding text.
  • single chain antibody “single domain antibody”, “sdAb” (also called Nanobody) is an antibody fragment consisting of a single monomeric variable antibody domain. Like a whole antibody, it is able to bind selectively to a specific antigen. Single domain antibodies are much smaller than common antibodies.
  • the invention relates to a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide, said nucleic acid comprising a first nucleic acid sequence of maximum 99 nt encoding a polypeptide operably linked to the a second heterologous nucleic acid sequence encoding a polypeptide, wherein said first nucleic acid sequence comprises a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b). Length of the first nucleic acid sequence
  • the size of the first nucleic acid sequence is preferably shorter while maintaining an effect as an expressivity tag.
  • the maximum length of the first nucleic acid sequence is 90 nucleotides, such as 80 nucleotides, 70 nucleotides, 60 nucleotides, 50 nucleotides, 40 nucleotides, 30 nucleotides, or such as 24 nucleotides,
  • Substitution may be introduced in the nucleic acid sequences of the invention provided that the modified nucleic acid sequence is a functional equivalent of the nucleic acid sequences from which it originates.
  • the modified nucleic acid sequence encodes an expressivity tag.
  • the modification of the coding sequences may be advantageous in order to optimize codons of the coding region for optimal translation in the host of production.
  • minor modification of the nucleic acid sequences may generate restrictions site, which facilitate the cloning of the sequences.
  • the first nucleic acid sequence has a sequence identity of at least 80% such as at least 85%, at least 90%, at least 95% or such as 100% to a amino acid sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4.
  • nucleic acid sequences which are subsequences of one or more of the sequences of the invention may also be used to obtain an effect similar to the ones disclosed in this invention. Especially sequences having a length between 15 and 99 nucleotides are likely to have an effect similar to the claimed sequences.
  • the sub-sequences of SEQ ID NO: 1 have a length of at least 23 nucleotides, such as 33 nucleotides, 43 nucleotides, 53 nucleotides, 63 nucleotides, 73 nucleotides, or such as 83 nucleotides.
  • second heterologous nucleic acid sequences encoding polypeptides which could be operably linked to the first nucleic acid sequences are GFP, streptavidin, single chain antibodies such as L32 and MMP's. See e.g. examples 3, 9 and 10 and corresponding figures. By having increased expression of GFP more sensitive reporter systems may be constructed. Similar, streptavidin is used in many assys, especially in vitro assays, thus by having a high expression level cheaper assays may be constructed.
  • the second heterologous nucleic acid sequence is selected from the group consisting of reporter constructs such as GFP, streptavidin, MMP's and single chain antibodies such as L32.
  • the first nucleic acid sequence is coding for an expressivity tag.
  • the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, wherein said first nucleic acid sequence encodes an expressivity tag.
  • Length of expressivity tag A minimum length is required of the first nucleic acid sequence of the expressivity tag to obtain an increased expression of the hybrid nucleic acid sequence, compared to when an expressivity tag is not operably linked to a second nucleic acid sequence.
  • the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, wherein said first nucleic acid sequence of c) has a minimum length of 15 nucleotides.
  • This minimum length could e.g. also be 16 nucleotides, such as 17 nucleotides, or such as 18 nucleotides.
  • the tag can be positioned in different positions in relation to the rest of the expressed polypeptide.
  • the invention relates to a recombinant hybrid nucleic acid sequence of the invention, wherein said first nucleic acid sequence is positioned 5' to the second heterologous nucleic acid sequence.
  • the first nucleic acid sequence is positioned in such a way that the expressivity tag will be positioned in the N-terminal end of the expressed fusion polypeptide. Similar, this means that preferably the first nucleic acid sequence is positioned at the 5'-end of the hybrid nucleic acid sequence.
  • the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, wherein said second heterologous nucleic acid sequence has been modified to reflect the codon preference of a host organism selected to express the nucleic acid molecule.
  • hybrid nucleic acid sequence may further comprise one or more additional elements which could be advantageous for detection, solubility or affinity purification of the expressed hybrid polypeptide.
  • the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, further comprising a nucleic acid sequence encoding a polypeptide tag for detection, solubility or affinity purification.
  • this/these one or more tag-elements is positioned in the hybrid nucleic acid sequence depending on the precise use and functionality of the element. It/they are e.g. positioned 3' to the first nucleic acid sequence, between the first and second nucleic acid sequences or 3' to the second nucleic acid sequence. In special cases it/they are positioned 5' to the first nucleic acid sequence.
  • a wide-range of different elements may be employed for detection, solubility or affinity purification.
  • a skilled person will be able to choose the tag which is optimal for his/her preferred use.
  • the skilled person will furthermore, be able to choose both primary and secondary antibodies which can be used for either purification or detection.
  • the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, wherein said tag for affinity purification, solubility or detection is selected from the group consisting of BCCP-tag, c-myc-tag, calmodulin-tag (SEQ ID NO:33), FLAG-tag (SEQ ID NO:45), HA-tag (SEQ ID NO:50), His-tag (SEQ ID NO:29), Maltose binding protein-tag (SEQ ID NO:27), Nus-tag (SEQ ID NO:60), Glutathione-S-transferase-tag (SEQ ID NO:31), Green fluorescent protein-tag (SEQ ID NO:62), Thioredoxin-tag (SEQ ID NO:64), S-tag (SEQ ID NO:41), Strep-tag (SEQ ID NO:35), human protein C tag (SEQ ID NO:37), Chitin binding protein tag (SEQ ID NO:39), T7-tag (SEQ ID NO:43), My
  • the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, further comprising a nucleic acid sequence encoding a spacer positioned between said first and said second nucleic acid sequence.
  • the first nucleic acid sequence may be divided into two or more sections using one or more spacer sequences.
  • the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, further comprising a nucleic acid sequence encoding a spacer positioned between the first and second codon of said first nucleic acid sequence and operably linking said codons.
  • a restriction at the DNA level may be useful for cloning purposes. Restriction sites can e.g. be cleaved using either restriction or nicking enzymes. Cleavage at the peptide level may be useful if part of the hybrid amino acid sequence is not desired in the final peptide.
  • a sequence encoding for specific proteases can be introduced between the first and second nucleic acid sequences. In this way the encoded first amino acid sequence can be removed from the final peptide. This removal can take place both before or after a purification step. Examples of useful proteases are Factor Xa (fXa), TEV, PreScission Protease.
  • Introduction of one or more spacer sequences may also be favourable to avoid structural interference between the first and second amino acid sequences in the hybrid polypeptide.
  • the hybrid nucleic acid sequence of the invention is transcribed generating a transcript from which the hybrid polypeptide is translated. Accordingly, one aspect the invention relates to a recombinant hybrid polypeptide, said polypeptide comprising a first polypeptide amino acid sequence and a second heterologous amino acid sequence, wherein said first amino acid sequence of maximum 33 amino acids is selected from the group consisting of a) an amino acid sequence selected from the group consisting of SEQ ID
  • the maximum length of the first amino acid sequence could be shorter while maintaining an effect as an expressivity tag.
  • the invention relates to a maximum length of the first amino acid sequence being 30 amino acids, such as 26 amino acids, 23 amino acids, 20 amino acids, 17 amino acids, 14 amino acids, 11 amino acids, or such as 9 amino acids.
  • substitution may be introduced in the amino acids in the sequences of the invention provided that the modified amino acid the sequence is a functional equivalent of the amino acid the sequence from which it originates.
  • the modified amino acid the sequence is an expressivity tag.
  • the invention relates a first amino acid sequence having a sequence identity of at least 80% such as at least 85%, at least 90%, at least 95% or such as 100% to a amino acid sequence selected from the group consisting of SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11 and SEQ ID NO: 12. Subsequences
  • sequences which are subsequences of one or more of the sequences of the invention may also be used to obtain an effect similar to the ones disclosed in this invention.
  • sequences having a length between 4 and 33 amino acids are likely to have an effect similar to the claimed sequences.
  • the sub-sequences of SEQ ID NO: 5 and SEQ ID NO:6, are at least 9 amino acids, such as 15 amino acids, or such as 21 amino acids, or such as 27 amino acids.
  • one embodiment of the present invention relates to an isolated hybrid polypeptide comprising a first polypeptide amino acid sequence selected from the group consisting of SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, SEQ ID NO: 12.
  • the first polypeptides of the embodiment reflect the removal of the N- terminal methionine from the corresponding first polypeptides SEQ ID NO: 5, SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO: 11 respectively.
  • the invention relates to a recombinant hybrid polypeptide according to the invention, wherein said first amino acid sequence is an expressivity tag.
  • a minimum length is naturally required of the amino acid sequence of the expressivity tag to obtain an increased expression of the hybrid polypeptide sequence, compared to when an expressivity tag is not operably linked to a second nucleic amino acid sequence.
  • the invention relates to a recombinant hybrid polypeptide according the invention, wherein said first amino acid sequence of c) has a minimum length of 4 amino acids.
  • This minimum length could e.g. also be 5 amino acids, such as 6 amino acids, such as 7 amino acids, such as 8 amino acids, such as 9 amino acids, or such as 10 amino acids. It is to be understood that the minimum lengths cited above relates to said first amino acid sequence when the N-terminal methionine has been removed, as described above.
  • the invention relates to a recombinant hybrid polypeptide according to the invention, wherein said first amino acid sequence is positioned N-terminal to the second heterologous amino acid sequence.
  • the first amino acid sequence is positioned in such a way that the expressivity tag will be positioned at the N-terminal end of the expressed fusion polypeptide.
  • the hybrid amino acid sequence may further comprise one or more additional elements which could advantageous for detection, solubility or affinity purification of the expressed hybrid polypeptide.
  • the invention relates to a recombinant hybrid polypeptide according to the invention, further comprising a polypeptide tag for detection, solubility or affinity purification.
  • a polypeptide tag for detection, solubility or affinity purification can be positioned in the hybrid amino acid sequence depending on the precise use and functionality of the element.
  • the tags for detection, solubility or affinity purification may be positioned C-terminal to the first amino acid sequence, between the first and second amino acid sequences or C-terminal to the second nucleic acid sequence.
  • the one or more tag-elements are positioned N- terminal to the first amino acid sequence
  • a recombinant hybrid polypeptide sequence according to the invention wherein said tag for affinity purification, solubility or detection is selected from the group consisting of BCCP-tag, c-myc- tag, calmodulin-tag (SEQ ID NO:34), FLAG-tag (SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48 and SEQ ID NO:49), HA-tag (SEQ ID NO:51), His-tag (SEQ ID NO:30), Maltose binding protein-tag (SEQ ID NO:28), Nus-tag (SEQ ID NO:61), Glutathione-S-transferase-tag (SEQ ID NO:32), Green fluorescent protein-tag (SEQ ID NO:63), Thioredoxin-tag (SEQ ID NO:65), S-tag (SEQ ID NO:42), Strep- tag (SEQ ID NO:36), human protein C tag (SEQ ID NO:38), Chitin binding protein tag (SEQ ID NO:
  • the recombinant hybrid nucleic acid of the invention may be incorporated in a vector for efficient transfer to and expression of the recombinant gene in a suitable host cell. Therefore in a further aspect the invention relates to a vector comprising a first nucleic acid sequence of maximum 99 nt, said first sequence comprising a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b), said vector.
  • the maximum length of the first nucleic acid sequence comprised in a vector is 90 nucleotides, such as 80 nucleotides, 70 nucleotides, 60 nucleotides, 50 nucleotides, 40 nucleotides, 30 nucleotides, or such as 24 nucleotides,
  • the first nucleic acid sequence has a sequence identity of at least 80% such as at least 85%, at least 90%, at least 95% or such as 100% to a amino acid sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4.
  • the sub-sequences of SEQ ID NO: 1 comprised in a vector have a length of at least 23 nucleotides, such as 33 nucleotides, 43 nucleotides, 53 nucleotides, 63 nucleotides, 73 nucleotides, or such as 83 nucleotides.
  • the first nucleic acid sequences of the invention encoding the expressivity tag may be cloned into a vector:
  • the vector comprising the first nucleic acid sequences may serve as a cloning vector for subsequent cloning of the second heterologous nucleic acid sequence encoding the polypeptide of interest.
  • the second heterologous nucleic acid sequence is inserted in a position that operably links the sequence to the first nucleic acid sequences encoding the expression tag.
  • the cloning vector further comprises a least one promoter element or a complete promoter operably linked to said first nucleic acid sequence.
  • the invention relates to a vector according to the invention, wherein said vector is further adapted for insertion of a second heterologous nucleic acid sequence encoding a polypeptide operably linked to said first nucleic acid sequence.
  • Such second heterologous nucleic acid sequence could be incorporated into the vector by standard cloning techniques known to the skilled person such as PCR- based techniques.
  • a vector comprising the second nucleic acid sequence operably linked to the first nucleic acid sequence (the expression tag) encodes a hybrid polypeptide for efficient expression in a suitable host.
  • the invention relates to a vector comprising a recombinant hybrid nucleic acid sequence according to the invention.
  • the vector may be further adapted for expression of the hybrid polypeptide by insertion of transcriptional control elements such as a promoter or a promoter including additional promoter/enhancer elements.
  • transcriptional control may be inserted thereby generating a expression vector.
  • the use transcriptional control elements depend on the expression system and host.
  • the promoter of the expression vector may comprise several independent sequence elements. These elements may both be elements which can promote expression, elements which can repress expression or a combination of both.
  • the expression control sequence comprises transcriptional control elements suitable for conditional expression of the recombinant hybrid nucleic acid sequence.
  • the invention relates to a vector according to the invention, wherein said recombinant hybrid nucleic acid sequence is operably linked to an expression control sequence, such as the araBAD promoter system such as the Ara-promoter (SEQ ID NO: 13), Ara-operator 1 (SEQ ID NO: 14), Ara- operator 2 (SEQ ID NO: 15) and the AraC repressor (SEQ ID NO: 16), the lac promoter system such as the Lac Promoter (SEQ ID NO: 17), the Lac-operator (SEQ ID NO: 18) and the Lad repressor (SEQ ID NO: 19), Tac/Trc-promoter system (SEQ ID NO:20), Lambda PL - promoter (SEQ ID NO:21), Lambda PR promoter (SEQ ID NO:22), T7 promoter (SEQ ID NO:23), T3 promoter (SEQ ID NO:24), SP6 promoter (SEQ ID NO:25) and the Amp
  • an expression control sequence
  • the invention relates to a vector according to the invention, wherein the vector is selected from the group consisting of plasmids, cosmids, phages, bacterial artificial chromosomes (BAC), Phagemids and Pl- derived artificial chromosomes.
  • the vector is selected from the group consisting of plasmids, cosmids, phages, bacterial artificial chromosomes (BAC), Phagemids and Pl- derived artificial chromosomes.
  • nucleic acid sequences of the invention may also be positioned in oligonucleotides for in vitro translation and/or transcription purposes.
  • the nucleic acid sequences of the invention may also be positioned in PCR-products for cloning purposes and for in vitro transcription and/or translation purposes.
  • the vectors may both be single stranded or double stranded. Furthermore, they may be linear circular or nicked. Especially when using in vitro translation systems it may be advantageous to use nicked or linear plasmids.
  • the invention relates to a vector according to the invention, wherein the vector is an in vitro translation vector.
  • in vitro translations systems are compatible with most in vivo expression system and vice versa. Therefore it is to be understood that the in vitro vectors can also be used for in vivo expression, and that in vivo expression vectors can be used for in vitro expression.
  • the in vitro translation system and or PCR fragments and vectors are selected from the list consisting of:
  • DNA or linear DNA templates with T7 polymerase promoters are used.
  • plasmid vector As described above many different vectors can be used to carry out different aspects of the invention.
  • a plasmid vector is preferred.
  • the invention relates to a vector according to the invention, wherein said plasmid is selected from the group consisting of TA cloning vectors, Gateway cloning vectors, restriction cloning vectors, Topo cloning vectors, pET vector system and pBAD vector systems.
  • the vector is selected from the group consisting of Charon 4, EZ-Tn5 pMOD-2, EZ-Tn5 pMOD-3, EZ-Tn5 pMOD-4, EZ-Tn5 pMOD-5, LITMUS28i, LITMUS38i, PinPoint Control, PinPoint Xa Control, PinPoint Xa-I, PinPoint Xa-2, PinPoint Xa-3, pACYC177, pACYC184, pACYCDuet-1, pALTER-EXl, pALTER-EX2, pBAC-2cp, pBACgus-2cp, pBAD-DEST49, pBAD-TOPO, pBAD- TOPO/LacZ, pBAD-TOPO/l_acZ/V5-His, pBAD/His A, pBAD/His B, pBAD/His C, pBAD/His/LacZ,
  • Host cell For cloning and expression purposes it may be useful to have the vector cloned in a host cell. In this way the vector can be maintained, cloned and purified and thus provide a source for almost unlimited numbers of the vectors. Moreover, stable producer cell lines expressing the hybrid polypeptide of interest may also be advantageous. In order to maintain a producer cell selection marker may be incorporated in the expression vector.
  • the invention relates to a host cell comprising a vector according to the invention.
  • bacteria are in many instances optimal host cells for e.g. expression purposes.
  • the invention relates a host cell according to the invention, wherein said host cell is a bacterium.
  • Bacillus licheniformis Bacillus subtilis and E.coli.
  • Bacillus licheniformis and Bacillus subtilis is e.g. cultured in order to obtain enzymes, such as proteases, for use in biological washing powder, and it may increase production if the expressivity tag according to the invention is positioned in frame in front of such proteases.
  • Enzymes produced by B. subtilis and B. licheniformis are widely used as additives in laundry detergents.
  • the bacterium used for hosting a vector is an E. coli strain.
  • the invention relates a host cell according to the invention, wherein said host cell is E. coli.
  • the host cell is selected among AGl, AB1157, BL21(AI), BL21(DE3), BL21 (DE3) pLysS, BNN93, BNN93 ( ⁇ gtll), BNN97, BW26434 (CGSC Strain # 7658), B5615, B834, B834 (DE3), B834(DE3) pLysS, BLR(DE3), BLR(DE3) pLysS, C600, C600 hflA150 , Y1073, BNN102, CSH50, D1210, DB3.1, DHl, DH5 ⁇ , DHlOB, DH12S, DM1ER2566, ER2267, HBlOl,
  • the host cell is E. coli JM109 and E. coli BL21(DE3).
  • the host cells are selected from non-bacterial origin such as fungi or molds. In a more specific embodiment the host cells are selected among Aspergillus niger and Aspergillus oryzae.
  • A. niger glucoamylase is used in the production of high fructose corn syrup, and pectinases are used in cider and wine clarification, ⁇ -galactosidase, an enzyme that breaks down certain complex sugars, is a component of Beano and other medications which the manufacturers claim can decrease flatulence.
  • pectinases are used in cider and wine clarification
  • ⁇ -galactosidase an enzyme that breaks down certain complex sugars
  • Another use for A. niger within the biotechnology industry is in the production of magnetic isotope-containing variants of biological macromolecules for NMR analysis.
  • A. niger is also cultured for the extraction of the enzymes glucose oxidase (GO) and Alpha-galactosidase (AGS).
  • Glucose oxidase is used in the design of glucose biosensors, due to its high affinity for ⁇ -D-glucose.
  • Alpha-galactosidase can be produced by A. niger fermentation; it is used to hydrolyze alpha 1-6 bonds found in melibiose, raffinose, and stachyose.
  • AN-PEP A. niger prolyl endoprotease
  • AN-PEP a microbial-derived prolyl endoprotease which cleaves gluten. This has strong implications in the treatment of Coeliac (Celiac) disease or other metabolic gluten sensitivity disease processes.
  • Aspergillus oryzae has strong secretion of amylases ( ⁇ -Amylase and glucoamylase), some carboxypeptidases and tyrosinases.
  • the first polypeptide amino acid sequence disclosed in the invention may be used as a target for one or more antibodies.
  • the expressivity tag may also be used for both detection and purification of the hybrid polypeptide using standard molecular techniques, such as affinity purification, and immuno detection.
  • the invention relates to an isolated antibody or antigen-binding fragment thereof which selectively binds to the first polypeptide amino acid sequence of the invention.
  • the invention relates to an antibody, isolated antibody or antigen- binding fragment thereof according to the invention, wherein the antibody is a purification tag or an epitope tag.
  • the recombinant hybrid nucleic acid sequences of the invention are used to express recombinant hybrid polypeptides. This may be especially relevant where large amount of protein is required, e.g. for the production of commercial enzymes such as restriction enzymes, therapeutic proteins, drugs, etc.
  • the presence of the expressivity tag increases the yield of the polypeptide of interest during the fermentation process.
  • one aspect of the invention relates to the use of a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide according to the invention for the expression of a recombinant polypeptide.
  • the invention relates to the use of a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide according to the invention, wherein said recombinant hybrid nucleic acid sequence is comprised in a vector according to the invention adapted for expression of said recombinant polypeptide.
  • the invention relates to a method for expressing a recombinant polypeptide comprising a) introducing into an expression system the hybrid nucleic acid sequence according to the invention encoding a recombinant polypeptide, b) expressing said nucleic acid sequence in the expression system, and optionally c) purifying the expressed recombinant polypeptide.
  • the above disclosed method may be performed by a method, wherein a vector disclosed above is used.
  • the invention relates to a method according to the invention, wherein the hybrid nucleic acid sequence is comprised in a vector according to the invention is used.
  • the invention relates to a method according to the invention, wherein the expression system is a host according to the invention.
  • the invention relates to a method according to the invention, wherein said expression system is an in vitro expression system.
  • nucleic acid sequences of the invention are at some stage RNA. Therefore, it is to be understood that the nucleic acid sequences of the invention can also be RNA sequences.
  • one or more components are normally obtained from a cell extract.
  • Different bacterial strains can be used as the source, but often the extract composition is obtained from E. coli.
  • the invention relates to a method according to the invention, wherein the in vitro translation system comprises extract from E. coli.
  • an antibody which binds to the recombinant polypeptide of the invention may be employed.
  • the invention relates to a method according to the invention, wherein an antibody according to claims the invention is used for purification of said recombinant polypeptide.
  • the invention relates to a method according to the invention, wherein an antibody according to the invention is used for detecting said recombinant polypeptide.
  • cleavage at the peptide level may be useful if part of the hybrid amino acid sequence is not desired in the final peptide.
  • a sequence encoding for specific proteases can be introduced between the first and second nucleic acid sequences. In this way the encoded first amino acid sequence can be removed from the final peptide. This removal can take place both before or after a purification step.
  • useful proteases are Factor Xa (fXa), TEV and PreScission Protease.
  • the invention relates to method according to the invention, wherein said first polypeptide amino acid sequence is subsequently removed from the expressed target nucleic acid sequence by proteolytic cleavage.
  • a vector, as disclosed above, may be very useful in a kit, such as an in vivo translation kit, an in vitro translation kit, and a PCR-kit, a transcription kit.
  • the invention relates to a kit comprising a vector according to the invention.
  • An expressivity tag could be cloned 5' of an essential chromosomal gene of a phage. The gene should be vital for the fitness of the phage.
  • Optimal expressivity tags would be evolved and the fittest phages could be harvested as they would be most abundant.
  • phage display could be used if a selection mechanism could be constructed favoring high expression of the binding protein. Depending on similarity of sequences produced a possible expressivity tag motif could hopefully be defined from such experiments.
  • the invention relates to a method for producing and or optimizing expressivity tags, wherein
  • nucleic acid sequences encoding an expressivity tag of the invention is inserted upstream of operably linked to an essential gene in a phage
  • phages with high fitness is selected optionally
  • phage display would know the different important parameters to adjust to provide a selection pressure.
  • InfB ⁇ l-471) 5' of two different reporter genes; InfA and gfp. Deletions were next made from the 3' end of Inf B ⁇ 1-471) and amounts of protein evaluated by SDS-PAGE, western immunoblotting and quantitative fluorescence measurements. The deletions in InfB(l-471) were placed in regions without protein secondary structure. If the functional part of the expressivity tag was a protein structure, it would by this design not be compromised.
  • Figure 1 shows the secondary structure elements of domain I of IF2 and the 3' InfB ⁇ l-471) deletions. Results showed that the expressivity tag region was located on InfB ⁇ l-21).
  • E. coli JM109 was from New England Biolabs and used for propagation of plasmids.
  • Genotype F ' traD36 proA + B + lacF ⁇ (lacZ)M15/ ⁇ (lac-proAB) glnV44 el4 ⁇ gyrA96 recAl relAl endAl thi hsdR17
  • E. coli BL21(DE3) was from Novagen and used for protein expression combined with the pET vector system.
  • the strain is deficient in the Lon protease and lacks the OmpT outer membrane protease which improves the yield of intact recombinant protein. It contains a copy of the T7 polymerase inserted chromosomally with the DE3 lysogen.
  • Genotype F " , ompT, /7sc/SB(rB ⁇ mB ⁇ ), gal, dcm, (DE3)
  • E. coll BL21 was from Novagen and used for protein expression with the pBAD vector system. This strain is like BL21(DE3) also deficient in the Lon protease and lacks the OmpT outer membrane protease.
  • Genotype F " , ompT, /7sc/SB(rB ⁇ mB ⁇ ), gal, dcm
  • Plasmids pET15b was from Novagen and used for both cloning and expression. It contains a bla gene coding for the ⁇ -lactamase enzyme giving resistance to ampicillin. Transcription from its multiple cloning site is under control of the T7 promoter, requiring the T7 polymerase.
  • pET24d was from Novagen and used for both cloning and expression. It contains a kanamycin resistance gene and transcription from its multiple cloning site is controlled by the T7 promoter.
  • pBAD was from Invitrogen and used for both cloning and expression. It contains an ampicillin resistance gene and transcription from its multiple cloning site is controlled by the araBAD promoter.
  • pLysS was from Novagen and used to avoid leaky expression from pET15b and pET24d plasmids.
  • pLysS has a weak constitutive expression of T7 lysozyme which binds to the T7 polymerase inhibiting its transcription activity.
  • pET15b[IF2-I] was provided by H. P. S ⁇ rensen and was used as a template to create the JnfiB( 1-471) fragments.
  • pGLO was from Bio-Rad and was used as a template to obtain gfp.
  • pET15b[IF2domI-IFl] was provided by T. B Kallehauge and was used as a template to obtain InfA. Relevant enzymes, antibodies and reagents
  • Restriction enzymes ( ⁇ spHI, Sad, Bamftl, EcoRl, Ncol, sphl) was supplied by New England Biolabs ® with its relevant reaction buffer. T7 ligase and reaction buffer was from Pharmacia. Plasmids were purified by QIAGENs Qiaprep Spin Miniprep Kit. Restriction digested DNA fragments were purified by the Wizard ® SV Gel and PCR clean-up system from Omega. Polyclonal rabbit anti-mouse immunoglobulin HRP conjugated antibody was from DAKO. Monoclonal mouse anti- Escherichia coli IFl antibody no. 1 was produced in the laboratory of biodesign. A complete list of buffer and solutions can be found in Appendix B.
  • InfA was first constructed by one PCR fragment which was ligated into a pET15b vector.
  • PCR InfA I was made with a pET15b[IF2-IFl] template which introduced restrictions sites BspHI and BamHI and six consecutive histidines.
  • the BspHl/BamHl digested InfA I PCR fragment was ligated into an Ncol/BamHl digested pET15b plasmid to make InfA.
  • InfB ⁇ l-471)-InfA, InfB(l-282)-InfA, InfB ⁇ l-177)-InfA and InfB(l-120)-InfA was next constructed by two PCR fragments, a InfB part and a InfA named InfA II, which was ligated into pET15b.
  • PCR fragments InfB ⁇ l-471), InfB ⁇ l-282), InfB ⁇ l-177) and 7nf ⁇ (l-120) were prepared with a pET15b[IF2-I] template.
  • the primers introduced restriction BspHI and Sad.
  • An InfA II PCR was made with a pET15b[IF2-IFl] template and introduced restrictions sites Sad and BamHI and six consecutive histidines.
  • InfB ⁇ l-9)-InfA, InfB ⁇ l-12)-InfA, InfB ⁇ l-15)-InfA, InfB ⁇ l-18)-InfA, InfB ⁇ l-21)-InfA and InfB(l-90)-InfA were constructed each by one PCR fragment which was ligated into pET15b.
  • PCR fragments 7nf ⁇ (l-9), JnJB(I- 12), Jn/B(l-15), 7nf ⁇ (l-18), 7nf ⁇ (l-21) and 7nf ⁇ (l-90) were prepared with a pET15b[7nf ⁇ (1-471)-Jn/5A] template and introduced restriction sites Sphl and Sad.
  • Sphl/Sacl digested 7nf ⁇ (l-9), InfB ⁇ l- 12), 7nf ⁇ (l-15), 7nf ⁇ (l-18), InfB ⁇ l-21) and 7n ⁇ B(l-90) were ligated with SacI/ ⁇ amHI digested 7nfiA II into a Sphl/BamHl digested pET15b plasmid to make InfB ⁇ l-9)-InfA, InfB ⁇ l-12)-InfA, InfB ⁇ l-15)-InfA, InfB ⁇ l-18)-InfA, InfB ⁇ l-21)-InfA and InfB ⁇ l-90)-InfA.
  • gfp was first constructed by one PCR which was ligated into pBAD.
  • PCR gfp I was made with a template pGLO plasmid. This introduced restriction sites ⁇ spHI and EcoRI along with six consecutive histidines.
  • the ⁇ spHI/EcoRI digested PCR fragment was ligated into an NcoI/EcoRI digested pBAD plasmid to make gfp.
  • PCR fragments InfB ⁇ l-471), InfB ⁇ l-282), InfB ⁇ l-177), InfB ⁇ l-120) were prepared 5 with a pET15b[IF2-I] template.
  • the primers introduced restriction ⁇ spHI and Sad.
  • a PCR of gfp II was made with pGLO as template which introduced restriction sites Sad and EcoRI and six consecutive histidines.
  • SacI/EcoRI digested gfp II was ligated with ⁇ spHI/Sad digested 7nf ⁇ (l-471), 7nf ⁇ (l-282), 7nf ⁇ (l-177) or InfB(l- 120) into a NcoI/EcoRI digested pBAD plasmid to make InfB ⁇ l-471)-gfp, InfB ⁇ l-
  • InfB ⁇ l-9)-gfp, InfB ⁇ l-12)-gfp, InfB ⁇ l-15)-gfp, InfB ⁇ l-18)-gfp, InfB ⁇ l-21)-gfp were next constructed by one PCR fragment each which was ligated into pBAD.
  • Primers contained a 5' non-annealing overhang of the InfB part along with an annealing gfp part.
  • InfB(l-18)-gfp, InfB(l-21)-gfp were made using pGLO as template. It introduced restriction sites ⁇ spHI, EcoRI and six consecutive histidines. The five ⁇ spHI/EcoRI digested PCR fragments were ligated into ⁇ /coI/EcoRI digested pBAD plasmids to make InfB ⁇ l-9)-gfp, InfB ⁇ l-12)-gfp, InfB ⁇ l-15)-gfp, InfB ⁇ l-18)-gfp, InfB ⁇ l-21)- gfp.
  • the InfB ⁇ l-90)-gfp construct was made by 3 PCR reactions.
  • First PCR produced a fragment of 7nf ⁇ (l-90) with a gfp 3' overhang named 7nf ⁇ (l-90) I.
  • Second PCR produced a fragment of gfp named 7nf ⁇ (l-90) II.
  • Finally a PCR named 7nf ⁇ (l-90) III was made using the two PCR fragments as primers, as they annealed to each other.
  • the 7nf ⁇ (l-90) III PCR product was ligated into pBAD.
  • 25 I was made with pET15b[IF2-I] as template which introduced restriction site ⁇ spHI and a 3' gfp overhang.
  • Next PCR 7nf ⁇ (l-90) II was made with a pGLO template which introduced restriction site EcoRI and six consecutive histidines.
  • a PCR using 7nf ⁇ (l-90) I and 7nf ⁇ (l-90) II as primers was made which produced a InfB ⁇ l-90)-gfp PCR fragment with ⁇ spHI and EcoRI restriction sites.
  • Plasmids with correctly sized inserts were sequenced by DNA-technology A/S using the screening primers.
  • pET15b and pET24d constructs were transformed into BL21(DE3) pLysS cells while pBAD constructs were transformed into BL21 for subsequent expression.
  • BL21(DE3) pLysS harboring pET15b or pET24d constructs and BL21 harboring pBAD constructs were grown in 3 ml_ 2xTY media with appropriate antibiotics (cone. 100 ⁇ g/mL ampicillin, 34 ⁇ g/mL chlorampenicol, 50 ⁇ g/mL kanamycin)
  • Amount of protein expressed from InfB(l-x)-gfp constructs was analyzed, by measuring the fluorescence from a fixed amount of cells.
  • 16 ml_ culture was centrifuged (3.700 x g, 10 min, 4 0 C), resuspended in 1 ml_ 2xTY and inoculated into 1,6 L 2xTY with antibiotics.
  • Cells were grown at 37° C, until the optical density at 600 nm reached 0.8, at which they were induced with 0.8 mM IPTG. Growth was continued 3.5 hours at 30° C.
  • Cells were harvested by centrifugation (9.500 x g, 10 min, 4 0 C), resuspended in 0.9 % NaCI and centrifuged again (3.700 x g, 25 min, 4 0 C). Cell pellets were stored at -20° C.
  • Fusion proteins expressed from InfB(l-9)-gfp, InfB(l-12)-gfp, InfB(l-15)-gfp, InfB(l-18)-gfp and InfB(l-21)-gfp contained a C-terminal His-tag, and were purified by Immobilized Metal Affinity Chromatography (IMAC).
  • IMAC Immobilized Metal Affinity Chromatography
  • Cells previously grown for expression were resuspended in 2 ml_ buffer A per gram cells. Cell suspensions were sonicated in a 50 ml_ Nunc tube on ice, for a total of 15 min (0,5 cycle, 100 amplitude). Lysed cells were ultracentrifuged (184.000 x g, 75 min, 4° C) to sediment cell debris and insoluble protein.
  • Soluble protein was 0,2 ⁇ m filtered and loaded onto a 8 ml_ chelating sepharose column, pre-loaded with Ni 2+ and calibrated in Buffer A. Bound protein was eluted by a stepwise 9-15-100 % gradient of Buffer C against Buffer A. Relevant fractions were analysis by SDS-PAGE. Identified fractions were pooled in a dialysis sack (MWCO: 6-8 kDa), and dialysed against 1 L of storage buffer over night at 4°C. Protein was stored at -20 0 C.
  • MALDI-TOF Matrix-assisted laser desorption/ionization time-of-flight
  • Fusion proteins expressed from InfB ⁇ l-9)-gfp, InfB ⁇ l-12)-gfp, InfB ⁇ l-15)-gfp, InfB ⁇ l-18)-gfp and InfB ⁇ l-21)-gfp constructs and purified as described were analyzed by MALDI-TOF, to confirm their masses. All five types of protein were dialysed against 2 L 50 mM Tris-HCI pH 7.5 overnight, 4°C. 0.5 ⁇ l_ of protein (20 pmol/ ⁇ L) was mixed with 0.5 ⁇ l_ sinapinic acid (lOmg/mL) on a MALDI plate. Samples were left to crystallize for 30 min. at room temperature.
  • Mass spectra were acquired by a Voyager-DE PRO (MALDI-TOF) instrument (Applied Biosystems). Final spectra were obtained by averaging 150-200 single shot spectra. Spectra were externally calibrated with bovine insulin (5734.59 Da) and E.coli thioredoxin (11674,48 Da).
  • Protein expressed from InfB(l-18)-gfp, InfB(l-21)-gfp and InfB(l-90)-gfp was analyzed by N-terminal sequencing by Prof. Claus Oxvig. Protein was excised from SDS-PAGE bands run with total cellular extracts from the small-scale expression.
  • RNA Sequence type Linaer Temperature: 37° C Percent: 5 Na 2+ cone: 1.0 Mg 2+ cone: 1.0
  • InfA was expressed at lower levels than either of the InfB ⁇ l- x)-InfA constructs. It was not possible to see much difference between any of the InfB deletions, except InfB(l-177)-InfA which seemed to have a slightly higher expression.
  • the molecular weight of the proteins expressed from the InfB(l-9)- InfA, InfB(l-12)-InfA, InfB(l-15)-InfA, InfB(l-18)-InfA did not correspond to the molecular weight, of the surrounding proteins in the SDS-PAGE.
  • Figure 6a showed gradually increasing amounts of protein expressed from gfp to InfB(l-21)-gfp.
  • InfB(l-9)-gfp and InfB(l-12)-gfp were comparable to gfp, while InfB ⁇ l-15)-gfp, InfB ⁇ l-18)-gfp and InfB ⁇ l-21) produced more protein.
  • InfB(l-90)-gfp and larger constructs seemed comparable to InfB(l-
  • the amount of protein expressed from InfB ⁇ l-x)-gfp was measured by the fluorescence from a volume of induced cells having a fixed cell density (figure 7).
  • InfB(l-90)-gfp, InfB(l-120)-gfp and InfB( 1-471) -gfp had 2-3 fold lower fluorescence counts than InfB(l-177)-gfp and InfB(l-282)-gfp, which did not match with the intensity of their corresponding protein bands in the SDS-PAGE.
  • Results in figure 7b also showed a deviation from the corresponding results in the SDS-PAGE (figure 6b).
  • gfp had the highest fluorescence count, while 7n ⁇ B(l-90)- gfp, InfB ⁇ l-120)-gfp and InfB ⁇ l-471)-gfp had very low counts as in figure 6a.
  • the difference in the amount of insoluble protein fractions produced for pBAD and pET24 was caused by a difference in transcriptional levels, and thereby overall protein expressed.
  • the pET system produced more protein than pBAD, and therefore also had higher amounts of inclusion bodies.
  • BL21 pBAD [n ⁇ 3(l-x)-g/p] and BL21(DE3) pLysS pET24d[/nf ⁇ (l-x)-g/p] cells were grown at 30° C and 20° C and analyzed by SDS-PAGE (figure 10).
  • protein synthesis initiates with a formyl-methionine.
  • the formyl group is removed post-translationally by a peptide deformylase (PDF).
  • PDF peptide deformylase
  • the methionine residue is, depending on the identity of the second amino acid in the peptide chain, removed by a methionine aminopeptidase (MAP).
  • MAP methionine aminopeptidase
  • the C-terminal breakage fragment did not produce any clear peak on the spectra. It had a size of approximately 20 kDa, and is marked on all spectra even though only a minor peak was seen.
  • the fragmentation relates to the nature of the MALDI-TOF method.
  • the crystallized protein was excited by laser photons which fragmented the GFP protein, as a result of the high intensity.
  • Figure 14 summarizes the values obtained from all spectra.
  • the theoretical mass is calculated from the assumption that the N-formyl-methionine was cleaved off. It can be seen, that the measured mass for each construct was close to its theoretical value.
  • InfB(l-9)-InfA, InfB(l-12)-InfA and InfB(l-18)-InfA furthermore had comparable free energies of only a marginal higher value than InfA.
  • InfB ⁇ l-21)-InfA had the highest free energy. The foldings did not correspond well with the amount of each protein produced (figure 4 and 5). Interpreting how base paired the SD-sequence and initiation start codon was in each construct, was not found feasible.
  • Trx encodes the Thioredoxin Reductase protein and its sequence is seen in figure 18.
  • the tags of the invention can be considered better expressional tags than thioredoxin.
  • the expressional levels were measured using fluorescence. Three codons out of six and twelve out of eighteen bases were the same. In order to check these possibly conserved positions were of particular importance to the expressivity of In fB ⁇ l-21), point mutations were made at position G7, A13 and A16. The mutants were expressed in BL21 pBAD and the fluorescence measured and compared to the wildtype InfB ⁇ l-21)-gfp construct.
  • Fluorescence count is plotted in figure 19 together with the frequency of the specific codon which the point mutation caused, to see if codon frequency could be correlated with expression. It could be seen that all mutations to some extent compromised the optimal expressivity of the wildtype InfB(l-21).
  • Figure 28 shows that a number of mutations result in expression levels comparable to InfB ⁇ l-21) (seq ID NO: 2).
  • the first nucleic acid sequence of the invention is mutated in at least one the following positions, as indicated, selected from the group consisting of positions 5C ⁇ 5A (SEQ ID: NO 68), 7G ⁇ 7A (SEQ ID: NO
  • SEQ ID: NO 85 21A ⁇ 21G
  • SEQ ID: NO 86 21A ⁇ 21T
  • SEQ ID: NO 87 21A ⁇ 21T
  • Approximately 15 nucleotides of the mRNA are placed in the downstream tunnel. Position +11 to +15 of the mRNA is surrounded by a layer of protein consisting of S3, S4 and S5. A dense array of basic Arg amino acids is projected into the tunnel from these proteins. This is thought to interact with backbone phosphates of the mRNA. From position +7 to +10 the mRNA passes through a layer of rRNA. Cytosine 1397 and Uracil 1196 from this layer is oriented towards the mRNA around position +7 and +9 on the mRNA. It is though that these bases might help to position the mRNA immediately downstream of the A codon position. Base C1397 is shown in figure 22 in interaction with position +7 of the mRNA. The base U1196 is spaced from the mRNA in an intermediate position between +8 and +9. It is placed on the internal side of the downstream tunnel but not oriented directly towards the +9 mRNA base
  • the +1 codon frame is where the AUG codon is in the P-site. It can be seen that in this frame InfB ⁇ l-21) contains a G at the first codon position, which neither InfA or gfp does. In the next codon frames, InfA and gfp contains more G's at the first codon position than InfB. It could be speculated, that a G at position +7 in the mRNA is of particular importance for the coordination of the transcript in the tunnel. As the first and second peptide bond is formed, it is no longer of such importance. This is however only speculations, trying to fit the expressivity of InfB(l-21) to the structural arrangement. The analysis does also not explain why the entire sequence of In fB ⁇ l-21), is of importance to its expressivity.
  • the main aim of this project was to determine a expressivity tag region on the 471 nucleotide transcript of In fB ⁇ l-471) and an expressivity tag motif should be defined which was general for the Escherichia coli organism.
  • Odd migration pattern At first, the puzzle of the odd migration pattern observed in the SDS-PAGEs in figures 4 and 6 needs to be addressed. A lot of the results confirmed the integrity of the constructs produced. All constructs were sequenced correctly; an anti-IFl monoclonal antibody clearly recognized IFl protein expressed from InfB ⁇ l-x)-InfA constructs (figure 5). Fluorescence was measured from protein expressed from InfB(l-x)-gfp constructs (figure 11), and MALDI-TOF measurements of a selection of proteins involved in the odd migration pattern verified their mass within a very small margin of error (figure 13). The idea that an error in the primary sequence was the cause of the migration pattern was ruled out.
  • the lower InfB ⁇ 1-471 )-gfp expression from pET24d can be explained by the fraction of insoluble protein produced at 20° C (figure 9c ).
  • NGG codons at certain positions are known to cause peptidyl tRNA drop off, could be an example of the contrary to InfB(l-21), NGG codons causing unfavorable interactions. Since they cause peptidyl-tRNA drop off if from codon position +2 to +5 it seems reasonable that the A-site is of interest. If NGG is placed e.g. position + 5, after four peptide bonds are formed it is in the A-site and could cause drop off. More work should be put into analysis of the structural movements in the decoding process between ribosome, mRNA and tRNA.
  • PCR amplification PCR reactions were made in 100 ⁇ l_ aliquots by the following buffer concentrations and reaction conditions. Annealing temperatures for each PCR reactions along with primers can be found in figures 25 and 26.
  • Plamids were purified by QIAGENs Qiaprep Spin Miniprep Kit by the supplied protocol. 3 ml_ of bacterial cell culture grown over night at 37° with appropriate 20 antibiotics was used per spin column and DNA was eluted with 60 ⁇ l_ H2O per spin column.
  • PCR fragments and digested plasmids were run on a 1 % agarose gel using 25 marker lanes to avoid UV-exposure of the DNA. Bands were excised and the DNA extracted from the gel by the Wizard® SV Gel and PCR clean-up system according to the supplied protocol. Digested PCR fragments were directly purified by the Wizard® SV Gel and PCR clean-up system.
  • Digestion reactions were performed by protocols from New England Biolabs using relevant restriction enzymes and buffers. A standard 75 ⁇ l_ scheme was used as shown below. Restriction reactions were incubated 90 minutes at 37 0 C. If double digest reactions were not possible single digests were performed with an
  • Ligation Digested inserts and plasmids were ligated by one of two schemes below depending on whether one or two inserts were needed. Ligation reactions were incubated at room temperature for 45 minutes.
  • JM109, BL21 or BL21(DE3) pLysS cells were first made competent for transformation.
  • Cells were grown in 3 mL 2xTY media overnight at 37° C. 600 ⁇ L overnight culture was collected by centrifugation (15.000 x g, 1 min) and redissolved in 100 ⁇ L fresh 2xTY.
  • the cell suspension was inoculated into 60 mL fresh 2xTY media and grown at 37 0 C until OD600 reached 0,5.
  • Cells were collected by centrifugation (3.700 x g, 10 min, 4 0 C), resuspended in 10 mL TFl solution and incubated on ice for 30 min.
  • Clones were picked from transformation plates and grown in 3 ml_ 2xTY at 37° C over night with appropriate antibiotics. Plasmid from each clone was purified by the QIAGENs Qiaprep Spin Miniprep Kit and eluted into 60 ⁇ l_ H2O per spin
  • a PCR screen was performed with T7 screen primers or pBAD screen primers according to the general PCR amplification scheme above with the primers listed in figure 26. Clones with right sized inserts were sent to sequencing at DNA- technology A/S. Sequencing primers used were T7 screen primers and pBAD screen primers.
  • Loading buffer (3x) 0.13 % bromophenol blue (v/v), 0.13 % xylene cyanol (v/v) , 30 30 % glycerol (v/v) TAE-buffer (5Ox) 2 M Tris-HCI (pH 8.0), 2 M acetic acid, 50 mM EDTA Molecular weight marker ⁇ -DNA, BstEII digested (Bilabs)
  • Lysis buffer 50 mM Tris-HCI pH 8, 150 mM NaCI, 1 mM EDTA, 0.1 mM PMSF
  • PCR buffer (1Ox) 750 mM Tris-HCI (pH 9.0), 200 mM (NH4)2SO4, 0.1 % tween (v/v)
  • NEB buffer 1 (1Ox) 100 mM Bis-Tris-Propane-HCI, 100 mM MgCI2, 10 mM dithiothreitol, pH 7
  • NEB buffer 3 (1Ox) 1 M NaCI, 500 mM Tris-HCI, 100 mM MgCI2, 10 mM dithiothreitol, pH 7.9
  • NEB buffer 4 (1Ox) 500 mM potassium acetate, 200 mM tris-acetate, 100 mM magnesium acetate, 10 mM dithiothreitol, pH 7.9 Sad 20 U/ ⁇ L
  • T4 ligase buffer (1Ox) 0.5 M Tris-HCI (pH 7.5), 0.1 MgCI2, 0.01 M ATP, 0.1 M
  • Buffer A 50 mM Tris-HCI (pH 7.6), 500 mM NaCI, 5 mM Imidazole
  • Buffer C 50 mM Tris-HCI (pH7.6), 500 mM NaCI, 500 mM Imidazole.
  • Electrophoresis buffer (1Ox) 0.25 M Tris-HCI (pH 8.3), 1.92 M glycine, 1% SDS Molecular weight marker PageRulerTM Prestained protein ladder (fermentas).
  • the ECL Plus Western Blotting Detection System (GE healthcare) was used for westernblotting .
  • the kit consists of two solutions ECLA and ECLB, compositions of these is unknown.
  • Streptavidin (Genbank: X03591.1) is from the organism Streptomyces avidinii. Its 24 amino acid signal peptide was excluded (Argarana, C. E., et al. 1986, Miksch G, et al. 2008).
  • L32 is a single chain antibody fragment (scFv) of human origin (Sanz, L., et al. 2001 and Sanz, L., et al. 2002).
  • Streptavidin has been expressed in E. coli in several vector contexts, and is included in this example, since it is a very important biotechnological protein product. In most of the reported expression systems for streptavidin, the system has been optimally tuned for high expression. Many of these optimized systems for heterologous expression often results in premature translational abortion and protein folding problems. In order to avoid these problems it is often a necessity to employ sub-optimal expression conditions.
  • This example is carried out under such sub-optimal conditions, e.g. less strong promotor and low expression temperature. Under these conditions streptavidin and L32 is expressed at levels that is barely or not detectable on a SDS-PAGE.
  • InfB ⁇ l-21)-l32, InfB ⁇ l-21)-streptavidin, 132 and streptavidin were amplified from a template containing the appropriate gene by standard PCR.
  • the infB ⁇ l-21) sequence was introduced through a non-annealing part of the forward primer (see Figure 30).
  • Forward primer contained a BspHl restriction site, followed by the infB ⁇ l-21) sequence and the annealing region of the appropriate gene.
  • the forward primer for genes without an infB ⁇ l-21) sequence contained a ⁇ /coI restriction site, followed by the annealing region of the appropriate gene ⁇ 132 or streptavidin).
  • Reverse primer contained an EcoRI restriction site followed by the annealing region of the appropriate gene ⁇ 132 or strep).
  • the PCR amplification included 30 cycles of 1 min at 94 0 C (DNA denaturation), 1 min at 56 0 C (primer hybridization) and 1 : 10 min at 72 0 C (DNA elongation), followed by a single step for 5 min at 72 0 C (final elongation).
  • the PCR products were digested with BspHl or ⁇ /coI, respectively, and EcoRI and ligated into a ⁇ /coI and EcoRI treated pBAD plasmid, followed by transformation of ligation mixture into chemical competent E. coli JM109 cells.
  • Protein expression pBAD/n ⁇ 3(l-21)-/32, InfB(l-21)-strep and pBAD-132, -streptavidin) was transformed into chemical competent E. coli BL21 cells.
  • two 50 ml_ cultures were grown in 2x TY (containing 100 ⁇ g/ml ampicillin) at 37 0 C.
  • the gene expression was induced by adding L-arabinose to a final concentration of 0,04 %.
  • one culture was incubated at 30 0 C and another at 37 0 C for 3,5 h post induction. Analysis of expression was performed by a SDS-PAGE.

Landscapes

  • Genetics & Genomics (AREA)
  • Health & Medical Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Wood Science & Technology (AREA)
  • Organic Chemistry (AREA)
  • Biotechnology (AREA)
  • General Engineering & Computer Science (AREA)
  • Biomedical Technology (AREA)
  • Zoology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Molecular Biology (AREA)
  • Microbiology (AREA)
  • Plant Pathology (AREA)
  • Physics & Mathematics (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Biophysics (AREA)
  • Micro-Organisms Or Cultivation Processes Thereof (AREA)
  • Peptides Or Proteins (AREA)

Abstract

The present invention relates to tags which increase the yield of protein expression. In particular the present invention relates to hybrid nucleic acid encoding a hybrid polypeptide comprising anexpressivity tag, which increases the yield of the hybrid polypeptide during the manufacturing of the hybrid polypeptide.The invention furthermore relates to expressed polypeptides, vectors, host cells and methods for using the hybrid nucleic acid sequences for expression purposes.

Description

Expressivity tag and use thereof
Technical field of the invention
The present invention relates to tags which increase the yield of protein expression. In particular the present invention relates to hybrid nucleic acid encoding a hybrid polypeptide comprising an expressivity tag, which increases the yield of the hybrid polypeptide during the manufacturing of the hybrid polypeptide.
Background of the invention
The gram-negative bacterium Escherichia coli is one of biotechnologies most commonly used organisms for heterologous protein expression. It has many advantages justifying its use. It can grow on inexpensive substrates and still rapidly reach high cell densities. In the bacterial cell, production of the protein of interest can reach sixty to seventy percent of total protein synthesized, giving a much wanted high product to cost ratio. The genome is well characterized and numerous different strains with specific properties are commercially available. Many biotechnological tools are also compatible with the organism, making it easy to manipulate for individual use of interest. In conclusion it is an ideal protein factory for small to medium scaled productions. Being such a widely used expression host several issues also arise in using it for routine applications. Besides its limitations in not being able to make post-translational modifications and large complex proteins with multiple disulfide bonds, a remarkable difference exist in the efficiency in which both heterologous and homologous proteins are produced. The bias ranges from super efficient expression of some proteins, to a total shut-down in others. The development of new and improved methods which secures a stable expression of any homologous and heterologous gene E. coli is therefore much wanted.
Protein synthesis is immensely complicated and a vast amount of variables are to be taken into account when evaluating expression profiles. Four main areas influence the absolute level of a protein : Transcription, mRNA stability, translation and protein half-life. Expressivity tags are known to be able to enhance the yield of heterologous proteins expressed in E. coli. The tag, often placed 5' of the gene of interest, is usually a natural E. coli gene which expresses well, and for unknown reasons promotes a higher level of translation of its 3' carrier gene (Wolfgang Peti a and Rebecca Page, 2007).
The first 471 nucleotides of the InfB gene termed InfB(l-471) encodes the domain I of E. coli Initiation Factor II (IF2). When fused to the aggregation prone streptavidin protein and expressed in E. coli, domain I of IF2 is known to enhance its solubility (Sørensen et. al, 2003).
Improved expressivity tags which can increase expression of most proteins would be advantageous.
Summary of the invention
An object of the present invention relates to manufacturing an expression system providing a high yield of protein. In particular, it is an object of the present invention to provide a short expressivity tag, which may incorporated into a heterologous polypeptide sequence of interest in order to increase the yield in a manufacturing process.
Expressivity tags are generally long (above 150 nucleotides) which is generally not favourable when you want to have a specific polypeptide expressed in high amounts. 1) Long expressivity tags have a higher probability of interfering with the polypeptide of interest. 2) Long expressivity tags consume a larger proportion of the overall amino acids available.
Recombinant hybrid nucleic acid sequence Thus, one aspect of the invention relates to a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide, said nucleic acid comprising a first nucleic acid sequence of maximum 99 nt encoding a polypeptide operably linked to the a second heterologous nucleic acid sequence encoding a polypeptide, wherein said first nucleic acid sequence comprises a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b).
Operably linked to a suitable promoter, the above mentioned hybrid nucleic acid sequence encoding a hybrid polypeptide sequence may be transcribed and translated into said hybrid polypeptide in a suitable expression system.
Recombinant hybrid polypeptide
In another aspect the present invention relates to a recombinant hybrid polypeptide, said polypeptide comprising a first polypeptide amino acid sequence and a second heterologous amino acid sequence, wherein said first amino acid sequence of maximum 33 amino acids is selected from the group consisting of a) an amino acid sequence selected from the group consisting of SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11 and SEQ ID NO: 12, and b) an amino acid sequence having at least 75% sequence identity to a), and c) an amino acid sequence which is a sub-sequence of a) and b).
In order to express the hybrid polypeptide of the invention in a suitable host a vector encoding the nucleic acid sequence of the invention incorporated in a vector. In particular such vector could be an expression vector.
Vector
Therefore, a third aspect of the present invention is to provide a vector comprising a first nucleic acid sequence of maximum 99 nt, said first sequence comprising a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b).
Host cell
An optimized vector can be introduced into a suitable host cell, which can be used for expression the hybrid polypeptide.
Thus, a fourth aspect of the present invention is to provide a host cell comprising a vector according to the invention.
Apart from increasing the yield during manufacturing of the hybrid polypeptide, the expressivity tag be used as an affinity tag for an antibody binding.
Antibody
Thus, a fifth aspect of the present invention is to provide an isolated antibody or antigen-binding fragment thereof which selectively binds to the first polypeptide amino acid sequence of the invention.
The incorporation of the expressivity tag of the invention into and operably linking the tag sequences to the nucleic acid sequences encoding a protein of interest may increase the yield of the polypeptide in a manufacturing process.
Use of a recombinant hybrid nucleic acid sequence Therefore, in a sixth aspect the present invention relates to the use of a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide according to the invention for the expression of a said polypeptide.
Method for expressing a recombinant polypeptide
An expression system capable of expressing a polypeptide in high amounts by the use of an expressivity tag would be advantageous.
Thus, in a seventh aspect the present invention relates to a method for expressing a recombinant polypeptide comprising a) introducing into an expression system the hybrid nucleic acid sequence according to the invention encoding a recombinant hybrid polypeptide, b) expressing said nucleic acid sequence in the expression system, and optionally c) purifying the expressed recombinant hybrid polypeptide .
Kit
A final aspect relates to a kit comprising a vector according of the invention.
Brief description of the figures
Figure 1 Figure 1 is a graphical representation of the protein secondary structure of domain I of IF2 and deletions in the InfB(l-471) gene. Deletions were placed in loop or random coil regions. Structure was obtained from PDB entry 1ND9 and tools at expasy.org.
Figure 2
Figure 2 shows InfB{l-x)-InfA constructs. Numbering refers to nucleotide position in coding region of the genes. InfA gene encodes the Initiation Factor I protein from E.coli. InfB(l-471) encodes the first domain of Initiation Factor II from E.coli. The six consecutive CAC codons encode a polyhistidine tag. The Sad site codes for Glutamic acid-Leucine amino acids. GenBank. InfB GenelD: 5586604. InfA GenelD: 945500.
Figure 3
Figure 3 shows InfB{l-x)-gfp deletion constructs. Numbering refers to nucleotide position in coding region of the genes. The gfp gene encodes the Green Fluorescence Protein (GFP) from Aequorea victoria. The six consecutive CAC codons encode a polyhistidine tag. The Sad site codes for Glutamic acid-Leucine amino acids. GenBank. InfB GenelD: 5586604. gfp GenelD 193239069. Figure 4
Figure 4 is a representation of SDS-PAGE of expression of InfB(l-x)-InfA by BL21(DE3) pLysS pET15b cells. Samples were taken after 3.5 hours of 37° C growth with 0.8 mM IPTG. Commassie stained. Lanes: 1. InfA, 2. In f B(IS)-In fA, 3. InfB{l-12)-InfA, 4. InfB{l-15)-InfA, 5. InfB{l-18)-InfA, 6. InfB{l-21)-InfA, 7. InfB{l-90)-InfA, 8. InfB{l-120)-InfA, 9. InfB{l-177)-InfA, 10. InfB{l-282)-InfA, and 11. In fB{l-471)-In fA. 12. Proten molecular mass marker.
Figure 5
Figure 5 is a representation of a Western blot of expression of InfB(l-x)-InfA by BL21(DE3) pLysS pET15b cells. Primary antibody used was targeted against IFl.
Lanes: 1. InfA, 2. InfB(l-9)-InfAf 3. InfB{l-12)-InfA, 4. InfB(l-15)-InfA, 5.
InfB(l-18)-InfA, 6. InfB{l-21)-InfA, 7. InfB(l-90)-InfA, 8. InfB(l-120)-InfA, 9.
InfB{l-177)-InfA, 10. InfB{l-282)-InfA, 11. InfB{l-471)-InfA, and 12. IFl marker protein.
Figure 6
Figure 6 is a representation of SDS-PAGE of expression of InfB(l-x)-gfp by (A) BL21 pBAD and (B) BL21(DE3) pLysS pET24d. Samples were taken after 3.5 hours of 37° C growth with (A) 0.02 % L-arabinose and (B) 0.8 mM IPTG. Commassie stained. Lanes: 1. Un-induced cells, 2. gfp, 3. InfB(l-9)-gfp, 4.
InfB(l-12)-gfpf 5. InfB(l-15)-gfp, 6. InfB(l-18)-gfp, 7. InfB(l-21)-gfpf 8. InfB(l- 90)-gfp, 9. InfB{l-120)-gfp, 10. InfB{l-177)-gfp, 11. InfB{l-282)-gfp, and 12. InfB{ 1-471) -gfp.
Figure 7
Figure 7 shows fluorescence from GFP moiety of fusion proteins expressed from InfB(l-x)-gfp in (A) BL21 pBAD and (B) BL21(DE3) pLysS pET24d. "÷" represents background. The fluorescence count of the gfp construct for each expression system is set to 1. A and B are therefore not directly comparable. Figure 8
Figure 8 shows insoluble protein fractions of expression of InfB(l-x)-gfp from BL21 pBAD grown with 0.02 % L-arabinose at (A) 37° C and (B) 30° C. The bands observed in all lanes are chromosomal gene products. Insoluble proteins of interest are marked with arrows. Lanes: 1. gfp, 2. InfB{l-9)-gfp, 3. InfB{l-12)- gfp, 4. InfB{l-15)-gfp, 5. InfB{l-18)-gfp, 6. InfB{l-21)-gfp, 7. InfB{l-90)-gfp, 8. InfB{l-120)-gfp, 9. InfB{l-177)-gfp, 10. InfB{l-282)-gfp, 11. InfB{l-471)-gfp.
Figure 9 Figure 9 shows insoluble protein fractions from expression of InfB(l-x)-gfp by BL21(DE3) pET24d pLysS with 0.8 mM IPTG at (A) 37° C, (B) 30° C and (C) 20° C. The bands observed in all lanes are chromosomal gene products. Lanes: 1. gfp, 2. InfB{l-9)-gfp, 3. InfB{l-12)-gfp, 4. InfB{l-15)-gfp, 5. InfB{l-18)-gfp, 6. InfB{l-21)-gfp, 7. InfB{l-90)-gfp, 8. InfB{l-120)-gfp, 9. InfB{l-177)-gfp, 10. InfB{l-282)-gfp, 11. InfB{ 1-471) -gfp.
Figure 10
Figure 10 is a representation of SDS-PAGE of expression of InfB(l-x)-gfp by (A) BL21 pBAD and (B) BL21(DE3) pLysS pET24d. Samples were taken after 3.5 hours of growth with (A) 0.02 % L-arabinose at 30° C and (B) 0.8 mM IPTG at 20° C. Commassie stained. Lanes: 1. gfp, 2. InfB{l-9)-gfp, 3. InfB{l-12)-gfp, 4. InfB{l-15)-gfp, 5. InfB{l-18)-gfp, 6. InfB{l-21)-gfp, 7. InfB{l-90)-gfp, 8. InfB{l-120)-gfp, 9. InfB{l-177)-gfp, 10. InfB{l-282)-gfp, 11. InfB{l-471)-gfp.
Figure 11
Figure 11 shows fluorescence from GFP moiety of fusion proteins expressed from InfB(l-x)-gfp in (A) BL21 pBAD at 30 ° C and (B) BL21(DE3) pLysS pET24d at 20 0 C. "÷" represents the background fluorescence. The fluorescence count of the gfp construct for each expression system is set to 1. A and B are therefore not directly comparable. Figure 12
Figure 12 shows the purification of protein expressed from (1) InfB(l-9)-gfp (2) InfB(l-12)-gfp (3) InfB(l-15)-gfp (4) InfB(l-18)-gfp (5) InfB(l-21)-gfp constructs by IMAC on a 8 ml_ column of chelating sepharose, charged with Ni2+. Relevant fractions were pooled and are shown in the SDS-PAGE.
Figure 13
Figure 13 is a representation of the MALDI-TOF spectra of purified protein from (A) InfB(l-9)-gfp (B) InfB(l-12)-gfp (C) InfB(l-15)-gfp (D) InfB(l-18)-gfp (E) InfB(l-21)-gfp constructs. Mass spectra were acquired by a Voyager-DE PRO instrument. Final spectra obtained by averaging 150-200 single shot spectra which were externally calibrated with Bovine Insulin (5734,6 Da) and E.coli Thioredoxin (11674,5 Da).
Figure 14
Figure 14 is a representation of the values obtained from all spectra. The theoretical mass is calculated from the assumption that the N-formyl-methionine was cleaved off.
Figure 15
Figure 15 is a representation of the MALDI-TOF zoomed picture of [M + H+] ion of protein expressed from (A) InfB{l-9)-gfp and (B) InfB{l-21)-gfp. The peak with the lowest molecular weight (1) is the analyzed protein without an N-formyl- methionine group. The middle peak (2), seen only clearly in A, is the protein with the methionine residue still attached. The highest molecular weight peak (3) is the protein with the formyl-methionine group attached.
Figure 16
Figure 16 is a graphical representation of the folded mRNA sequences of 5'-end of pET15b vector transcript (87 nt) followed by the first 50 nucleotides of /nfβ(l-x)- InfA. The online available MFOLD program was used. SD = Shine-Dalgarno sequence. Start = Initiation start codon. dG is in units of kcal/mol. The table summarizes the dG values.
Figure 17 Figure 17 is a graphical representation of the folded mRNA sequences of 5'-end of pBAD vector followed by 50 first nucleotides of InfB(l-x)-gfp. Folding was done by the MFOLD program. dG is in units of kcal/mol. The table summarizes the dG values.
Figure 18
Figure 18 shows the coding sequences of Trx and InfB having expressivity tag properties. Matching bases are marked in bold.
Figure 19 Figure 19 shows the fluorescence of GFP moiety of fusion proteins expressed from InfB(l-21)-gfp with point mutations at position (A) 7 (B) 13 and (C) 16. Fluorescence is plotted with the frequency of the codon which the mutation caused. Wildtype bases are indicated with an asterix and set to a fluorescence count of 1. The frequency of the codon plotted at the 2nd y-axis is obtained from a codon table of 8087 coding sequences in the E.coli genome. The fluorescence of gfp is represented with two dotted lines, each being the limits of the standard deviation on the dataset. Constructs were expressed in BL21 pBAD for 3.5 hours at 30° C with 0.02 % L-arabinose.
Figure 20
Figure 20 shows the numbers of A/U residues and CA repeat bases have been reported to be of importance for the efficiency of translation, and are shown for sequences of InfB, InfA and gfp. Figure 21
Figure 21 shows Thermus thermophilus 70S ribosome, fMet-tRNA and mRNA in complex, pdb entry: 2HGR . A sphere of rRNA bases is shown around the mRNA at a 16 Anstrom radius in the downstream tunnel.
Figure 22
Figure 22 shows the possible interaction between rRNA position C1397 and mRNA position +7.
Figure 23
Figure 23 shows the mRNA sequences of InfB(l-21), InfA(l-21) and g/p(l-21). Possible Watson-Crick base pair interactions between mRNA position +7 to rRNA C1397 are marked in bold.
Figure 24
Figure 24 shows the expression of LacZ gene as amount of β-galactosidase fusion protein in percentage of total protein synthesized. The codon producing the most protein is defined as 1, in both studies this was AAA (Lys).
Figure 25
Figure 25 shows primers used for making InfB(l-x)-InfA and InfB(l-x)-gfp and InfB(l-21)-gfp point mutants, constructs are listed in table I and II. Annealing temperatures for each PCR reaction is listed in the c°-column.
Figure 26
Figure 26 shows primers used for making InfB(l-21)-gfp (Table III). Screening primers are listed in table IV. Annealing temperatures for each PCR reaction is listed in the c°-column. Figure 27
Figure 27 shows the relative expression of the Thio6-GFP system compared to the infB(l-21)-gfp system, measured by the fluorescence of the expressed reporter protein.
Figure 28
Figure 28 shows the expression levels of point mutations of infB(l-21)-GFP relative to infB(l-21)-GFP. The mutated nucleic acids and the resulting new codons are indicated below each graph.
Figure 29
Figure 29 shows an increased expression of proteins when these proteins comprise the infB(l-21)-tag compared to control proteins. (A) Streptavidin +/- infB(l-21)-tag. (B) L32 +/- infB(l-21)-tag. For more information see example 9.
Figure 30
Figure 30 shows sequences and characteristics of the used primers for amplification of infB(l-21)-l32, InfB(l-21)-streptavidin, 132 and streptavidin. These constructs are used in example 9.
Primers listed in figures 25, 26 and 30 are also listed in the sequence list as SEQ ID NO: 92-161.
The present invention will now be described in more detail in the following.
Detailed description of the invention
Definitions
Prior to discussing the present invention in further details, the following terms and conventions will first be defined : Expressivity tag
The term "expressivity tag" refers to nucleic acid sequences or amino acid sequences which increase the expression and thus the yield of a polypeptide to which the tag is operably linked.
Operably linked
The term "operably linked" refers to the connection of elements being a part of a functional unit such as a gene or an open reading frame. Accordingly, by operably linking a promoter to a nucleic acid sequence encoding a polypeptide the two elements becomes part of the functional unit - a gene. The linking of the promoter to the nucleic acid sequence enables the transcription of the nucleic acid sequence directed by the promoter. By operably linking two heterologous nucleic acid sequences encoding a polypeptide the sequences becomes part of the functional unit - an open reading frame encoding a fusion protein comprising the amino acid sequences encoding by the by the heterologous nucleic acid sequences. By operably linking two amino acids sequences, the sequences become part of the same functional unit - a polypeptide. Operably linking two heterologous amino acid sequences generates a hybrid (fusion) polypeptide.
Recombinant
Recombinant gene, promoter, nucleic acid sequence (DNA), amino acid sequences (polypeptide) refers to the products generated by genetic engineering such as the combination or insertion of one or more nucleic acid sequences, thereby combining nucleic acid sequences that would not normally occur together (recombinant nucleic acid sequence, recombinant polynucleotide). Accordingly, recombinant expression refers to the expression of a RNA transcript and/or polypeptide from a coding region directed by a heterologous promoter or a promoter comprising heterologous promoter/enhancer elements.
Heterologous sequences
The term "heterologous" sequences refer to sequence elements of different sequences species. Thus, the fusion of nucleic acid sequences of one species of nucleic acids sequence to nucleic acid sequences of another species of nucleic acids generates a recombinant fusion nucleic acid sequence comprising heterologous sequences. Likewise, the fusion of amino acid sequences of one species of polypeptide to amino acid sequences of another species of amino acid sequences generates a recombinant fusion amino acid sequences comprising heterologous sequences. Accordingly, elements originating from the same species of nucleic acids sequence or polypeptide are not considered heterologous according to the invention. Thus, according to the invention a nucleic acid sequence originating from InfB encoding an expressivity tag and a 5' subsequence of InfB encoding a amino acid sequences are not heterologous sequences according to the invention. The molecule formed by operably linking genetic/protein elements of heterologous genetic/protein species is referred to as a "chimera", "chimeric", fusion, or hybrid molecule.
Hybrid nucleic acid sequence
"Hybrid nucleic acid sequence" refers to a nucleic acid sequence (polynucleotide) comprising at least two heterologous nucleic acid sequences. The molecule is also referred to as "chimeric" nucleic acid sequence.
Hybrid polypeptide
"Hybrid polypeptide" or "hybrid amino acid sequences" refer to an amino acid sequence (polypeptide) comprising at least two heterologous amino acid sequences. The molecule is also referred to as "chimeric" polypeptide.
Sequence identity
The term "sequence identity" indicates a quantitative measure of the degree of homology between two amino acid sequences or between two nucleic acid sequences of equal length. If the two sequences to be compared are not of equal length, they must be aligned to give the best possible fit, allowing the insertion of gaps or, alternatively, truncation at the ends of the polypeptide sequences or
(Nref-Nd,f)l00 nucleotide sequences. The sequence identity can be calculated as Nref , wherein Ndif is the total number of non-identical residues in the two sequences when aligned and wherein Nref is the number of residues in one of the sequences. Hence, the DNA sequence AGTCAGTC will have a sequence identity of 75% with the sequence AATCAATC (Ndif=2 and Nref=8). A gap is counted as non-identity of the specific residue(s), i.e. the DNA sequence AGTGTC will have a sequence identity of 75% with the DNA sequence AGTCAGTC (Ndif=2 and Nref=8). With respect all embodiments of the invention relating to nucleotide sequences, the percentage of sequence identity between one or more sequences may also be based on alignments using the clustalW software
(http:/www.ebi .ac.uk/clustalW/index.html) with default settings. For nucleotide sequence alignments these settings are: Alignment=3Dfull, Gap Open 10.00, Gap Ext. 0.20, Gap separation Dist. 4, DNA weight matrix: identity (IUB). Alternatively, and as illustrated in the examples, nucleotide sequences may be analysed using programme DNASIS Max and the comparison of the sequences may be done at http://www.paraliqn.org/. This service is based on the two comparison algorithms called Smith-Waterman (SW) and ParAlign. The first algorithm was published by Smith and Waterman (1981) and is a well established method that finds the optimal local alignment of two sequences. The other algorithm, ParAlign, is a heuristic method for sequence alignment. Default settings for score matrix and Gap penalties as well as E-values were used.
Vector
The term "vector" refers to a DNA molecule used as a vehicle to transfer recombinant genetic material into a host cell. The four major types of vectors are plasmids, bacteriophages and other viruses, cosmids, and artifical chromosomes. The vector itself is generally a DNA sequence that consists of an insert (a heterologous nucleic acid sequence, transgene) and a larger sequence that serves as the "backbone" of the vector. The purpose of a vector which transfers genetic information to the host is typically to isolate, multiply, or express the insert in the target cell. Vectors called expression vectors (expression constructs) are specifically adapted for the expression of the heterologous sequences in the target cell, and generally have a promoter sequence that drives expression of the heterologous sequences. Simpler vectors called transcription vectors are only capable of being transcribed but not translated : they can be replicated in a target cell but not expressed, unlike expression vectors. Transcription vectors are used to amplify the inserted heterologous sequences. The transcripts may subsequently be isolated and used in as templates suitable in vitro translations systems. Functional equivalent
The term "functional equivalent" relates in the present context to a nucleic acid sequence or polypeptide sequence which pertains substantially the same activity as nucleic acid sequence or polypeptide sequence to which it is compared. The same functional activity (or functional equivalent) may be an activity of at least 25%, such as at least 40%, such as at least 60 %, such as at least 80% of the sequence to which it is compared. It is of course to be understood that a functional equivalent may have a higher activity such as at the most 400%, such as at the most 350%, such as at the most 300%, such as at the most 200%, such as at the most 175%, such as at the most 150%, or such as at the most 125% when compared to a sequence. In the present context the same functional activity as an expressivity tag can be measured by comparing the expression levels of the protein or peptides to which the tags are operably linked. An example of such measurement can be seen in Figure 28 and the corresponding text.
Single chain antibody
In the present context "single chain antibody", "single domain antibody", "sdAb" (also called Nanobody) is an antibody fragment consisting of a single monomeric variable antibody domain. Like a whole antibody, it is able to bind selectively to a specific antigen. Single domain antibodies are much smaller than common antibodies.
Hybrid nucleic acid sequence
In one aspect the invention relates to a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide, said nucleic acid comprising a first nucleic acid sequence of maximum 99 nt encoding a polypeptide operably linked to the a second heterologous nucleic acid sequence encoding a polypeptide, wherein said first nucleic acid sequence comprises a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b). Length of the first nucleic acid sequence
The size of the first nucleic acid sequence is preferably shorter while maintaining an effect as an expressivity tag.
Thus, in one embodiment the maximum length of the first nucleic acid sequence is 90 nucleotides, such as 80 nucleotides, 70 nucleotides, 60 nucleotides, 50 nucleotides, 40 nucleotides, 30 nucleotides, or such as 24 nucleotides,
Substitution may be introduced in the nucleic acid sequences of the invention provided that the modified nucleic acid sequence is a functional equivalent of the nucleic acid sequences from which it originates. Thus, the modified nucleic acid sequence encodes an expressivity tag. The modification of the coding sequences may be advantageous in order to optimize codons of the coding region for optimal translation in the host of production. Moreover, minor modification of the nucleic acid sequences may generate restrictions site, which facilitate the cloning of the sequences.
Thus, in another embodiment the first nucleic acid sequence has a sequence identity of at least 80% such as at least 85%, at least 90%, at least 95% or such as 100% to a amino acid sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4.
Different nucleic acid sequences which are subsequences of one or more of the sequences of the invention may also be used to obtain an effect similar to the ones disclosed in this invention. Especially sequences having a length between 15 and 99 nucleotides are likely to have an effect similar to the claimed sequences.
Therefore, in yet another embodiment the sub-sequences of SEQ ID NO: 1 have a length of at least 23 nucleotides, such as 33 nucleotides, 43 nucleotides, 53 nucleotides, 63 nucleotides, 73 nucleotides, or such as 83 nucleotides.
It is to be understood that these subsequences are functional equivalent to SEQ ID NO: 1.
Examples of second heterologous nucleic acid sequences encoding polypeptides which could be operably linked to the first nucleic acid sequences are GFP, streptavidin, single chain antibodies such as L32 and MMP's. See e.g. examples 3, 9 and 10 and corresponding figures. By having increased expression of GFP more sensitive reporter systems may be constructed. Similar, streptavidin is used in many assys, especially in vitro assays, thus by having a high expression level cheaper assays may be constructed. Thus, in an embodiment according to the invention the second heterologous nucleic acid sequence is selected from the group consisting of reporter constructs such as GFP, streptavidin, MMP's and single chain antibodies such as L32. These are mere examples of nucleic acids (or proteins) which could be used in connection with the present invention. The person skilled in the art could come up with other relevant proteins.
Expressivity tag
To increase the expression of the hybrid polypeptide it is advantageous if the first nucleic acid sequence is coding for an expressivity tag.
Therefore, in an embodiment of the invention, the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, wherein said first nucleic acid sequence encodes an expressivity tag.
Length of expressivity tag A minimum length is required of the first nucleic acid sequence of the expressivity tag to obtain an increased expression of the hybrid nucleic acid sequence, compared to when an expressivity tag is not operably linked to a second nucleic acid sequence.
Therefore, in another embodiment the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, wherein said first nucleic acid sequence of c) has a minimum length of 15 nucleotides. This minimum length could e.g. also be 16 nucleotides, such as 17 nucleotides, or such as 18 nucleotides.
Position of Expressivity tag
To maximize the effect of the expressivity tag of the invention, the tag can be positioned in different positions in relation to the rest of the expressed polypeptide. Thus, in a further embodiment the invention relates to a recombinant hybrid nucleic acid sequence of the invention, wherein said first nucleic acid sequence is positioned 5' to the second heterologous nucleic acid sequence.
Preferably the first nucleic acid sequence is positioned in such a way that the expressivity tag will be positioned in the N-terminal end of the expressed fusion polypeptide. Similar, this means that preferably the first nucleic acid sequence is positioned at the 5'-end of the hybrid nucleic acid sequence.
Codon preference
It has been shown, that proteins that are expressed in high amounts within the cell, have a statistical bias towards the codons that are overall more frequently used within a genome. The different tRNA species does not exist in equal concentrations within the cell. Actually, the concentration of each tRNA species has been shown to correlate with its cognate codon usage. This relationship is thought to be a result of a optimized translation machinery utilizing a balanced codon to tRNA ratio to maximize cell growth. It is also thought to serve to regulate the rate at which different transcripts are translated. When a gene is introduced into E. coli from a foreign organism, its codon context might be unfavorable for high level expression. In general, the more codons a gene contains that are rarely used in the expression host, the less likely it is that the protein will be expressed at reasonable levels. Since optimal expression of a protein can depend on the codon preference of the host organism selected to express the hybrid nucleic acid sequence, it is advantageous to optimize the nucleic sequences to the host.
Thus, in yet a further embodiment the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, wherein said second heterologous nucleic acid sequence has been modified to reflect the codon preference of a host organism selected to express the nucleic acid molecule.
Nucleic acid comprising tag
The hybrid nucleic acid sequence may further comprise one or more additional elements which could be advantageous for detection, solubility or affinity purification of the expressed hybrid polypeptide. Thus, in an additional embodiment the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, further comprising a nucleic acid sequence encoding a polypeptide tag for detection, solubility or affinity purification.
In a further embodiment, this/these one or more tag-elements is positioned in the hybrid nucleic acid sequence depending on the precise use and functionality of the element. It/they are e.g. positioned 3' to the first nucleic acid sequence, between the first and second nucleic acid sequences or 3' to the second nucleic acid sequence. In special cases it/they are positioned 5' to the first nucleic acid sequence.
A wide-range of different elements may be employed for detection, solubility or affinity purification. A skilled person will be able to choose the tag which is optimal for his/her preferred use. The skilled person will furthermore, be able to choose both primary and secondary antibodies which can be used for either purification or detection.
Thus, in yet an additional embodiment the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, wherein said tag for affinity purification, solubility or detection is selected from the group consisting of BCCP-tag, c-myc-tag, calmodulin-tag (SEQ ID NO:33), FLAG-tag (SEQ ID NO:45), HA-tag (SEQ ID NO:50), His-tag (SEQ ID NO:29), Maltose binding protein-tag (SEQ ID NO:27), Nus-tag (SEQ ID NO:60), Glutathione-S-transferase-tag (SEQ ID NO:31), Green fluorescent protein-tag (SEQ ID NO:62), Thioredoxin-tag (SEQ ID NO:64), S-tag (SEQ ID NO:41), Strep-tag (SEQ ID NO:35), human protein C tag (SEQ ID NO:37), Chitin binding protein tag (SEQ ID NO:39), T7-tag (SEQ ID NO:43), Myc-tag (SEQ ID NO:52), V5-tag, VSV-tag, Avi.tag (SEQ ID NO:56), BioEase-tag (SEQ ID NO: 58), SNAP-tag, FlaSH-tag, Nus A-tag, DsbA-tag (SEQ ID NO:66) and the recombinant nucleic acid sequence itself.
Spacer sequence
In certain applications it is advantageous to position a spacer sequence somewhere in the hybrid nucleic acid sequence. Thus, in yet an embodiment the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, further comprising a nucleic acid sequence encoding a spacer positioned between said first and said second nucleic acid sequence.
In special cases the first nucleic acid sequence may be divided into two or more sections using one or more spacer sequences.
Thus, in yet an embodiment the invention relates to a recombinant hybrid nucleic acid sequence according to the invention, further comprising a nucleic acid sequence encoding a spacer positioned between the first and second codon of said first nucleic acid sequence and operably linking said codons.
Introduction of one or more spacer sequences is preferred for example if there is a need for a specific restriction site in either the nucleic acid sequence or the polypeptide sequence. A restriction at the DNA level may be useful for cloning purposes. Restriction sites can e.g. be cleaved using either restriction or nicking enzymes. Cleavage at the peptide level may be useful if part of the hybrid amino acid sequence is not desired in the final peptide. For example a sequence encoding for specific proteases can be introduced between the first and second nucleic acid sequences. In this way the encoded first amino acid sequence can be removed from the final peptide. This removal can take place both before or after a purification step. Examples of useful proteases are Factor Xa (fXa), TEV, PreScission Protease.
Introduction of one or more spacer sequences may also be favourable to avoid structural interference between the first and second amino acid sequences in the hybrid polypeptide.
Recombinant hybrid polypeptide
In a preferred embodiment, the hybrid nucleic acid sequence of the invention is transcribed generating a transcript from which the hybrid polypeptide is translated. Accordingly, one aspect the invention relates to a recombinant hybrid polypeptide, said polypeptide comprising a first polypeptide amino acid sequence and a second heterologous amino acid sequence, wherein said first amino acid sequence of maximum 33 amino acids is selected from the group consisting of a) an amino acid sequence selected from the group consisting of SEQ ID
NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11 and SEQ ID NO: 12, and b) an amino acid sequence having at least 75% sequence identity to a), and c) an amino acid sequence which is a sub-sequence of a) and b).
The maximum length of the first amino acid sequence could be shorter while maintaining an effect as an expressivity tag.
Length of first amino acid sequence
Thus, in an embodiment the invention relates to a maximum length of the first amino acid sequence being 30 amino acids, such as 26 amino acids, 23 amino acids, 20 amino acids, 17 amino acids, 14 amino acids, 11 amino acids, or such as 9 amino acids.
Sequence identity of amino acid sequence
It may also be possible to change some of the amino acids in the sequences of the invention and still maintain an effect as an expressivity tag. Substitution may be introduced in the amino acids in the sequences of the invention provided that the modified amino acid the sequence is a functional equivalent of the amino acid the sequence from which it originates. Thus, the modified amino acid the sequence is an expressivity tag.
Thus, one embodiment the invention relates a first amino acid sequence having a sequence identity of at least 80% such as at least 85%, at least 90%, at least 95% or such as 100% to a amino acid sequence selected from the group consisting of SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11 and SEQ ID NO: 12. Subsequences
Different amino acid sequences which are subsequences of one or more of the sequences of the invention may also be used to obtain an effect similar to the ones disclosed in this invention. Especially sequences having a length between 4 and 33 amino acids are likely to have an effect similar to the claimed sequences.
Therefore, in yet a further embodiment the sub-sequences of SEQ ID NO: 5 and SEQ ID NO:6, are at least 9 amino acids, such as 15 amino acids, or such as 21 amino acids, or such as 27 amino acids.
It is to be understood that these subsequences are functional equivalent to SEQ ID NO:5 and SEQ ID NO:6.
Removal of methionine In bacteria, protein synthesis initiates with a formyl-methionine. The formyl group is removed post-translationally by a peptide deformylase (PDF). The methionine residue is, depending on the identity of the second amino acid in the peptide chain, removed by a methionine aminopeptidase (MAP).
Accordingly, one embodiment of the present invention relates to an isolated hybrid polypeptide comprising a first polypeptide amino acid sequence selected from the group consisting of SEQ ID NO:6, SEQ ID NO:8, SEQ ID NO: 10, SEQ ID NO: 12. The first polypeptides of the embodiment reflect the removal of the N- terminal methionine from the corresponding first polypeptides SEQ ID NO: 5, SEQ ID NO:7, SEQ ID NO:9, SEQ ID NO: 11 respectively.
Expressivity tag
To increase the expression of the hybrid polypeptide it is advantageous if the first amino acid sequence is coding for an expressivity tag. Therefore, in an embodiment of the invention, the invention relates to a recombinant hybrid polypeptide according to the invention, wherein said first amino acid sequence is an expressivity tag. Minimum length of expressivity tag
A minimum length is naturally required of the amino acid sequence of the expressivity tag to obtain an increased expression of the hybrid polypeptide sequence, compared to when an expressivity tag is not operably linked to a second nucleic amino acid sequence.
Therefore, in another embodiment the invention relates to a recombinant hybrid polypeptide according the invention, wherein said first amino acid sequence of c) has a minimum length of 4 amino acids. This minimum length could e.g. also be 5 amino acids, such as 6 amino acids, such as 7 amino acids, such as 8 amino acids, such as 9 amino acids, or such as 10 amino acids. It is to be understood that the minimum lengths cited above relates to said first amino acid sequence when the N-terminal methionine has been removed, as described above.
Positioning of expressivity ag To obtain the full effect of the expressivity tag of the invention, correct positioning of the tag in relation to the rest of the polypeptide to be expressed may be relevant.
Thus, in a further embodiment the invention relates to a recombinant hybrid polypeptide according to the invention, wherein said first amino acid sequence is positioned N-terminal to the second heterologous amino acid sequence.
Preferably the first amino acid sequence is positioned in such a way that the expressivity tag will be positioned at the N-terminal end of the expressed fusion polypeptide.
Tag on amino acid sequence
The hybrid amino acid sequence may further comprise one or more additional elements which could advantageous for detection, solubility or affinity purification of the expressed hybrid polypeptide.
Thus, in an additional embodiment the invention relates to a recombinant hybrid polypeptide according to the invention, further comprising a polypeptide tag for detection, solubility or affinity purification. This/these one or more tag-elements can be positioned in the hybrid amino acid sequence depending on the precise use and functionality of the element. Accordingly, the tags for detection, solubility or affinity purification may be positioned C-terminal to the first amino acid sequence, between the first and second amino acid sequences or C-terminal to the second nucleic acid sequence. In particular embodiments the one or more tag-elements are positioned N- terminal to the first amino acid sequence
Many different elements can be used for detection, solubility or affinity purification. A skilled person will be able to choose the tag which is optimal for his/her preferred use. The skilled person will furthermore, be able to choose both primary and secondary antibodies which can be used for either purification or detection.
Thus, in an additional embodiment relates to a recombinant hybrid polypeptide sequence according to the invention, wherein said tag for affinity purification, solubility or detection is selected from the group consisting of BCCP-tag, c-myc- tag, calmodulin-tag (SEQ ID NO:34), FLAG-tag (SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48 and SEQ ID NO:49), HA-tag (SEQ ID NO:51), His-tag (SEQ ID NO:30), Maltose binding protein-tag (SEQ ID NO:28), Nus-tag (SEQ ID NO:61), Glutathione-S-transferase-tag (SEQ ID NO:32), Green fluorescent protein-tag (SEQ ID NO:63), Thioredoxin-tag (SEQ ID NO:65), S-tag (SEQ ID NO:42), Strep- tag (SEQ ID NO:36), human protein C tag (SEQ ID NO:38), Chitin binding protein tag (SEQ ID NO:40), T7-tag (SEQ ID NO:44), Myc-tag (SEQ ID NO:53), V5-tag (SEQ ID NO:54), VSV-tag (SEQ ID NO:55), Avi.tag (SEQ ID NO:57), BioEase-tag (SEQ ID NO:59), SNAP-tag, FlaSH-tag, Nus A-tag, DsbA-tag (SEQ ID NO:67) and the recombinant polypeptide sequence itself.
Vector
For recombinant expression of the hybrid polypeptide the recombinant hybrid nucleic acid of the invention may be incorporated in a vector for efficient transfer to and expression of the recombinant gene in a suitable host cell. Therefore in a further aspect the invention relates to a vector comprising a first nucleic acid sequence of maximum 99 nt, said first sequence comprising a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b), said vector.
In an embodiment the maximum length of the first nucleic acid sequence comprised in a vector is 90 nucleotides, such as 80 nucleotides, 70 nucleotides, 60 nucleotides, 50 nucleotides, 40 nucleotides, 30 nucleotides, or such as 24 nucleotides,
Thus, in another embodiment the first nucleic acid sequence has a sequence identity of at least 80% such as at least 85%, at least 90%, at least 95% or such as 100% to a amino acid sequence selected from the group consisting of SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4.
In a particular embodiment the sub-sequences of SEQ ID NO: 1 comprised in a vector have a length of at least 23 nucleotides, such as 33 nucleotides, 43 nucleotides, 53 nucleotides, 63 nucleotides, 73 nucleotides, or such as 83 nucleotides.
It is to be understood that these subsequences are functional equivalent to SEQ ID NO: 1.
The first nucleic acid sequences of the invention encoding the expressivity tag may be cloned into a vector: The vector comprising the first nucleic acid sequences may serve as a cloning vector for subsequent cloning of the second heterologous nucleic acid sequence encoding the polypeptide of interest. The second heterologous nucleic acid sequence is inserted in a position that operably links the sequence to the first nucleic acid sequences encoding the expression tag. In one embodiment, the cloning vector further comprises a least one promoter element or a complete promoter operably linked to said first nucleic acid sequence.
Thus, in a further embodiment the invention relates to a vector according to the invention, wherein said vector is further adapted for insertion of a second heterologous nucleic acid sequence encoding a polypeptide operably linked to said first nucleic acid sequence.
Such second heterologous nucleic acid sequence could be incorporated into the vector by standard cloning techniques known to the skilled person such as PCR- based techniques.
A vector comprising the second nucleic acid sequence operably linked to the first nucleic acid sequence (the expression tag) encodes a hybrid polypeptide for efficient expression in a suitable host.
Therefore, in a further embodiment the invention relates to a vector comprising a recombinant hybrid nucleic acid sequence according to the invention.
The vector may be further adapted for expression of the hybrid polypeptide by insertion of transcriptional control elements such as a promoter or a promoter including additional promoter/enhancer elements.
Thus, to direct the expression (transcription) of the hybrid nucleic acid sequence inserted in a vector transcriptional control may be inserted thereby generating a expression vector. The use transcriptional control elements depend on the expression system and host. The promoter of the expression vector may comprise several independent sequence elements. These elements may both be elements which can promote expression, elements which can repress expression or a combination of both. In one embodiment, the expression control sequence comprises transcriptional control elements suitable for conditional expression of the recombinant hybrid nucleic acid sequence.
Thus, in a further embodiment the invention relates to a vector according to the invention, wherein said recombinant hybrid nucleic acid sequence is operably linked to an expression control sequence, such as the araBAD promoter system such as the Ara-promoter (SEQ ID NO: 13), Ara-operator 1 (SEQ ID NO: 14), Ara- operator 2 (SEQ ID NO: 15) and the AraC repressor (SEQ ID NO: 16), the lac promoter system such as the Lac Promoter (SEQ ID NO: 17), the Lac-operator (SEQ ID NO: 18) and the Lad repressor (SEQ ID NO: 19), Tac/Trc-promoter system (SEQ ID NO:20), Lambda PL - promoter (SEQ ID NO:21), Lambda PR promoter (SEQ ID NO:22), T7 promoter (SEQ ID NO:23), T3 promoter (SEQ ID NO:24), SP6 promoter (SEQ ID NO:25) and the AmpR promoter (SEQ ID NO:26).
Different types of vectors may be used for incorporation of a tag sequence according to the invention. The choice of tag depends on the specific use of the vector, how to clone the vector and how to express the inserted sequence. Therefore, in yet an embodiment the invention relates to a vector according to the invention, wherein the vector is selected from the group consisting of plasmids, cosmids, phages, bacterial artificial chromosomes (BAC), Phagemids and Pl- derived artificial chromosomes.
Furthermore, the nucleic acid sequences of the invention may also be positioned in oligonucleotides for in vitro translation and/or transcription purposes. The nucleic acid sequences of the invention may also be positioned in PCR-products for cloning purposes and for in vitro transcription and/or translation purposes.
The vectors may both be single stranded or double stranded. Furthermore, they may be linear circular or nicked. Especially when using in vitro translation systems it may be advantageous to use nicked or linear plasmids.
In vitro translation vector
Thus, in yet a further embodiment the invention relates to a vector according to the invention, wherein the vector is an in vitro translation vector.
In vitro translations systems are compatible with most in vivo expression system and vice versa. Therefore it is to be understood that the in vitro vectors can also be used for in vivo expression, and that in vivo expression vectors can be used for in vitro expression. In a particular embodiment the in vitro translation system and or PCR fragments and vectors are selected from the list consisting of:
-Expressway™ Cell-Free Expression Systems from Invitrogen using vectors such as pEXP5-NT/TOPO®, pEXP5-CT/TOPO®,pEXPl-DEST, pEXP2-DEST, pEXP3- DEST, and pEXP4-DEST.
-The EasyXpress Linear Template system from Qiagen were transcription is directly from PCR-fragments followed by translation.
-The EasyXpress system, where no expression vector is specified
-EcoPro™ T7 System from Novagen where circular or linear DNA can be used under control of E. coli or T7 promoters.
-Single Tube Protein® System 3 (STP3) using plasmid DNA or linear DNA under
T7 or SP6 promoter control.
-S30 T7 High-Yield Protein Expression System from Promega where all vectors with a T7 promoter or strong E. coli promoter can be used. -PROTEINscript™ II from Applied Biosystems where PCR fragments, supercoiled
DNA or linear DNA templates with T7 polymerase promoters are used.
Plasmid vector
As described above many different vectors can be used to carry out different aspects of the invention. In some applications a plasmid vector is preferred. Thus, in a specific embodiment the invention relates to a vector according to the invention, wherein said plasmid is selected from the group consisting of TA cloning vectors, Gateway cloning vectors, restriction cloning vectors, Topo cloning vectors, pET vector system and pBAD vector systems.
In one particular embodiment the vector is selected from the group consisting of Charon 4, EZ-Tn5 pMOD-2, EZ-Tn5 pMOD-3, EZ-Tn5 pMOD-4, EZ-Tn5 pMOD-5, LITMUS28i, LITMUS38i, PinPoint Control, PinPoint Xa Control, PinPoint Xa-I, PinPoint Xa-2, PinPoint Xa-3, pACYC177, pACYC184, pACYCDuet-1, pALTER-EXl, pALTER-EX2, pBAC-2cp, pBACgus-2cp, pBAD-DEST49, pBAD-TOPO, pBAD- TOPO/LacZ, pBAD-TOPO/l_acZ/V5-His, pBAD/His A, pBAD/His B, pBAD/His C, pBAD/His/LacZ, pBAD/Myc-His/lacZ, pBAD/Thio, pBAD/Thio-E, pBAD/Thio-TOPO, pBAD/glll A, pBAD/glll B, pBAD/glll C, pBAD/glll Calmodulin, pBAD/myc-His A, pBAD/myc-His B, pBAD/myc-His C, pBAD102/D-TOPO, pBAD102/D/LacZ, pBAD18, pBAD202/D-TOPO, pBAD202/D/LacZ, pBEc-Q, pBEc-SBP, pBEc-SBP-Q, pBEc-SBP-SETl, pBEc-SBP-SETl-Q, pBEc-SBP-SET2, pBEc-SBP-SET2-Q, pBEc- SBP-SET3, pBEc-SBP-SET3-Q, pBEc-SETl, pBEc-SET2, pBEc-SET3, pBEn-SBP- SETIa, pBEn-SBP-SETlb, pBEn-SBP-SETlc, pBEn-SBP-SET2a, pBEn-SBP-SET2b, pBEn-SBP-SET2c, pBEn-SBP-SET3a, pBEn-SBP-SET3b, pBEn-SBP-SET3c, pBEn- 5 SBPa, pBEn-SBPb, pBEn-SBPc, pBEn-SETla, pBEn-SETlb, pBEn-SETlc, pBEn- SET2a, pBEn-SET2b, pBEn-SET2c, pBEn-SET3a, pBEn-SET3b, pBEn-SET3c, pBR322, pBiEx-1, pBiEx-2, pBiEx-3, pCAL-c, pCAL-kc, pCAL-n, pCAL-n-EK, pCAL- n-FLAG, pCCIBAC, pCDF-2 Ek/LIC, pCDFDuet-1, pCMV-GLuc, pCOLADuet-1, pCR2.1, pCRT7-E, pCX, pCX-TOPO, pDEST14, pDEST15, pDEST17, pDEST24,
10 pDONR201, pET-l la, pET-l lb, pET-llc, pET-l ld, pET-14b, pET-15b, pET-16b, pET-17b, pET-19b, pET-20b( + ), pET-21( + ), pET-21a( + ), pET-21b(+), pET- 21c( + ), pET-21d(+), pET-22b( + ), pET-23a(+), pET-23b( + ), pET-23c(+), pET- 23d( + ), pET-24( + ), pET-24a( + ), pET-24b(+), pET-24c( + ), pET-24d(+), pET- 25b( + ), pET-26b( + ), pET-27b( + ), pET-28a(+), pET-28b( + ), pET-28c(+), pET-
15 29a( + ), pET-29b( + ), pET-29c( + ), pET-30 Ek/LIC, pET-30 Xa/LIC, pET-30a(+), pET-30b( + ), pET-30c( + ), pET-31b( + ) DNA, pET-32 Ek/LIC, pET-32 Xa/LIC, pET- 32a( + ) DNA, pET-32b( + ) DNA, pET-32c(+) DNA, pET-33b( + ), pET-39b(+), pET- 3a, pET-3b, pET-3c, pET-3d, pET-40b( + ), pET-41 Ek/LIC, pET-41a( + ), pET- 41b( + ), pET-41c(+), pET-42a( + ), pET-42b(+), pET-42c( + ), pET-43.1 Ek/LIC,
20 pET-43.1a( + ), pET-43.1b( + ), pET-43.1c( + ), pET-44 Ek/LIC, pET-44a( + ), pET- 44b( + ), pET-44c(+), pET-45b( + ), pET-46 Ek/LIC, pET-47b( + ), pET-48b( + ), pET- 49b( + ), pET-50b( + ), pET-51 Ek/LIC, pET-51b( + ), pET-52 3c/LIC, pET-52b( + ), pET-5a, pET-5b, pET-5c, pET-9a, pET-9b, pET-9c, pET-9d, pET-DEST42, pETIOO/D-TOPO, pET100/D/LacZ, pET101/D-TOPO, pET101/D/LacZ, pET102/D-
25 TOPO, pET102/D/LacZ, pET151/D-T0P0, pET151/D/LacZ, pET160-GW/CAT, pET160/GW/D-TOPO, pET161-GW/CAT, pET161/GW/D-T0P0, pET200/D-TOPO, pET200/D/LacZ, pETBIue-1, pETBIue-2, pETDuet-1, pETcoco-1, pETcoco-2, pEpiFOS-5, pFIA T7, pFlK T7, pF3A WG (BYDV), pF3K WG (BYDV), pF4A CMV, pF4K CMV, pF5A CMV-neo, pF5K CMV-neo, pF9A CMV hRluc-neo, pFC7A (HQ),
30 pFC7K (HQ), pFC8A (HaloTag), pFC8K (HaloTag), pFNIOA (ACT), pFNUA (BIND), pFN2A (GST), pFN2K (GST), pFN6A (HQ), pFN6K (HQ), pG5luc, pGEM-1, pGEM- HZf( + ), pGEM-HZf(-), pGEM-13Zf( + ), pGEM-15Zf(-), pGEM-2, pGEM-3, pGEM- 3Z, pGEM-3Zf( + ), pGEM-3Zf(-), pGEM-4 Vector, pGEM-4Z, pGEM-5Zf( + ), pGEM- 5Zf(-), pGEM-7Zf( + ), pGEM-7Zf(-), pGEM-9Zf(-), pGEMEX-1, pGEMEX-2, pGEX-
35 2T, pGEX-2TK, pGEX-3X, pGEX-4T-l, pGEX-4T-2, pGEX-4T-3, pGEX-5X-l, pGEX- 5X-2, pGEX-5X-3, pGEX-6P-l, pGL2-Basic, pGL2-Control, pGL2-Enhancer, pGL2- Promoter, pGLuc-Basic Vector, pGPSl . l, pGPS2.1, pGPS3, pGPS4, pGPS5, pIndigoBAC-5, pKLACl, pKLACl-malE, pKO500, pLacI, pLysE, pLysS, pMAL-c4E, pMAL-c4G, pMAL-c4X, pMAL-p4E, pMAL-p4G, pMAL-p4X, pMAL-pIII, pNEB193, pNEB206A, pNEBR-Rl, pNEBR-Xl, pNEBR-XIGLuc, pNEBR-XlHygro, pRS1274, pRS308, pRS414, pRS415, pRS475, pRS550, pRS551, pRS552, pRS577, pRS591, pRSET A, pRSET B, pRSET C, pRSET T7, pRSET-E, pRSET/LacZ, pRSF-2 Ek/LIC, pRSFDuet-1, pSP64, pSP64 PoIy(A), pSP65, pSP70, pSP72, pSP73, pTWINl, pTXBl, pTXB3, pTYBl, pTYBl l, pThioHis A, pThioHis B, pThioHis C, pTrcHis A, pTrcHis B, pTrcHis C, pTrcHis-TOPO, pTrcHis-TOPO/lacZ, pTrcHis/CAT, pTrcHis2 A, pTrcHis2 B, pTrcHis2 C, pTrcHis2-TOPO, pTrcHis2-TOPO/LacZ, pTrcHis2/LacZ, pUC19, pWEB, and pWEB-TNC.
Host cell For cloning and expression purposes it may be useful to have the vector cloned in a host cell. In this way the vector can be maintained, cloned and purified and thus provide a source for almost unlimited numbers of the vectors. Moreover, stable producer cell lines expressing the hybrid polypeptide of interest may also be advantageous. In order to maintain a producer cell selection marker may be incorporated in the expression vector.
Thus in an embodiment the invention relates to a host cell comprising a vector according to the invention.
As described above bacteria are in many instances optimal host cells for e.g. expression purposes.
Bacteria
Thus, in yet an embodiment the invention relates a host cell according to the invention, wherein said host cell is a bacterium.
Examples of bacterial strains which may be suitable as host cells are Bacillus licheniformis, Bacillus subtilis and E.coli. Bacillus licheniformis and Bacillus subtilis is e.g. cultured in order to obtain enzymes, such as proteases, for use in biological washing powder, and it may increase production if the expressivity tag according to the invention is positioned in frame in front of such proteases. Enzymes produced by B. subtilis and B. licheniformis are widely used as additives in laundry detergents.
It may be preferred that the bacterium used for hosting a vector is an E. coli strain. Thus, in yet an embodiment the invention relates a host cell according to the invention, wherein said host cell is E. coli.
In an embodiment of the invention the host cell is selected among AGl, AB1157, BL21(AI), BL21(DE3), BL21 (DE3) pLysS, BNN93, BNN93 (λgtll), BNN97, BW26434 (CGSC Strain # 7658), B5615, B834, B834 (DE3), B834(DE3) pLysS, BLR(DE3), BLR(DE3) pLysS, C600, C600 hflA150 , Y1073, BNN102, CSH50, D1210, DB3.1, DHl, DH5α, DHlOB, DH12S, DM1ER2566, ER2267, HBlOl,
HMS174(DE3), HT96 NovaBlue, HT96 NovaBlueF-, IJ1126, IJ1127, JM83, JMlOl, JM103, JM105, JM106, JM107, JM108, JM109, JM109(DE3), JMIlO, JM2.300, LE392, Machl, MC1061, MC4100, MG1655, 0mniMAX2, Origami , Origami 2 , Origami B, Rosetta, RosettaBlue, Rosetta-gami, Rosetta-gami 2, Rosetta-gami B, Rosetta(DE3)pLysS, Rosetta-gami(DE3)pl_ysS, RRl, STBL2, STBL3, STBL4, SURE, SURE2, TOPlO, ToplOF1, Tuner, UT5600, W3110, XLl-Blue, XL2-Blue, XL2-Blue MRF1, XLl-Red, XLlO-GoId, and XLlO-GoId KanR.
In a specific embodiment the host cell is E. coli JM109 and E. coli BL21(DE3).
Fungi and molds
In another embodiment the host cells are selected from non-bacterial origin such as fungi or molds. In a more specific embodiment the host cells are selected among Aspergillus niger and Aspergillus oryzae.
Many useful enzymes are produced using industrial fermentation of A. niger. For example, A. niger glucoamylase is used in the production of high fructose corn syrup, and pectinases are used in cider and wine clarification, α-galactosidase, an enzyme that breaks down certain complex sugars, is a component of Beano and other medications which the manufacturers claim can decrease flatulence. Another use for A. niger within the biotechnology industry is in the production of magnetic isotope-containing variants of biological macromolecules for NMR analysis.
A. niger is also cultured for the extraction of the enzymes glucose oxidase (GO) and Alpha-galactosidase (AGS). Glucose oxidase is used in the design of glucose biosensors, due to its high affinity for β-D-glucose. Alpha-galactosidase can be produced by A. niger fermentation; it is used to hydrolyze alpha 1-6 bonds found in melibiose, raffinose, and stachyose. Research published in 2006-2008 investigated A. niger prolyl endoprotease (AN-PEP), a microbial-derived prolyl endoprotease which cleaves gluten. This has strong implications in the treatment of Coeliac (Celiac) disease or other metabolic gluten sensitivity disease processes.
Aspergillus oryzae has strong secretion of amylases (α-Amylase and glucoamylase), some carboxypeptidases and tyrosinases.
Thus, by using the expressivity tag of the invention it may also be possible to increase production of enzymes in host cells of non-bacterial origin.
Antibody The first polypeptide amino acid sequence disclosed in the invention may be used as a target for one or more antibodies. In this way the expressivity tag may also be used for both detection and purification of the hybrid polypeptide using standard molecular techniques, such as affinity purification, and immuno detection.
Therefore, in an embodiment the invention relates to an isolated antibody or antigen-binding fragment thereof which selectively binds to the first polypeptide amino acid sequence of the invention.
Similarly, the invention relates to an antibody, isolated antibody or antigen- binding fragment thereof according to the invention, wherein the antibody is a purification tag or an epitope tag. Use of a recombinant hybrid nucleic acid sequence
In an aspect of the invention, the recombinant hybrid nucleic acid sequences of the invention are used to express recombinant hybrid polypeptides. This may be especially relevant where large amount of protein is required, e.g. for the production of commercial enzymes such as restriction enzymes, therapeutic proteins, drugs, etc. The presence of the expressivity tag increases the yield of the polypeptide of interest during the fermentation process.
Thus, one aspect of the invention relates to the use of a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide according to the invention for the expression of a recombinant polypeptide.
Similarly, the invention relates to the use of a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide according to the invention, wherein said recombinant hybrid nucleic acid sequence is comprised in a vector according to the invention adapted for expression of said recombinant polypeptide.
Method for expressing a recombinant polypeptide
Furthermore, in yet an aspect the invention relates to a method for expressing a recombinant polypeptide comprising a) introducing into an expression system the hybrid nucleic acid sequence according to the invention encoding a recombinant polypeptide, b) expressing said nucleic acid sequence in the expression system, and optionally c) purifying the expressed recombinant polypeptide.
The above disclosed method may be performed by a method, wherein a vector disclosed above is used. Thus, in an embodiment the invention relates to a method according to the invention, wherein the hybrid nucleic acid sequence is comprised in a vector according to the invention is used.
Furthermore, in a further embodiment the invention relates to a method according to the invention, wherein the expression system is a host according to the invention. In addition, the invention relates to a method according to the invention, wherein said expression system is an in vitro expression system.
When performing in vitro translation the nucleic acid sequences of the invention are at some stage RNA. Therefore, it is to be understood that the nucleic acid sequences of the invention can also be RNA sequences.
When performing in vitro translation one or more components are normally obtained from a cell extract. Different bacterial strains can be used as the source, but often the extract composition is obtained from E. coli.
Therefore, in another embodiment, the invention relates to a method according to the invention, wherein the in vitro translation system comprises extract from E. coli.
As means of purification or detection an antibody which binds to the recombinant polypeptide of the invention may be employed. Thus, in an embodiment, the invention relates to a method according to the invention, wherein an antibody according to claims the invention is used for purification of said recombinant polypeptide.
Similarly, in an embodiment, the invention relates to a method according to the invention, wherein an antibody according to the invention is used for detecting said recombinant polypeptide.
As also mentioned earlier, cleavage at the peptide level may be useful if part of the hybrid amino acid sequence is not desired in the final peptide. For example a sequence encoding for specific proteases can be introduced between the first and second nucleic acid sequences. In this way the encoded first amino acid sequence can be removed from the final peptide. This removal can take place both before or after a purification step. Examples of useful proteases are Factor Xa (fXa), TEV and PreScission Protease.
Thus, in yet an embodiment the invention relates to method according to the invention, wherein said first polypeptide amino acid sequence is subsequently removed from the expressed target nucleic acid sequence by proteolytic cleavage. Kit
A vector, as disclosed above, may be very useful in a kit, such as an in vivo translation kit, an in vitro translation kit, and a PCR-kit, a transcription kit.
Thus, in a further embodiment the invention relates to a kit comprising a vector according to the invention.
As means of further optimization of the expressivity tag of the present invention directed evolution using phages may be used. An expressivity tag could be cloned 5' of an essential chromosomal gene of a phage. The gene should be vital for the fitness of the phage. Optimal expressivity tags would be evolved and the fittest phages could be harvested as they would be most abundant. Alternatively, phage display could be used if a selection mechanism could be constructed favoring high expression of the binding protein. Depending on similarity of sequences produced a possible expressivity tag motif could hopefully be defined from such experiments.
Thus, in a further embodiment the invention relates to a method for producing and or optimizing expressivity tags, wherein
1) the nucleic acid sequences encoding an expressivity tag of the invention is inserted upstream of operably linked to an essential gene in a phage
2) the phages are grown under a selection pressure
3) phages with high fitness is selected optionally
4) steps 1) to 3) are repeated
A person skilled in the art of e.g. phage display would know the different important parameters to adjust to provide a selection pressure.
It should be noted that embodiments and features described in the context of one of the aspects of the present invention also apply to the other aspects of the invention. All patent and non-patent references cited in the present application, are hereby incorporated by reference in their entirety.
The invention will now be described in further details in the following non-limiting examples.
Examples
The inventors fused InfB{l-471) 5' of two different reporter genes; InfA and gfp. Deletions were next made from the 3' end of Inf B{1-471) and amounts of protein evaluated by SDS-PAGE, western immunoblotting and quantitative fluorescence measurements. The deletions in InfB(l-471) were placed in regions without protein secondary structure. If the functional part of the expressivity tag was a protein structure, it would by this design not be compromised. Figure 1 shows the secondary structure elements of domain I of IF2 and the 3' InfB{l-471) deletions. Results showed that the expressivity tag region was located on InfB{l-21). Next, 3' deletions in InfB{l-21) was constructed, removing one codon at the time, and the expression analyzed. The results showed that the intact InfB{l-21) sequence contained the expressivity tag, as expression dropped when it was compromised. The expression of InfB(l-21) point mutants fused to gfp was next analyzed to define important positions in InfB(l-21). All the mutants produced compromised the efficient expression of InfB(l-21). No position of specific importance was therefore defined. It could be shown that there was no correlation between expression and codon frequency or number of downstream A/U bases. The mRNA structure in the TIR and possible interactions between mRNA and the ribosome in initiation of translation could be of importance for InfB{l-21).
Materials
Bacterial strains
E. coli JM109 was from New England Biolabs and used for propagation of plasmids.
Genotype: F ' traD36 proA+B+ lacF Δ(lacZ)M15/ Δ(lac-proAB) glnV44 el4~ gyrA96 recAl relAl endAl thi hsdR17
E. coli BL21(DE3) was from Novagen and used for protein expression combined with the pET vector system. The strain is deficient in the Lon protease and lacks the OmpT outer membrane protease which improves the yield of intact recombinant protein. It contains a copy of the T7 polymerase inserted chromosomally with the DE3 lysogen.
Genotype: F", ompT, /7sc/SB(rB~ mB~), gal, dcm, (DE3)
E. coll BL21 was from Novagen and used for protein expression with the pBAD vector system. This strain is like BL21(DE3) also deficient in the Lon protease and lacks the OmpT outer membrane protease.
Genotype: F", ompT, /7sc/SB(rB~ mB~), gal, dcm
Plasmids pET15b was from Novagen and used for both cloning and expression. It contains a bla gene coding for the β-lactamase enzyme giving resistance to ampicillin. Transcription from its multiple cloning site is under control of the T7 promoter, requiring the T7 polymerase.
pET24d was from Novagen and used for both cloning and expression. It contains a kanamycin resistance gene and transcription from its multiple cloning site is controlled by the T7 promoter.
pBAD was from Invitrogen and used for both cloning and expression. It contains an ampicillin resistance gene and transcription from its multiple cloning site is controlled by the araBAD promoter.
pLysS) was from Novagen and used to avoid leaky expression from pET15b and pET24d plasmids. pLysS has a weak constitutive expression of T7 lysozyme which binds to the T7 polymerase inhibiting its transcription activity.
pET15b[IF2-I] was provided by H. P. Sørensen and was used as a template to create the JnfiB( 1-471) fragments.
pGLO was from Bio-Rad and was used as a template to obtain gfp.
pET15b[IF2domI-IFl] was provided by T. B Kallehauge and was used as a template to obtain InfA. Relevant enzymes, antibodies and reagents
Restriction enzymes (βspHI, Sad, Bamftl, EcoRl, Ncol, sphl) was supplied by New England Biolabs® with its relevant reaction buffer. T7 ligase and reaction buffer was from Pharmacia. Plasmids were purified by QIAGENs Qiaprep Spin Miniprep Kit. Restriction digested DNA fragments were purified by the Wizard® SV Gel and PCR clean-up system from Omega. Polyclonal rabbit anti-mouse immunoglobulin HRP conjugated antibody was from DAKO. Monoclonal mouse anti- Escherichia coli IFl antibody no. 1 was produced in the laboratory of biodesign. A complete list of buffer and solutions can be found in Appendix B.
Oligonucleotides and sequencing
All oligonucleotides and sequencings were made by DNA-technology A/S. A complete list of primers is to be found in Figures 25 and 26.
Cloning
Detailed information on PCR amplification, restriction digestion, ligation, transformation, DNA purification and sequence verification can be found in Appendix A. A summary of the cloning procedure will be described.
The various InfB{l-471) 3' deletions fused to InfA were prepared using the primers listed in Figure 25 table I. Each PCR fragment produced is listed with its forward and reverse primer, and is in this cloning description referred to by its PCR fragment name. Figure 2 shows all the constructs cloned. The InfB(l-471) 3' deletion fragments and InfA are linked by a Sad restriction site coding for Glu-Leu amino acids. The initiator ATG of InfA was removed in all constructs where it was fused to any fragment of InfB{ 1-471). When all eleven constructs are mentioned at the same time in the text they will be referred to as InfB{l-x)-InfA.
InfA was first constructed by one PCR fragment which was ligated into a pET15b vector. PCR InfA I was made with a pET15b[IF2-IFl] template which introduced restrictions sites BspHI and BamHI and six consecutive histidines. The BspHl/BamHl digested InfA I PCR fragment was ligated into an Ncol/BamHl digested pET15b plasmid to make InfA. InfB{l-471)-InfA, InfB(l-282)-InfA, InfB{l-177)-InfA and InfB(l-120)-InfA was next constructed by two PCR fragments, a InfB part and a InfA named InfA II, which was ligated into pET15b.
PCR fragments InfB{l-471), InfB{l-282), InfB{l-177) and 7nfβ(l-120) were prepared with a pET15b[IF2-I] template. The primers introduced restriction BspHI and Sad. An InfA II PCR was made with a pET15b[IF2-IFl] template and introduced restrictions sites Sad and BamHI and six consecutive histidines. βspHI/SacI digested 7nfβ(l-471), 7nfβ(l-282), 7nfβ(l-177) and 7nfβ(l-120) was ligated with SacI/BamHI digested InfA II into a Λ/coI/βamHI digested pET15b plasmid to make InfB{l-471)-InfA, InfB(l-282)-InfA, InfB{l-177)-InfA and InfB{l-120)-InfA.
InfB{l-9)-InfA, InfB{l-12)-InfA, InfB{l-15)-InfA, InfB{l-18)-InfA, InfB{l-21)-InfA and InfB(l-90)-InfA were constructed each by one PCR fragment which was ligated into pET15b. PCR fragments 7nfβ(l-9), JnJB(I- 12), Jn/B(l-15), 7nfβ(l-18), 7nfβ(l-21) and 7nfβ(l-90) were prepared with a pET15b[7nfβ (1-471)-Jn/5A] template and introduced restriction sites Sphl and Sad. Sphl/Sacl digested 7nfβ(l-9), InfB{l- 12), 7nfβ(l-15), 7nfβ(l-18), InfB{l-21) and 7nΛB(l-90) were ligated with SacI/βamHI digested 7nfiA II into a Sphl/BamHl digested pET15b plasmid to make InfB{l-9)-InfA, InfB{l-12)-InfA, InfB{l-15)-InfA, InfB{l-18)-InfA, InfB{l-21)-InfA and InfB{l-90)-InfA.
The various 3' 7nΛB(l-471) deletions fused to g/p were prepared using the primers listed in Figure 25 table II. Figure 3 shows all the constructs made. The initiator ATG of gfp was removed in all constructs where it is fused to any fragment of 7nΛB(l-471). When all eleven constructs are mentioned at the same time in the text they will be referred to as InfB{l-x)-gfp.
gfp was first constructed by one PCR which was ligated into pBAD. PCR gfp I was made with a template pGLO plasmid. This introduced restriction sites βspHI and EcoRI along with six consecutive histidines. The βspHI/EcoRI digested PCR fragment was ligated into an NcoI/EcoRI digested pBAD plasmid to make gfp. Next InfB{l-471)-gfp, InfB{l-282)-gfp, InfB{l-177)-gfp and InfB{l-120)-gfp were constructed by two PCR fragments, a InfB part and a gfp part named gfp II, which was ligated into pBAD.
PCR fragments InfB{l-471), InfB{l-282), InfB{l-177), InfB{l-120) were prepared 5 with a pET15b[IF2-I] template. The primers introduced restriction βspHI and Sad. A PCR of gfp II was made with pGLO as template which introduced restriction sites Sad and EcoRI and six consecutive histidines. SacI/EcoRI digested gfp II was ligated with βspHI/Sad digested 7nfβ(l-471), 7nfβ(l-282), 7nfβ(l-177) or InfB(l- 120) into a NcoI/EcoRI digested pBAD plasmid to make InfB{l-471)-gfp, InfB{l-
10 282)-gfp, InfB{l-177)-gfp and InfB{l-120)-gfp.
InfB{l-9)-gfp, InfB{l-12)-gfp, InfB{l-15)-gfp, InfB{l-18)-gfp, InfB{l-21)-gfp were next constructed by one PCR fragment each which was ligated into pBAD. Primers contained a 5' non-annealing overhang of the InfB part along with an annealing gfp part. PCR fragments InfB{l-9)-gfp, InfB{l-12)-gfp, InfB{l-15)-gfp,
15 InfB(l-18)-gfp, InfB(l-21)-gfp were made using pGLO as template. It introduced restriction sites βspHI, EcoRI and six consecutive histidines. The five βspHI/EcoRI digested PCR fragments were ligated into Λ/coI/EcoRI digested pBAD plasmids to make InfB{l-9)-gfp, InfB{l-12)-gfp, InfB{l-15)-gfp, InfB{l-18)-gfp, InfB{l-21)- gfp.
20 The InfB{l-90)-gfp construct was made by 3 PCR reactions. First PCR produced a fragment of 7nfβ(l-90) with a gfp 3' overhang named 7nfβ(l-90) I. Second PCR produced a fragment of gfp named 7nfβ(l-90) II. Finally a PCR named 7nfβ(l-90) III was made using the two PCR fragments as primers, as they annealed to each other. The 7nfβ(l-90) III PCR product was ligated into pBAD. First PCR 7nfβ(l-90)
25 I was made with pET15b[IF2-I] as template which introduced restriction site βspHI and a 3' gfp overhang. Next PCR 7nfβ(l-90) II was made with a pGLO template which introduced restriction site EcoRI and six consecutive histidines. A PCR using 7nfβ(l-90) I and 7nfβ(l-90) II as primers was made which produced a InfB{l-90)-gfp PCR fragment with βspHI and EcoRI restriction sites. The
30 βspHI/EcoRI digested fragment was ligated into an Λ/coI/EcoRI digested pBAD plasmid.
The procedure described for making each construct was followed in cloning them into the pET24d plasmid. A point mutation analysis was made of InfB{l-21)-gfp changing the base pairs at downstream positions +7, +13 and +16 in the InfB{l-21) construct. This was done with the primers in Figure 26 table III. The primers were constructed with an InfB overhang and a gfp annealing part. The constructs were therefore made by 5 one PCR each which was ligated into pBAD. Eight PCRs named InfB(l-21)-gfp 7A, 7T, 7C, 13T, 13G, 16T, 16G, 16C was made with a pGLO template. This introduced restriction sites BspHI and EcoRI along with six consecutive histidines. The BspHI/EcoRI digested PCR fragment was ligated into an NcoI/EcoRI digested pBAD plasmid to make the constructs InfB{l-21)-gfp 7A, 7T, 7C, 13T, 13G, 16T, 10 16G, 16C.
All ligation reactions were transformed into E.coli JM 109 for further propagation. Clones were picked and grown separately for subsequent plasmid purification. A PCR screen was performed with either T7 screen primers or pBAD screen primers
15 (Figure 26 table IV) for verification of the approximate size of an insert. Plasmids with correctly sized inserts were sequenced by DNA-technology A/S using the screening primers. pET15b and pET24d constructs were transformed into BL21(DE3) pLysS cells while pBAD constructs were transformed into BL21 for subsequent expression.
20
Small-scale expression
BL21(DE3) pLysS harboring pET15b or pET24d constructs and BL21 harboring pBAD constructs were grown in 3 ml_ 2xTY media with appropriate antibiotics (cone. 100 μg/mL ampicillin, 34 μg/mL chlorampenicol, 50 μg/mL kanamycin)
25 overnight. 500 μl_ overnight culture was centrifuged (15.000 x g, 1 min), cells resuspended in 150 μl_ 2xTY and inoculated into 50 ml_ 2xTY media with appropriate antibiotics. Cells were grown at 37° C until OD60O reached 0.8. BL21(DE3) pLysS pET15b and pET24d was induced with a final cone, of 0.8 mM IPTG and BL21 pBAD with 0.02 % L-Arabinose. Different growth temperatures
30 were tested to optimize soluble protein expression. Cells were grown at 370C and 3O0C for 3.5 hours and at 2O0C for 16 hours. A sample was taken for SDS-PAGE by the following procedure. 1 ml_ cell culture was centrifuged (15.000 x g, 4 min) and resuspended in OD6oo-50μL disintegration buffer, boiled for 5 min and stored at -200C. The rest of the cell culture was recovered by centrifugation (3.700 x g,
35 25 min, 40C) and cell pellets stored at -2O0C. Western Immunoblotting
Expression of InfB{l-x)-InfA was analyzed by a western blot with anti-IFl antibodies. A sandwich western blot was constructed. On the lower plate (anode) the sandwich was built as following : Filter paper, nitrocellulose paper, SDS-PAGE run with total cell extracts from small-scale overexpression, filter paper. Filter papers were soaked in blotting buffer, nitrocellulose paper in distilled water. Blotting was carried out at 200 mA for 1 hour in a Kem Tec apparatus. After blotting the nitrocellulose membrane was transferred to a rolling glass and washed with 50 mL TBS-T for 5 min. 15 ml_ primary mouse anti-f.co// IFl monoclonal antibody no. 1 (lμg/mL in TBS-T) was added and incubated for 1 hour. The membrane was washed 3x5 min with 50 ml_ TBS-T. 10 ml_ secondary antibody (rabbit anti-mouse HRP-conjugated antiimmunoglobulin, 2.3 μg/mL in TBS-T) was added and incubated 1 hour. The membrane was next washed 6x5 min with 50 mL TBS-T. A mixture of 4 mL solution 1 and 100 μl_ solution 2 of the ECL western blotting kit was incubated with the membrane for 5 min at constant shaking. The membrane was wrapped in Vita Wrap and exposed to an X-ray film for only a few seconds, which gave the clearest picture.
Quantitative fluorescence measurements
Amount of protein expressed from InfB(l-x)-gfp constructs was analyzed, by measuring the fluorescence from a fixed amount of cells.
Frozen BL21(DE3) pLysS pET24d[/nfβ(l-x)-g/p] and BL21 pBAD[/nfβ(l-x)-g/p] cell pellets, previously grown for small-scale expression, were thawed on ice and resuspended in resuspension buffer. The optical density at 600 nm was measured, and cell densities adjusted according to the optimum range of measurement, for the fluorescence reader. Fluorescence counts were later normalized according to cell density. Cell suspensions were pipette in a 96-well tray in 10 % dilution series. The fluorescence was measured in a Synergy HTTR-I reader (Bio-tek), with an excitation wavelength of 400 nm (bandwidth 30) and an emission wavelength of 508 nm (bandwidth 20). Solubility Test
For the GFP protein to be active, it needs to be correctly folded. The amount of GFP inclusion bodies produced was therefore checked. Frozen BL21(DE3) pLysS pET24ά[InfB{l-x)-gfp] and BL21 pBAD[/nfβ(l-x)-g/p] cell pellets, previously grown for small-scale expression, were thawed on ice and resuspended in lysis buffer. The optical density at 600 nm was measured and adjusted to the same numerical value for all cell suspensions. A volume of 3 ml_ cells were sonicated in a 13 ml_ Falcon tube on ice, for a total of 15 min (0.7 cycles, 100 amplitude). Prior optimization of the sonication process, had shown that these conditions gave optimum cell disruption. The temperature of the sample and sonicator tip was checked regularly, to avoid overheating the sample. 750 μl_ disrupted cells were centrifuged (15.000 x g, 15 min, 4°C), and pellets resuspended in 750 μL water, boiled for 5 min and saved for SDS-PAGE analysis.
Large-scale expression
An odd SDS-PAGE migration pattern was observed for some of the constructs, and so protein was expressed, purified by chromatography and checked by mass spectrometry.
BL21(DE3) pLysS pET24d[/nfβ(l-9)-g/p], pET24d[/nfβ(l-12)-g/p], pET24d[/nΛ3(l-15)-g/p], pET24d[/nfβ(l-18)-g/p] and pET24d[/nfβ(l-21)-g/p] cells were grown in 50 ml_ 2xTY over night with antibiotics (34 μg/mL chlorampenicol, 50 μg/mL kanamycin). 16 ml_ culture was centrifuged (3.700 x g, 10 min, 40C), resuspended in 1 ml_ 2xTY and inoculated into 1,6 L 2xTY with antibiotics. Cells were grown at 37° C, until the optical density at 600 nm reached 0.8, at which they were induced with 0.8 mM IPTG. Growth was continued 3.5 hours at 30° C. Cells were harvested by centrifugation (9.500 x g, 10 min, 40C), resuspended in 0.9 % NaCI and centrifuged again (3.700 x g, 25 min, 40C). Cell pellets were stored at -20° C.
Protein purification for mass spectrometry
Fusion proteins expressed from InfB(l-9)-gfp, InfB(l-12)-gfp, InfB(l-15)-gfp, InfB(l-18)-gfp and InfB(l-21)-gfp contained a C-terminal His-tag, and were purified by Immobilized Metal Affinity Chromatography (IMAC). Cells previously grown for expression were resuspended in 2 ml_ buffer A per gram cells. Cell suspensions were sonicated in a 50 ml_ Nunc tube on ice, for a total of 15 min (0,5 cycle, 100 amplitude). Lysed cells were ultracentrifuged (184.000 x g, 75 min, 4° C) to sediment cell debris and insoluble protein. Soluble protein was 0,2 μm filtered and loaded onto a 8 ml_ chelating sepharose column, pre-loaded with Ni2+ and calibrated in Buffer A. Bound protein was eluted by a stepwise 9-15-100 % gradient of Buffer C against Buffer A. Relevant fractions were analysis by SDS-PAGE. Identified fractions were pooled in a dialysis sack (MWCO: 6-8 kDa), and dialysed against 1 L of storage buffer over night at 4°C. Protein was stored at -200C.
Matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF)
Fusion proteins expressed from InfB{l-9)-gfp, InfB{l-12)-gfp, InfB{l-15)-gfp, InfB{l-18)-gfp and InfB{l-21)-gfp constructs and purified as described were analyzed by MALDI-TOF, to confirm their masses. All five types of protein were dialysed against 2 L 50 mM Tris-HCI pH 7.5 overnight, 4°C. 0.5 μl_ of protein (20 pmol/μL) was mixed with 0.5 μl_ sinapinic acid (lOmg/mL) on a MALDI plate. Samples were left to crystallize for 30 min. at room temperature. Mass spectra were acquired by a Voyager-DE PRO (MALDI-TOF) instrument (Applied Biosystems). Final spectra were obtained by averaging 150-200 single shot spectra. Spectra were externally calibrated with bovine insulin (5734.59 Da) and E.coli thioredoxin (11674,48 Da).
N-terminal sequencing
Protein expressed from InfB(l-18)-gfp, InfB(l-21)-gfp and InfB(l-90)-gfp was analyzed by N-terminal sequencing by Prof. Claus Oxvig. Protein was excised from SDS-PAGE bands run with total cellular extracts from the small-scale expression.
mRNA folding by MFOLD
In order to investigate the base pairing in the TIR of mRNA sequences for InfB{l- x)-InfA and InfB{l-x)-gfp constructs, mRNA sequences were folded by the online available program MFOLD.
The following standard settings were applied. Nucleic acid type: RNA Sequence type: Linaer Temperature: 37° C Percent: 5 Na2+ cone: 1.0 Mg2+ cone: 1.0
Maximum bulge/interior loop size: 30 Maximum asymmetry of a bulge/interior loop: 30 Maximum number of foldings to be computed : 50
Codon frequency
Codon frequencies are obtained from kazusa. or.jp/codon/cgi-bin/showcodon. cgi?species=37762
Example 1
Analysis of the expression of InfB(l-x)-InfA
Total cellular extracts of induced cells expressing InfB{l-x)-InfA (figure 2) were analyzed by SDS-PAGE (figure 4).
It was evident, that InfA was expressed at lower levels than either of the InfB{l- x)-InfA constructs. It was not possible to see much difference between any of the InfB deletions, except InfB(l-177)-InfA which seemed to have a slightly higher expression. The molecular weight of the proteins expressed from the InfB(l-9)- InfA, InfB(l-12)-InfA, InfB(l-15)-InfA, InfB(l-18)-InfA did not correspond to the molecular weight, of the surrounding proteins in the SDS-PAGE. All four proteins migrated slower than the protein expressed from InfB{l-21)-InfA and 7nΛB(l-90)- InfA, which should not be the case. All constructs were verified by DNA sequencing, and it thus seemed like an odd migration pattern that had to be addressed. To analyze the integrity of the InfA moiety, and further evaluate amounts of protein expressed, a western blot with anti-IFl antibodies was made of the total cellular extracts (figure 5).
IFl protein was clearly recognized by the primary antibody in all constructs, verifying the integrity of the InfA part. It was quite obvious from the western blot, that InfA expressed less than either of the fusion constructs. InfB{l-177)-InfA produced the most intensive band, as in the SDS-PAGE, but it was not much more intense than the band produced by InfB{l-90)-InfA. It seemed that a fragment as small as 7nΛB(l-90) or possibly even InfB(l-21), could contain the expressivity tag region. Reliably quantifying the amount of protein is not feasible from either western blot or SDS-PAGE. Proteins of high molecular weight stain more intensively than low molecular weight proteins in SDS-PAGE. Western blots also have some quantitative discrepancies, as smaller proteins transfers more easily 5 from the SDS-PAGE to the nitrocellulose filter in the blotting procedure. Definitive conclusions were therefore troublesome, as proteins of differing molecular weights were compared. A fundamental problem realized with the InfB(l-x)-InfA constructs, was that the Sad restriction site (figure 2), which could interfere with results as the InfA reference did not have this site. Also when analyzing fragments 10 as small as InfB(l-9), the Sad site could also have just as much influence on expression as the actual InfB part. As results indicated, that the area of interest could be a region on 7nΛB(l-90), it was decided to leave out the Sad site for fragments of 7nΛB(l-90) and smaller for InfB(l-x)-gfp constructs (figure 4).
15 Example 2
Analysis of expression of InfB(l-x)-gfp
Total cellular extracts of cells expressing pBAD[/nfβ(l-x)-g/p] and pET24d[7nfβ(l- x)-gfp] were analyzed by SDS-PAGE (figure 6).
20 The odd migration pattern observed in figure 4 was still evident in both figure 6a and 6b. The reason had to stem from the common InfB moiety, even though this made little sense. Protein samples were boiled in sample buffer containing SDS and β-mercaptoethanol. This left out the possibility of only partial denaturation or persisting di-sulfidebridges, which could interfere with migration. Increasing the
25 concentration of SDS and β-mercaptoethanol in the sample buffer, did not change the migration pattern. Figure 6a showed gradually increasing amounts of protein expressed from gfp to InfB(l-21)-gfp. InfB(l-9)-gfp and InfB(l-12)-gfp were comparable to gfp, while InfB{l-15)-gfp, InfB{l-18)-gfp and InfB{l-21) produced more protein. InfB(l-90)-gfp and larger constructs seemed comparable to InfB(l-
30 21)-gfp, except InfB{l-471)-gfp which produced a slightly bigger band. This could however be attributed to more intensive staining, because of a higher molecular weight. The results in figure 6b showed a less pronounced enhancement of expression. InfB{l-9)-gfp expressed less than gfp. InfB{l-12)-gfp and 7nΛB(l-15)- gfp expressed approximately the same amounts as gfp. InfB(l-18)-gfp and larger fragments expressed more than gfp, even though this was not as pronounced as previously seen. The observation from the InfA experiments, that a fragment as small as 7nΛB(l-90) or possibly InfB{l-21) contained the expressivity was confirmed by the results from gfp.
Example 3
The amount of protein expressed from InfB{l-x)-gfp, was measured by the fluorescence from a volume of induced cells having a fixed cell density (figure 7).
The results seen in figure 7a deviated from the corresponding results in the SDS- PAGE (figure 6a).
InfB(l-90)-gfp, InfB(l-120)-gfp and InfB( 1-471) -gfp had 2-3 fold lower fluorescence counts than InfB(l-177)-gfp and InfB(l-282)-gfp, which did not match with the intensity of their corresponding protein bands in the SDS-PAGE. Results in figure 7b also showed a deviation from the corresponding results in the SDS-PAGE (figure 6b). gfp had the highest fluorescence count, while 7nΛB(l-90)- gfp, InfB{l-120)-gfp and InfB{l-471)-gfp had very low counts as in figure 6a. Cells were grown at 37° C, a temperature unfavorable for high level soluble protein expression. It was therefore suspected, that a part of the total amount of protein seen in the SDS-PAGEs (figure 6), was expressed as inactive inclusion bodies. The amount of insoluble protein produced at different growth temperatures (20° C, 30° C, 37° C) was analyzed (figure 8 and 9).
From figure 8a it can be seen, that a fraction of the protein expressed from InfB{l-471)-gfp, InfB{l-120)-gfp, InfB{l-90)-gfp at 37° C, was insoluble protein. This explains the deviation between the fluorescence measurements (figure 7a) and SDS-PAGE (figure 6a). InfB(l-177)-gfp , InfB(l-21)-gfp and InfB(l-18)-gfp also expressed some degree of insoluble protein. At a growth temperature of 30° C (figure 8b) none of the constructs produced an observable insoluble fraction, and this condition was concluded to be optimal for measuring the fluorescence. In figure 9 the insoluble fractions of protein expressed from BL21(DE3) pLysS pET24d[/nΛ3(l-x)-g/p] is seen.
In figure 9a it is seen how almost all constructs expressed at 37° C produced a considerable fraction of insoluble protein. This explains the deviation between the fluorescence measurements (figure 7b) and the SDS-PAGE (figure 6b). At a temperature of 30° C (figure 9b), insoluble proteins were still evident for some of the constructs. Lanes 1-3 had more protein loaded, which can be explained by incomplete sonication of cells. It does however not change the conclusion, that this growth temperature could not be used. At 20° C (figure 9c) the amount of insoluble protein was minimized even though bands could still be seen for protein expressed from InfB{l-471)-gfp, InfB{l-282)-gfp and InfB{l-177)-gfp constructs. It was decided to measure the fluorescence, for the constructs expressed from the BL21(DE3) pLysS pET24d cells at 20° C, keeping in mind that InfB{l-471)-gfp, InfB{l-282)-gfp and InfB{l-177)-gfp would produce lower readings, because of low solubility. The difference in the amount of insoluble protein fractions produced for pBAD and pET24 was caused by a difference in transcriptional levels, and thereby overall protein expressed. The pET system produced more protein than pBAD, and therefore also had higher amounts of inclusion bodies.
BL21 pBAD[/nΛ3(l-x)-g/p] and BL21(DE3) pLysS pET24d[/nfβ(l-x)-g/p] cells were grown at 30° C and 20° C and analyzed by SDS-PAGE (figure 10).
In both figure 10a and 10b it is seen that the overall expression of the constructs were lower than at 37° C (figure 6). The expression profile however, was the same and the fluorescence of the constructs measured (figure 11).
From figure 11a it is seen that the InfB(l-21) promoted the highest expression. As one codon is removed in InfB(l-18)-gfp to InfB(l-15)-gfp, the expression drops until InfB(l-12)-gfp, where it reaches the level of gfp. The fluorescence measurements correspond well with the SDS-PAGE result (figure 10a). The results in figure lib showed that InfB{l-21) promoted the highest expression, correlating well the results in figure 11a. The expression drops a bit with InfB{l-18)-gfp while InfB{l-15)-gfp was at the level of gfp. The expression of InfB{l-15)-gfp deviated from the results in figure 11a, where it was expressed 2.25 fold compared to gfp. The expression of InfB(l-120)-gfp in figure lib was lower than its surrounding constructs. If the SDS-PAGE in figure 10b is inspected, the InfB(l-120)-gfp band is less intensive than its surrounding bands, corresponding with the fluorescence results. In conclusion, a significant difference in expression was shown between gfp and InfB{l-21)-gfp of 3.5 and 5 fold in the two expression systems. It is confirmed by two independent systems, that InfB{l-21) is the expressivity tag region in InfB{l-471).
Example 4 Protein analysis by N-terminal sequencing
In bacteria, protein synthesis initiates with a formyl-methionine. The formyl group is removed post-translationally by a peptide deformylase (PDF). The methionine residue is, depending on the identity of the second amino acid in the peptide chain, removed by a methionine aminopeptidase (MAP). (Spector S. et. al, 2003. Expression of N-formylated protein in Escherichia coli. Protein expr purif. 32:317- 322). The N-terminal part of protein from InfB(l-18)-gfp , InfB(l-21)-gfp and InfB(l-90)-gfp was analyzed by N-terminal sequencing. Peptides and proteins with blocked N-terminal amino groups (e.g. acetylated, formylated) were not sequenced. Total cellular extracts was run on a SDS-PAGE and protein bands excised. The extracted protein was analyzed and eight cycles of sequencing showed that the eight N-terminal amino acids of protein from InfB{l-18)-gfp, InfB(l-21)-gfp and InfB(l-90)-gfp was correct. The methionine group was removed to great extent but a fraction of protein still had the residue attached. InfB(l-18)-gfp (10.6 %), InfB(l-21)-gfp (6.8 %) and InfB(l-90)-gfp (10.4 %), respectively. Differences were evident in the methionine removal; this could be of influence to the half-life of the proteins.
Example 5 Protein analysis by MALDI-TOF
The odd SDS-PAGE migration pattern observed for InfB{l-x)-InfA (figure 4) and InfB{l-x)-gfp (figure 6), protein from five constructs were analyzed by MALDI- TOF. InfB(l-9)-gfp, InfB(l-12)-gfp, InfB(l-15)-gfp, InfB(l-18)-gfp and InfB(l- 21)-gfp constructs were expressed by BL21(DE3) pLysS pET24d cells, and protein purified by immobilized metal affinity chromatography (IMAC) with the aid of their C-terminal his-tags (figure 12). All five proteins showed the same elution profile thus only one chromatogram is shown. Fractions of purified protein were very pure, with only a minor contaminant just below the 26 kDa marker band, seen on the SDS-PAGE especially in lane 4 and 5 (figure 12). It should also be noted, that the odd migration pattern persisted after the purification process.
The mass of each protein were analyzed by MALDI-TOF (figure 13).
All five MALDI-TOF spectra gave a scheme of four major peaks. Three of these peaks corresponded to the analyzed protein with 1-3 charges (z), respectively. A fourth peak was seen on all spectra, which did not immediately have a mass (m) that corresponded to any of the analyzed proteins. It was noted that the m/z value for this peak, varied by approximately 100-120 Da between spectra, corresponding to the approximate mass of an amino acid. By closer inspections it was found that the peak matched an N-terminal fragment, of each construct broken at the fluorophore site of GFP. The fluorophore site is made up of an internal Ser-Tyr-Gly sequence, and the fragment matched a breakage between the Tyr-Gly residues. The C-terminal breakage fragment did not produce any clear peak on the spectra. It had a size of approximately 20 kDa, and is marked on all spectra even though only a minor peak was seen. The fragmentation relates to the nature of the MALDI-TOF method. The crystallized protein was excited by laser photons which fragmented the GFP protein, as a result of the high intensity. Figure 14 summarizes the values obtained from all spectra. The theoretical mass is calculated from the assumption that the N-formyl-methionine was cleaved off. It can be seen, that the measured mass for each construct was close to its theoretical value. Any deviations between the measured data and theoretical values could relate to the fact, that calibration was done externally with only two proteins, Insulin (5734,6 Da) and Thioredoxin (11674,5 Da) which both had lower masses than the analyzed proteins. Deviations between measured and theoretical values are however acceptable. The odd migration pattern seen in the SDS-PAGE could not be attributed to a false nature of any of the analyzed proteins.
Each peak on the spectra had "shoulder" peaks. The shoulders were two different versions of the analyzed protein. One version contained the N-terminal methionine and the other contained the N-formyl-methionine. The degree of deformylation was quite similar for all constructs analyzed. There was a difference in the efficiency with which the methionine was cleaved off between constructs. For protein of InfB{l-21)-gfp it was efficiently removed while for InfB{l-9)-gfp it was not. The last three constructs had intermediate removal of the methionine. Figure.15 illustrates the difference between protein from InfB(l-21)-gfp and InfB(l-9)-gfp constructs.
The fact that some of the constructs had their N-terminal methionine more efficiently cleaved off could contribute to the odd migration pattern seen in the SDS-PAGEs, however in no way explain it. The migrational difference of the bands were of multiple kDa while the mass of a methionine is 149,2 Da. The MALDI-TOF results corresponded well with the N-terminal sequencing. Protein expressed from InfB(l-18)-gfp had a higher fraction of N-terminal methionine residues than InfB(l-21)-gfp. It was noted that this could be of influence for the half-life of the proteins.
Example 6
Structure of mRNA folded by MFOLD
In order to try and explain the optimal expressivity tag properties of InfB(l-21), folding schemes of mRNA sequences were analyzed. Any difference in structure in the TIR was noted. A loose structure in the TIR could be connected to a higher expression than a highly base paired TIR. In order to be able to compare the free energy between folded sequences, they had to be of the same length. It was therefore decided to minimize transcripts to nucleotides from the 5' end of the mRNA, and 50 nucleotides into the coding sequence of each construct. Due to these settings the constructs InfB(l-90)-InfA, InfB(l-120)-InfA, InfB(l-177)-InfA, InfB{l-282)-InfA, In fB{l-471)-In fA were not distinguished. The 5'-end of the pET15b vector transcript (87 nt) and 50 nts into each InfB(l- Λ71)-InfA construct were folded by the MFOLD program (figure 16). An array of folding schemes was obtained for each construct, and the ones with the lowest energy analyzed. A problem with comparing the schematic results was that a common scaffold was not applied, and similar sequences therefore ended up having quite different overall structure. This disturbed the interpretation of base pairs formed by the SD-sequence and initiation start codon. Due to parts of each sequence being removed, the results were only indicative. It was evident from figure 16, that no major difference in free energy was seen between constructs. If the structure of the mRNA should be attributed to the property of InfB{l-21), it was expected that InfA had the most base paired structure and 7nΛB(l-21) and InfB(50) the least base paired. This was not the case, as InfB(l- 50) and InfB(l-15)-InfA actually had a more base paired structure, looking at the free energy, than InfA. InfB(l-9)-InfA, InfB(l-12)-InfA and InfB(l-18)-InfA furthermore had comparable free energies of only a marginal higher value than InfA. InfB{l-21)-InfA had the highest free energy. The foldings did not correspond well with the amount of each protein produced (figure 4 and 5). Interpreting how base paired the SD-sequence and initiation start codon was in each construct, was not found feasible.
The InfB(l-x)-gfp constructs were also analyzed by the MFOLD program (figure 17).
From these folding schemes a tendency was evident, gfp had the lowest free energy while InfB{l-21)-gfp had the highest. This corresponded well with the relative amount of protein observed by the fluorescent measurements (figure 11). The rest of the constructs had intermediate free energies. If the folding schemes was to match the fluorescence results, InfB{l-9)-gfp and InfB{l-12)-gfp should have comparable free energies to gfp, which was not seen. Nevertheless it seemed that some degree of correspondence was seen, between structure in the TIR and the amount of protein expressed. It could in no way be ruled out, that a low level of structure in the TIR caused the property of the expressivity tag.
Example 7
Analysis of the expression of point mutants of InfB(l-21)-gfp
An intensive literature survey was conducted, in order see if any sequences were reported having similar properties of InfB(l-21). Only one sequence of a comparable length was found, being the first 18 nucleotides of the E. coli Trx gene. (Peti W. and Page R. 2007). expression in Escherichia coli with minimal cost. Protein Expr Purif. 51 : 1-10). Trx encodes the Thioredoxin Reductase protein and its sequence is seen in figure 18. To compare the effect of expression between the /nfβ(l-21)-tag and the Thiredoxin-tag disclosed in figure 18, the expressional levels of GFP positioned in frame behind each of the tags were measured (figure 27). With a relative expressional level of on 27% of Thioredoxin-GFP compared to /nfβ(l-21)-GFP, the tags of the invention can be considered better expressional tags than thioredoxin. The expressional levels were measured using fluorescence. Three codons out of six and twelve out of eighteen bases were the same. In order to check these possibly conserved positions were of particular importance to the expressivity of In fB{l-21), point mutations were made at position G7, A13 and A16. The mutants were expressed in BL21 pBAD and the fluorescence measured and compared to the wildtype InfB{l-21)-gfp construct. Fluorescence count is plotted in figure 19 together with the frequency of the specific codon which the point mutation caused, to see if codon frequency could be correlated with expression. It could be seen that all mutations to some extent compromised the optimal expressivity of the wildtype InfB(l-21).
The fact that point mutations at position 7 (figure 19a) had an influence on expression was interesting. The fluorescence counts in figure 11 showed, that InfB{l-9)-gfp expressed a little less than gfp. I might thus seem like a possible discrepancy between results, that position 7 in figure 19 influenced the expression. This was however not necessarily the case because of the complexity of the system. As deletions were from the 3'-end, the importance of the 5'-end had not been evaluated. The fact that InfB(l-9) did not have an expressivity tag property on its own, did not mean that the first 9 nucleotides of InfB(l-21) had no influence on its property. The results in figure 19 thus indicated, that it was the entire InfB{l-21) sequence, and not only the actual positions of bases 15-21 (showing a high expression in figure 11) that was the expressivity tag region. Quite a drastic drop in expression was seen in figure 19c for the InfB{l-21) 16G mutation, changing the codon from ACG (Thr) to GCG (Ala). The expressivity of InfB(l-21) could, according to this result be abolished by a single base pair substitution. It was noted, that many mutations caused a drop in expression to 0.4-0.6, compared to the wildtype. This drop was not below the expression of gfp and so some effect of the expressivity tag was still intact. No correlation between expression and codon frequency was found. For instance in figure 19b, the G mutation made a codon frequency of above 2, but still produced lower fluorescence counts than a mutation to T, which only had a codon frequency of 1.
Contrary in figure 19c, the wildtype A both had a higher fluorescence count and codon frequency than the mutant T.
The three positions analyzed all had an importance for the property of InfB{l-21). 5 Depending on which wildtype base was manipulated and which base was inserted, the expression could decrease to levels below gfp, while most of them caused a intermediate drop to 0.4-0.6.
A comprehensive point mutation analysis of every position in InfB(l-21) from position +7 to +21 was initiated, in order to understand the importance of each 10 individual base position. In addition a study of positions +4 to +6 was initiated.
The result of this study is shown in figure 28. Figure 28 shows that a number of mutations result in expression levels comparable to InfB{l-21) (seq ID NO: 2).
These mutations are position 5C→5A, 7G→7A, 7G→7T, 7G→7C, 8A→8C, 8A→8G,
8A→8T, 9T→9C, 12A→12T, 13A→13T, 14C→14A, 15G→15A, 17T→17C, 17T→17G, 15 17T→17A, 20A→20C, 20A→20G, 20A→20T, 21A→21G, and 21A→21T.
It is to be understood that the above list of mutations are also applicable to other versions of the expressivity tags of the invention (such as SEQ ID: NO 1, 3, 4 and sub-sequences of these sequences), wherein the specific positions are present. It 20 is also to be understood that the different mutations may be combined to create expressivity tags with multiple mutations compared to InfB(l-21) (and the other versions such as e.g. SEQ ID: NO 1, 3, 4 and sub-sequences of these sequences).
Similar, it is to be understood that the peptides corresponding to the above list of 25 mutation also form part of the present invention.
Thus, in an embodiment of the invention, the first nucleic acid sequence of the invention is mutated in at least one the following positions, as indicated, selected from the group consisting of positions 5C→5A (SEQ ID: NO 68), 7G→7A (SEQ ID: NO
30 69), 7G→7C (SEQ ID: NO 70), 7G→7T (SEQ ID: NO 71), 8A→8C (SEQ ID: NO 72),
8A→8T (SEQ ID: NO 73), 8A→8G (SEQ ID: NO 74), 9T→9C (SEQ ID: NO 75), 12A→12T (SEQ ID: NO 76), 13A→13T (SEQ ID: NO 77), 14C→14A (SEQ ID: NO 78), 15G→15A (SEQ ID: NO 79), 17T→17C (SEQ ID: NO 80), 17T→17A (SEQ ID: NO 81),17T→17G (SEQ ID: NO 82), 20A→20C (SEQ ID: NO 83), 20A→20G (SEQ ID: NO 84), 20A→20T
35 (SEQ ID: NO 85), 21A→21G (SEQ ID: NO 86), and 21A→21T (SEQ ID: NO 87). It is to be understood that the specific mutations present in SEQ ID: NO 68-87 are also applicable to SEQ ID: NO 1, 3, 4 and sub-sequences of SEQ ID: NO 1, 3, 4. Of course this requires that the specific position is present in said sequences.
Sequence analysis
The numbers of A/U residues and CA repeat bases have been reported to be of importance for the efficiency of translation, and are shown for sequences of In fB, InfA and gfp in figure 20. InfB{l-21) contained one more A/U residue than InfA{l-21) and g/p(l-21). This seems like a small difference. The point mutants of InfB{l-21) 7A and 7U increased the amount of A/U bases but decreased expression (figure 19a). There were no repeats of CA bases, but one and two in the entire sequence of InfB(l- 21) and InfA(l-21), respectively. No correlation was therefore found between the efficiency of expression of InfB(l-21) and A/U or CA bases.
Example 8
Downstream interactions between mRNA and 70S ribosome
In an attempt to explain the properties of InfB(l-21), interactions between the mRNA and ribosome in initiation of translation was analyzed. Interactions were visualized by crystal structures and the PyMoI program. In figure 21, mRNA is seen with the fMet-tRNA bound to the AUG initiation codon. A sphere of rRNA base residues, spaced within a 12 Angstrom radius around position +4 to +12 position on the mRNA is shown, to visualize the downstream tunnel.
Approximately 15 nucleotides of the mRNA are placed in the downstream tunnel. Position +11 to +15 of the mRNA is surrounded by a layer of protein consisting of S3, S4 and S5. A dense array of basic Arg amino acids is projected into the tunnel from these proteins. This is thought to interact with backbone phosphates of the mRNA. From position +7 to +10 the mRNA passes through a layer of rRNA. Cytosine 1397 and Uracil 1196 from this layer is oriented towards the mRNA around position +7 and +9 on the mRNA. It is though that these bases might help to position the mRNA immediately downstream of the A codon position. Base C1397 is shown in figure 22 in interaction with position +7 of the mRNA. The base U1196 is spaced from the mRNA in an intermediate position between +8 and +9. It is placed on the internal side of the downstream tunnel but not oriented directly towards the +9 mRNA base
When the coordination of C1397 of is linked to the sequence of InfB{l-21) an interesting notion can be made. At position +7 InfB contains a G. The possible interactions of InfB{l-21), InfA{l-21) and g/p(l-21) to C1393 is shown for each codon frame in figure 23.
The +1 codon frame is where the AUG codon is in the P-site. It can be seen that in this frame InfB{l-21) contains a G at the first codon position, which neither InfA or gfp does. In the next codon frames, InfA and gfp contains more G's at the first codon position than InfB. It could be speculated, that a G at position +7 in the mRNA is of particular importance for the coordination of the transcript in the tunnel. As the first and second peptide bond is formed, it is no longer of such importance. This is however only speculations, trying to fit the expressivity of InfB(l-21) to the structural arrangement. The analysis does also not explain why the entire sequence of In fB{l-21), is of importance to its expressivity.
Discussion The main aim of this project was to determine a expressivity tag region on the 471 nucleotide transcript of In fB{l-471) and an expressivity tag motif should be defined which was general for the Escherichia coli organism.
Odd migration pattern At first, the puzzle of the odd migration pattern observed in the SDS-PAGEs in figures 4 and 6 needs to be addressed. A lot of the results confirmed the integrity of the constructs produced. All constructs were sequenced correctly; an anti-IFl monoclonal antibody clearly recognized IFl protein expressed from InfB{l-x)-InfA constructs (figure 5). Fluorescence was measured from protein expressed from InfB(l-x)-gfp constructs (figure 11), and MALDI-TOF measurements of a selection of proteins involved in the odd migration pattern verified their mass within a very small margin of error (figure 13). The idea that an error in the primary sequence was the cause of the migration pattern was ruled out. The possibility that some of the constructs were modified post-translationally was considered. The number of post-translational modifications seen in E. coli is however limited. Any such modification of a covalent nature should also have been observed in the MALDI- TOF spectra, which it was not. A possible cleavage by a protease would also have been seen in the MALDI-TOF spectra. The possibility that another protein, could bind to some of the constructs, only after heat denaturation which was done before SDS-PAGE analysis, was considered. The presence of the anionic surfactant SDS should however disrupt any non-covalent interactions, and thereby exclude this possibility. Increasing the concentration of SDS and β-mercaptoethanol in the sample buffer did not change the migration pattern. No logical explanation could be reasoned for the odd migration pattern, and it can therefore for the moment only be characterized as an SDS-PAGE artifact.
Expressivity tag region on In fβ( 1-471)
At lot of work was put into analyzing and proving the authenticity of the expression promoted by the InfB(l-471) 3' deletions. Three different vector systems were made, and two different reporter genes analyzed in order to thoroughly support all results produced.
The analysis of the InfB{l-x)-InfA constructs revealed, that every 3' deletion produced could increase the expression of InfA. Differences in expression were seen as InfB{l-471) was shortened. According to SDS-PAGE (figure 4) and especially western blot (figure 5), the shortest fragment containing the region of interest was 7nΛB(l-90). The InfB(l-21) fragment induced an expression almost as high as 7nΛB(l-90), but it could not be definitively concluded if the region could be narrowed down to this small fragment, because of the limited sensitivity of SDS-PAGE and western blotting. It was interesting to see that InfA, a homologous gene, could have its expressivity enhanced by a part of another E.coli gene InfB.
Since GFP could be analyzed directly from cellular suspensions by fluorescence emission, protein amounts could be quantified in a more sensitive way. The results from the pBAD and pET24d expression system both showed, that InfB(l- 21) indeed was the minimum fragment of the ones produced, inducing optimal expressivity. Few differences between the expression from pBAD and pET24d were seen. A drop in expression was seen for constructs smaller than InfB(l-21)- gfp in both expression systems (figure 11). For the constructs expressed from pBAD, it was a linear drop per codon removed. In the pET24d system, a steep drop was seen for InfB{l-15)-gfp which only expressed at the level of gfp. This was a discrepancy between results not readily explained. It can therefore not from the results in figure 11 be concluded, if codon no. 5 is of higher importance as indicated by the pET24d system, or if every codon is equally important as indicated by the pBAD results. A comparison with the InfB(l-x)-InfA constructs is dubious, as this system contained the interfering Sad restriction site. Some of the construct with a InfB fragment longer than InfB(l-21) expressed lower amounts of protein than InfB(l-21)-gfp, which might seem puzzling as they contain the expressivity tag region of InfB{l-21). When the deviation on the dataset is taken into account, InfB{l-471)-gfp and InfB{l-120)-gfp for pET24d and InfB{l-90)-gfp, InfB{l-120)-gfp, InfB{l-177)-gfp and InfB{l-471)-gfp for pBAD expressed less than InfB{l-21)-gfp. The lower InfB{ 1-471 )-gfp expression from pET24d can be explained by the fraction of insoluble protein produced at 20° C (figure 9c ). Also, when the SDS-PAGE at figure 10 is inspected, it is seen that InfB(l-120)-gfp expresses less than its surrounding constructs indicating that the fluorescence measurements are reliable. The lower expressivity seen for the constructs from pBAD, could be a result of the crudeness of the fluorescence measurements, however also easily attributed to the complexity of the system as a whole. As reviewed in the introduction, small changes in the coding sequence of a gene can have a huge impact on its expression, e.g. because of a change in mRNA stability, structure etc. Observations might be attributed to this complexity, even though this is only speculations. It should be noted, that samples have been measured more than once and no difference in the expression has been seen, outside the standard deviation reported on the dataset in figure 11. Testing the expressivity of a InfB fragment of an intermediate length between InfB(l-21) and 7nΛB(l-90) might be a very relevant case, as the expressivity for the 3' deletions were not constant until the point of the smallest fragment with maximum expressivity; InfB{l-21). A fragment inducing an higher expressivity than InfB{l-21) could possibly be found when including an extra region. In conclusion, from the results presented it was shown that InfB(l-21) was the expressivity tag region on InfB(l- 471). Furtheremore, this region was shown to be able to enhance the expressivity of a heterologous gene, gfp, along with a homologous gene, InfA. Explanation for Infβ(l-21) properties
Due to the nature of the vector systems, cellular regulation of expression at the DNA level e.g. involving promoters of different strength was eliminated. Reasons for the efficient expressivity of InfB{l-21) was to be found in some of the factors presented in the introduction. The fact that the region of interest had been narrowed down to quite a small area on the transcript, eliminated some possibilities.
Neither InfA nor gfp contained a NGG codon at any of the position +2 to +5, which could cause abnormally low amounts of protein (Gonzales de Valdivia E. et. al, 2004. A codon window in mRNA downstream of the initiation codon where NGG codons give strongly reduced gene expression in Escherichia coli. Nucleic acids res. 32: 17:5198-5205). Also, none of the point mutations caused a NGG codon, so falsely low protein level can be excluded for any of the constructs. The 2nd codon position of a gene has been shown to have a great impact on protein synthesis. From the results in figure 11, it could be excluded that the 2nd codon position of InfB{l-21) was of particular importance for its expressivity as InfB{l-9) and InfB{l-12) did not increase expression of gfp. It could however be speculated, that the second codon position of InfB caused an increased of InfA. This would be a result of InfA having a poorly expressed second codon. It is speculated because InfB{l-9) and InfB{l-12) had an enhancing effect on InfA as seen in figure 5. An evaluation of the efficiency of the second codon according to the literature is noted in figure 24. The effect of the second codon on the expression of the lacZ gene in E. coli is shown for two different studies (Looman A. C. et. al and Stenstrom CM. et al.). In both studies the expression is shown relative to the codon which mediated the highest expression in the studies, AAA (Lys), which is set to 1. As the expression is related to a fixed maximum in both studies, the relation between codons in one study should be comparable to the relation of second study.
There is a deviation between the two reports. This is explained by Looman et al. using an expression system with a canonical SD sequence while Steenstrom et al. uses a system without. The study by Looman et al. is most comparable to our system as it also contained a SD sequence and will therefore be considered. The second codon of InfB could not have a positive effect on expression relative to InfA or gfp as seen in figure 20. It should actually decrease their expression. This correlated well with the results for gfp where e.g. InfB{l-9)-gfp expressed less than gfp. The slight increase in expression of InfB{l-9)-InfA relative to InfA could not be attributed to the second codon.
In conclusion, the 2nd codon position of InfB(l-21) is not a determining factor for the expressivity tag.
The structure of mRNA for all constructs produced was investigated in two folding schemes figures 16 and 17. For InfB{l-x)-InfA no convincing correlation could be found between expression and the structure of mRNA. It should however be noted, that InfB{l-21)-InfA had the highest free energy of the constructs, which should promote a higher expression. For InfB(l-x)-gfp a noteworthy correlation was seen, and again InfB(l-21)-gfp had the highest free energy. The free energy for InfB(l-9)-gfp and InfB(l-12)-gfp did not match their expression. If only InfB(l-21)-InfA/gfp were compared to InfA/gfp, a nice correlation between expression and free energy would be seen, however as it did not match for the other constructs, the results are questionable. It is difficult to draw definitive conclusions from MFOLD folding, because of questionable accuracies in its predictions. An evaluation of the 3.1 version of MFOLD, found an average prediction accuracy of 41 % for the 16S and 23S rRNA compared to crystal structures. Smaller RNA structures like 5S rRNA and tRNA, both of which are more comparable in length to the sequences folded in this thesis, all had prediction accuracies greater than 60 %. Noting the inaccuracy of MFOLD, its seems that InfB(l-21) could possibly promote a loose structure in the TIR, which could be of importance for its efficient expressivity.
The correlation between codon frequency and expression was analyzed by point mutations in InfB(l-21) (figure 19). It was speculated that a favorable codon composition in the immediate downstream region provided by InfB(l-21), could mediate a more efficient translation by eg. minimizing peptidyl tRNA drop-off. No convincing relation could however be seen between the expression of InfB(l-21) point mutants and codon frequency. The fact that codon frequency is not the determining factor is not a unique, as the second codon of a gene is not dictated by codon frequency. Additional information was gained from the point mutants. It was shown in figure 11 that InfB{l-9) and InfB{l-12) did not mediate an increased expressivity. It could therefore be speculated, that it was only the region InfB{15-21) that was the region of interest. It was in this regard interesting to note, that the positions +7 and +13 both was of influence to InfB(l- 21). These results indicated, that it was the entire 21 nucleotide sequence and not only the actual positions of nucleotides InfB(15-21) which were important to its expressivity. In conclusion, codon frequency is not a determining factor for the efficient expressivity of InfB{l-21).
When the ribosome is anchored at the Shine-Dalgarno sequence, it physically covers approximately 15 nt downstream of the initiation start codon. It is quite interesting how InfB(l-21) is close to this region also in length. Intuitively, this leads one to think of a mechanism explaining this coincidence. It could be speculated that the sequence of InfB(l-21) has motifs favorably interacting with one or two of the ribosomal subunits. A perfect match in the downstream tunnel involving multiple favorable interactions without steric clashes could promote an efficient initial expression. This could be of special importance in the initiation of translation, as a peptide chain has not made its way into the ribosomal exit tunnel stabilizing the complex. It could be speculated that once the peptide chain reaches a critical length, the amount of peptidyl tRNA drop-off is minimized. The fact that NGG codons at certain positions are known to cause peptidyl tRNA drop off, could be an example of the contrary to InfB(l-21), NGG codons causing unfavorable interactions. Since they cause peptidyl-tRNA drop off if from codon position +2 to +5 it seems reasonable that the A-site is of interest. If NGG is placed e.g. position + 5, after four peptide bonds are formed it is in the A-site and could cause drop off. More work should be put into analysis of the structural movements in the decoding process between ribosome, mRNA and tRNA.
An interpretation of the interactions between mRNA and the downstream tunnel is complex. The crystal structure analyzed (figures 21 and 22) has captured the complex in one state, while it in vivo is moving, constantly breaking and forming bonds. It is not known how much the mRNA can wiggle inside the tunnel which is of importance for the dynamics of its interactions. It was e.g. noted from two different crystal structures (pdb: 2HGR, 2HGP) that when the complex went from initiation to elongation, the residues spaced within a fixed Angstrom radius of the mRNA changed. Another problem was how it was not evident which interactions were acting favorably and unfavorably. It is e.g. likely, that a too strong interaction between mRNA and the ribosome could slow down translation. No clear motif in the downstream tunnel was found, which on some level was expected. The ribosome should be able to translate a broad combination of transcripts, and therefore primarily interacts with the mRNA backbone in the downstream tunnel. From crystal structures it can be seen that C1397 could coordinate G7 in InfB(l- 21) in the first codon frame position (figure 23). It could be of importance in positioning the mRNA as a supplement to the SD sequence. The coordination made by C1397 is not a new report. There is a statistical preference for guanosine at first codon position in highly expressed E.coli genes. It has been suggested, that a GNN pattern in these genes may be responsible for monitoring the correct reading frame during translation by complementary to C periodical sites in the E.coli 16 S rRNA. The point mutants for position +7 of InfB(l-21) (figure 8a) did not show that this residues was of extra importance. The expression of the two dropped to 0.5-0.6 compared to the wildtype which was a general pattern for the mutants.
The structural explanation still provides a theory why position +7 is of importance to InfB{l-21).
To conclude, the structural arrangement of InfB(l-21) mRNA on the ribosome was analyzed and favorable interactions between the two in the downstream tunnel could be of importance to the efficient expressivity of InfB(l-21). Analysis of a wider repertoire of structures would be preferred before conclusions are drawn for this subject.
Appendix A
PCR amplification PCR reactions were made in 100 μl_ aliquots by the following buffer concentrations and reaction conditions. Annealing temperatures for each PCR reactions along with primers can be found in figures 25 and 26.
0.75 μL Template (50-100 ng/uL) 0.6 μM Forward primer 0.6 μM Reverse primer 3 μl_ Pfu polymerase (cone. N. D.) 0.2 mM dNTP 5 2 mM MgCI2
75 mM Tris-HCI (pH 9.0) 20 mM (NH4)2SO4 0.01% Tween
10 30 cycles
94° C - I min
See figures 25-26 for annealing temperatures - 1 min.
720 c - 1-2 min
1 cycle
Figure imgf000064_0001
Plasmid purification
Plamids were purified by QIAGENs Qiaprep Spin Miniprep Kit by the supplied protocol. 3 ml_ of bacterial cell culture grown over night at 37° with appropriate 20 antibiotics was used per spin column and DNA was eluted with 60 μl_ H2O per spin column.
DNA fragment purification
PCR fragments and digested plasmids were run on a 1 % agarose gel using 25 marker lanes to avoid UV-exposure of the DNA. Bands were excised and the DNA extracted from the gel by the Wizard® SV Gel and PCR clean-up system according to the supplied protocol. Digested PCR fragments were directly purified by the Wizard® SV Gel and PCR clean-up system.
30 Digestion of DNA
Digestion reactions were performed by protocols from New England Biolabs using relevant restriction enzymes and buffers. A standard 75 μl_ scheme was used as shown below. Restriction reactions were incubated 90 minutes at 370C. If double digest reactions were not possible single digests were performed with an
35 intermediate DNA purification step as already described. 55 μl_ purified plasmid DNA/PCR fragment (100-200 ng/uL) 6.6 μL NEB buffer (1Ox) 2 μl_ Enzyme I (1011/μL) 2 μL Enzyme II (20U/μL)
1 μL BSA (lOmg/mL)
Ligation Digested inserts and plasmids were ligated by one of two schemes below depending on whether one or two inserts were needed. Ligation reactions were incubated at room temperature for 45 minutes.
10 μL Plasmid (200 ng/uL) 7 μL Insert (400 ng/uL)
2 μL T7 ligase buffer (1Ox)
1 μL T7 ligase (400.000 U/ μL)
10 μL Plasmid (200 ng/uL) 6 μL Insert I (400 ng/uL) 6 μL Insert II (400 ng/uL) 2.4 μL T7 ligase buffer (1Ox) 1 μL T7 ligase (400.000 U/ μL)
Transformation
JM109, BL21 or BL21(DE3) pLysS cells were first made competent for transformation. Cells were grown in 3 mL 2xTY media overnight at 37° C. 600 μL overnight culture was collected by centrifugation (15.000 x g, 1 min) and redissolved in 100 μL fresh 2xTY. The cell suspension was inoculated into 60 mL fresh 2xTY media and grown at 370C until OD600 reached 0,5. Cells were collected by centrifugation (3.700 x g, 10 min, 40C), resuspended in 10 mL TFl solution and incubated on ice for 30 min. Cells were again collected by centrifugation (3.700 x g, 10 min, 40C) and resuspended in 1 mL TF2 solution. Competent cells not used immediately were frozen in liquid nitrogen and stored at -80° C. Transformation reactions were carried out as follows. Ligation reaction of 20-24 μl_ or 5 μl_ purified plasmid was added to 100 μl_ aliqotes of competent cells and incubated on ice for 5 min. The reaction was given a 30 sec. heat shock at 42° C and incubated on ice for 1 min. 500 μl_ SOC media was added and the culture was 5 incubated at 370C for 1 hour. Cells were collected by centrifugation (15.000 x g, 1 min) redissolved in 150 μl_ 2xTY media and streaked on an agar plate with appropriate antibiotics (cone. 100 μg/mL ampicillin, 34 μg/mL chlorampenicol, 50 μg/mL kanamycin). Agar plates were incubated 12 hours at 37° C before clones were picked.
10
DNA sequence verification
Clones were picked from transformation plates and grown in 3 ml_ 2xTY at 37° C over night with appropriate antibiotics. Plasmid from each clone was purified by the QIAGENs Qiaprep Spin Miniprep Kit and eluted into 60 μl_ H2O per spin
15 column. A PCR screen was performed with T7 screen primers or pBAD screen primers according to the general PCR amplification scheme above with the primers listed in figure 26. Clones with right sized inserts were sent to sequencing at DNA- technology A/S. Sequencing primers used were T7 screen primers and pBAD screen primers.
20
Appendix B
Buffers and solutions
All buffers and solutions are in ddH2O unless otherwise stated. Fractional ratios (%) are indicated in weight per volume (w/v) for all buffers and 25 solutions unless otherwise stated.
Agarose gel electrophoresis Ethidium bromide 10 mg/mL ethidium bromide
Loading buffer (3x) 0.13 % bromophenol blue (v/v), 0.13 % xylene cyanol (v/v) , 30 30 % glycerol (v/v) TAE-buffer (5Ox) 2 M Tris-HCI (pH 8.0), 2 M acetic acid, 50 mM EDTA Molecular weight marker λ-DNA, BstEII digested (Bilabs)
Calcium competent cells TFl 10 mM NaAc pH 5.8, 50 mM MnCI2, 5 mM NaCI
TF2 10 mM NaAc pH 5.8, 5 mM MnCI2, 70 mM CaCI, 5 % glycerol
Cell growth and lysis 2xTY medium 1.6 % peptone from casein, 1 % yeast extract, 0.5 % NaCI
(autoclaved)
2xTY agar 2xTY medium with 1.5 % bacto agar (autoclaved)
Ampicillin 25 mg/mL, 0.2 μM filtered
Chloramphenicol 34 mg/mL in ethanol and 0.2 μM filtered Kanamycin 25 mg/mL, 0.2 μM filtered
L-arabinose 20 % L-arabinose, 0.2 μM filtered
IPTG 200 mg/mL IPTG (0.8 M), 0.2 μM filtered
0.9 % NaCI 0.9 % NaCI (autoclaved)
Lysis buffer 50 mM Tris-HCI pH 8, 150 mM NaCI, 1 mM EDTA, 0.1 mM PMSF
Redissolvment buffer 10 mM Tris-HCI pH 8, 0.1 mM PMSF
DNA amplification
PCR buffer (1Ox) 750 mM Tris-HCI (pH 9.0), 200 mM (NH4)2SO4, 0.1 % tween (v/v)
MgCI2 25 mM MgCI2 dNTP-mix 10 mM dATP, 10 mM dTTP, 10 mM dCTP, 10 mM dGTP
Pfu DNA polymerase N. D
All primers are listed in Figures 25 and 26
DNA digestion
BamHI 20 U/μL
BspHI 10 U/μL
BSA (1Ox) 10 mg/mL EcoRI 20 U/μL
Ncol 10 U/μL
NEB buffer 1 (1Ox) 100 mM Bis-Tris-Propane-HCI, 100 mM MgCI2, 10 mM dithiothreitol, pH 7
NEB buffer 3 (1Ox) 1 M NaCI, 500 mM Tris-HCI, 100 mM MgCI2, 10 mM dithiothreitol, pH 7.9 NEB buffer 4 (1Ox) 500 mM potassium acetate, 200 mM tris-acetate, 100 mM magnesium acetate, 10 mM dithiothreitol, pH 7.9 Sad 20 U/μL
Sphl 40 U/μL
DNA Ligation
T4 ligase buffer (1Ox) 0.5 M Tris-HCI (pH 7.5), 0.1 MgCI2, 0.01 M ATP, 0.1 M
DTT
T4 DNA ligase 400.000 U/mL
DNA purification
Qiaprep® Spin miniprep Kit (Qiagen)
Wizard® SV Gel and PCR Clean-Up system (Omega)
Protein purification
Buffer A 50 mM Tris-HCI (pH 7.6), 500 mM NaCI, 5 mM Imidazole
Buffer C 50 mM Tris-HCI (pH7.6), 500 mM NaCI, 500 mM Imidazole.
H(O) 50 mM HEPES (pH 7.6), 10 mM MgCI2, 1 mM DTT, 0.1 mM PMSF, 15 mM NaN3
Storage Buffer 10 % 10x H(O), 100 mM NaCI, 50 % glycerol
SDS-PAGE
Cracking buffer 50 mM Tris-HCI (pH 6.8), 15 % glycerol, 10 % SDS, 140 mM β-mercaptoethanol, 0.1 % bromophenol blue,
2 mM EDTA Sample buffer (5x) 125 mM Tris-HCI (pH 6.8), 50 % glycerol, 10 % SDS,
760 mM β-mercaptoethanol, 0.13 % bromophenol blue Separation buffer 0.75 M Tris-HCI (pH 8.8), 0.2 % SDS Stacking buffer 0.75 M Tris-HCI (pH 6.8), 0.2 % SDS
Electrophoresis buffer (1Ox) 0.25 M Tris-HCI (pH 8.3), 1.92 M glycine, 1% SDS Molecular weight marker PageRuler™ Prestained protein ladder (fermentas).
(170, 130, 95, 72, 55, 43, 34, 26, 17, 11 kDa)
Acrylamide-bis solution 30 % 37.5: 1 of acrylamide:N,N'-Methylene Bis- acrylamide APS 10 % ammoniumpersulfate
Staining solution Gelcode ® Blue stain reagent
Western Blotting Blocking buffer 10 % skim milk in TBS-T
Electrode blotting buffer 48 mM Tris-HCI (pH 8.9), 39 mM glycine, 0.0375 %
SDS, 20 % ethanol (v/v) Primary antibody 0.74 mg/mL monoclonal mouse anti-E.coli IFl antibody no. 1 produced in Biodesign lab Secondary antibody 1.3 mg/mL polyclonal rabbit anti-mouse immunoglobulin antibody HRP-conjugated TBS-T 0.9 % NaCI, 50 mM Tris-HCI (pH 7.4), 1 % tween (v/v)
The ECL Plus Western Blotting Detection System (GE healthcare) was used for westernblotting . The kit consists of two solutions ECLA and ECLB, compositions of these is unknown.
Example 9
Aim: To test whether the InfB(l-21) expressivity tag can increase the expression of other genes than gfp and InfA. Two poorly expressed genes streptavidin (SEQ ID NO : 88 (without tag) and SEQ ID NO : 89 (with tag)) and 132 (SEQ ID NO : 90 (without tag) and SEQ ID NO : 91 (with tag)) was chosen as reporter proteins.
Intro: Streptavidin (Genbank: X03591.1) is from the organism Streptomyces avidinii. Its 24 amino acid signal peptide was excluded (Argarana, C. E., et al. 1986, Miksch G, et al. 2008).
L32 is a single chain antibody fragment (scFv) of human origin (Sanz, L., et al. 2001 and Sanz, L., et al. 2002).
Streptavidin has been expressed in E. coli in several vector contexts, and is included in this example, since it is a very important biotechnological protein product. In most of the reported expression systems for streptavidin, the system has been optimally tuned for high expression. Many of these optimized systems for heterologous expression often results in premature translational abortion and protein folding problems. In order to avoid these problems it is often a necessity to employ sub-optimal expression conditions.
Expression of murin or humanized single chain antibodies are notoriously difficult to express in E. coli (Arbabi-Ghahroudi M et. Al. 2005). In this example, the L32 single chain antibody raised against human laminin through phagedisplay technology, has been expressed with and without the 21nt expressivity tag.
This example is carried out under such sub-optimal conditions, e.g. less strong promotor and low expression temperature. Under these conditions streptavidin and L32 is expressed at levels that is barely or not detectable on a SDS-PAGE.
Methods:
Cloning : InfB{l-21)-l32, InfB{l-21)-streptavidin, 132 and streptavidin were amplified from a template containing the appropriate gene by standard PCR. The infB{l-21) sequence was introduced through a non-annealing part of the forward primer (see Figure 30). Forward primer contained a BspHl restriction site, followed by the infB{l-21) sequence and the annealing region of the appropriate gene. The forward primer for genes without an infB{l-21) sequence contained a Λ/coI restriction site, followed by the annealing region of the appropriate gene {132 or streptavidin). Reverse primer contained an EcoRI restriction site followed by the annealing region of the appropriate gene {132 or strep). The PCR amplification included 30 cycles of 1 min at 94 0C (DNA denaturation), 1 min at 56 0C (primer hybridization) and 1 : 10 min at 72 0C (DNA elongation), followed by a single step for 5 min at 72 0C (final elongation). The PCR products were digested with BspHl or Λ/coI, respectively, and EcoRI and ligated into a Λ/coI and EcoRI treated pBAD plasmid, followed by transformation of ligation mixture into chemical competent E. coli JM109 cells.
Recombinant transformants on agar plate were identified using PCR screen. Identified plasmids with a correct insertion were verified by sequencing (DNA sequencing facility, MBI, Aarhus University, Denmark). Verified plasmids were transformed into chemical competent E. coli BL21 cells for expression. In figure 30 the restriction site for BspHI and EcoRI in the forward-primer and the reverse-primer, respectively, is underlined. The hybridization region of the primers is shown by italic and a bold for the infB(l-21) sequence and the annealing sequence with the appropriate gene, respectively.
Protein expression: pBAD/nΛ3(l-21)-/32, InfB(l-21)-strep and pBAD-132, -streptavidin) was transformed into chemical competent E. coli BL21 cells. For each construct, two 50 ml_ cultures were grown in 2x TY (containing 100 μg/ml ampicillin) at 37 0C. At an OD60O of 0.8, the gene expression was induced by adding L-arabinose to a final concentration of 0,04 %. For each construct, one culture was incubated at 30 0C and another at 37 0C for 3,5 h post induction. Analysis of expression was performed by a SDS-PAGE.
Results:
Streptavidin expression
The data presented in figure 29 clearly shows that an increase in the expression of streptavidin is observed when the tag is present. Arrows indicate the expected size of the proteins in question. This example is carried out under such sub- optimal conditions, e.g. less strong promotor and low expression temperature. Under these conditions streptavidin (no tag) is expressed at a level that is not detectable on a SDS-PAGE. Under the same conditions, but using the 5' 21nt expressivity tag a significant expression level is observed (more than 20 fold higher than with no tag, extimated using gel-scanning technology).
L32 expression
The data presented in figure 29 clearly shows that an increase in the expression of L32 takes place when the tag is present. Arrows indicate the expected size of the proteins in question. This example is carried out under such sub-optimal conditions, e.g. less strong promotor and low expression temperature. Under these conditions L32 (no tag) is expressed at a level that is barely detectable on a SDS-PAGE. Under the same conditions, but using the 5' 21nt expressivity tag a significant expression level is observed (more than 10 fold higher than with no tag, extimated using gel-scanning technology). Example 10
In a pilot study expression of subset of matrix metalloproteases were tested for expression using the 21 nt tag according to the invention. (MMP: Matrix metalloprotease; MT: Membran-type)
Tested constructs:
Human MTl-MMP catalytic domain
Human MTl-MMP catalytic domain (E250A)
Human MTl-MMP hemopexin domain Human MTl-MMP hemopexin domain + linker
Human MTl-MMP catalytic and hemopexin domain + linker
Human MMP-9 catalytic domain
Human MMP-2 catalytic domain
Human MMP-3 catalytic domain Human MMP-12 catalytic domain
Human MMP-8 catalytic domain
Human MMP-I catalytic domain
Human MMP-7 catalytic domain
None of these constructs gave any detectable expression in E. coli without the tag. By including the 21 nt expressivity tag they were expressed with up to 50 mg protein/L culture. Futher information on these protein can be found in (Maeda H, 1998)
References
Wolfgang Peti a and Rebecca Page. Strategies to maximize heterologous protein expression in Escherichia coli with minimal cost. Protein Expression and Purification 51 (2007) 1-10.
Sørensen et. al, 2003. A favorable solubility partner for the recombinant expression of streptavidin. Protein Expr Purif. 32:252-259 Looman A. C. et. al, 1987. Influence of the codon following the AUG initiation codon on the expression of a modified lacZ gene in Escherichia coli. The EMBO jour. 6:8:2489-2492.
Stenstrom CM. et al., 2000. Codon bias at the 3'-side of the initiation codon is correlated with translation initiation in Escherichia coli. Gene. 263 :273-284.
Argarana, C. E., et al., Molecular cloning and nucleotide sequence of the streptavidin gene. Nucleic acids research, 1986. 14(4) : p. 1871.
Sanz, L., et al., Generation and characterization of recombinant human antibodies specific for native laminin epitopes: potential application in cancer therapy. Cancer Immunology, Immunotherapy, 2001. 50(10) : p. 557-565.
Sanz, L., et al., Single-chain antibody-based gene therapy: inhibition of tumor growth by in situ production of phage- derived human antibody fragments blocking functionally active sites of cell-associated matrices. Gene therapy, 2002. 9(15) : p. 1049.
Miksch G, Ryu S, Risse JM, Flaschel E. Factors that influence the extracellular expression of streptavidin in Escherichia coli using a bacteriocin release protein. Appl Microbiol Biotechnol. 2008; 81(2) :pp 319-26.
Arbabi-Ghahroudi M, Tanha J, MacKenzie R. Prokaryotic expression of antibodies. Cancer Metastasis Rev. 2005;24(4) :pp 501-19.
Maeda H, Okamoto T, Akaike T. Human matrix metalloprotease activation by insults of bacterial infection involving proteases and free radicals. Biol Chem. 1998 Feb;379(2) : 193-200.

Claims

Claims
1. A recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide, said nucleic acid comprising a first nucleic acid sequence of maximum 99 nt encoding a polypeptide operably linked to the a second heterologous nucleic acid sequence encoding a polypeptide, wherein said first nucleic acid sequence comprises a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b).
2. The recombinant hybrid nucleic acid sequence according to claim 1, wherein said first nucleic acid sequence encodes an expressivity tag.
3. The recombinant hybrid nucleic acid sequence according to any of the claims 1- 2, wherein said first nucleic acid sequence of c) has a minimum length of 15 nt nucleotides.
4. The recombinant hybrid nucleic acid sequence according to any of the preceding claims 1-3, wherein said first nucleic acid sequence is positioned 5' to the second heterologous nucleic acid sequence.
5. The recombinant hybrid nucleic acid sequence according to any of the preceding claims 1-4, wherein said second heterologous nucleic acid sequence has been modified to reflect the codon preference of a host organism selected to express the nucleic acid molecule.
6. The recombinant hybrid nucleic acid sequence according to any of the preceding claims, further comprising a nucleic acid sequence encoding a polypeptide tag for detection, solubility or affinity purification.
7. The recombinant hybrid nucleic acid sequence according to any of the preceding claims, wherein said tag for affinity purification, solubility or detection is selected from the group consisting of
BCCP-tag, c-myc-tag, calmodulin-tag (SEQ ID NO:33), FLAG-tag (SEQ ID NO:45), HA-tag (SEQ ID NO:50), His-tag (SEQ ID NO:29), Maltose binding protein-tag (SEQ ID NO:27), Nus-tag (SEQ ID NO:60), Glutathione-S-transferase-tag (SEQ ID NO:31), Green fluorescent protein-tag (SEQ ID NO:62), Thioredoxin-tag (SEQ ID NO:64), S-tag (SEQ ID NO:41), Strep-tag (SEQ ID NO:35), human protein C tag (SEQ ID NO:37), Chitin binding protein tag (SEQ ID NO:39), T7-tag (SEQ ID NO:43), Myc-tag (SEQ ID NO:52), V5-tag, VSV-tag, Avi.tag (SEQ ID NO:56), BioEase-tag (SEQ ID NO: 58), SNAP-tag, FlaSH-tag, Nus A-tag, DsbA-tag (SEQ ID NO:66)and the recombinant nucleic acid sequence itself.
8. The recombinant hybrid nucleic acid sequence according to any of the preceding claims, further comprising a nucleic acid sequence encoding a spacer positioned between said first and said second nucleic acid sequence.
9. The recombinant hybrid nucleic acid sequence according to any of the preceding claims 1-8, further comprising a nucleic acid sequence encoding a spacer positioned between the first and second codon of said first nucleic acid sequence and operably linking said codons.
10. A recombinant hybrid polypeptide, said polypeptide comprising a first polypeptide amino acid sequence and a second heterologous amino acid sequence, wherein said first amino acid sequence of maximum 33 amino acids is selected from the group consisting of a) an amino acid sequence selected from the group consisting of SEQ ID NO:5, SEQ ID NO:6, SEQ ID NO:7, SEQ ID NO:8, SEQ ID NO:9, SEQ ID NO: 10, SEQ ID NO: 11 and SEQ ID NO: 12, and b) an amino acid sequence having at least 75% sequence identity to a), and c) an amino acid sequence which is a sub-sequence of a) and b).
11. The recombinant hybrid polypeptide according to claim 10, wherein said first amino acid sequence is an expressivity tag.
12. The recombinant hybrid polypeptide according to any of the claims 10-11, wherein said first amino acid sequence of c) has a minimum length of 4 amino acids.
13. The recombinant hybrid polypeptide according to any of the claims 10-12, wherein said first amino acid sequence is positioned N-terminal to the second heterologous amino acid sequence.
14. The recombinant hybrid polypeptide according to any of the claims 10-13, further comprising a polypeptide tag for detection, solubility or affinity purification.
15. The recombinant hybrid polypeptide according to any of the claims 10-14, wherein said tag for affinity purification, solubility or detection is selected from the group consisting of
BCCP-tag, c-myc-tag, calmodulin-tag (SEQ ID NO: 34), FLAG-tag (SEQ ID NO:46, SEQ ID NO:47, SEQ ID NO:48 and SEQ ID NO:49), HA-tag (SEQ ID NO:51), His- tag (SEQ ID NO:30), Maltose binding protein-tag (SEQ ID NO:28), Nus-tag (SEQ ID NO:61), Glutathione-S-transferase-tag (SEQ ID NO:32), Green fluorescent protein-tag (SEQ ID NO:63), Thioredoxin-tag (SEQ ID NO:65), S-tag (SEQ ID NO:42), Strep-tag (SEQ ID NO:36), human protein C tag (SEQ ID NO:38), Chitin binding protein tag (SEQ ID NO:40), T7-tag (SEQ ID NO:44), Myc-tag (SEQ ID NO:53), V5-tag (SEQ ID NO:54), VSV-tag (SEQ ID NO:55), Avi.tag (SEQ ID NO: 57), BioEase-tag (SEQ ID NO: 59), SNAP-tag, FlaSH-tag, Nus A-tag, DsbA-tag (SEQ ID NO:67) and the recombinant polypeptide sequence itself.
16. A vector comprising a first nucleic acid sequence of maximum 99 nt, said first sequence comprising a nucleic acid sequence selected from the group consisting of a) SEQ ID NO: 1, SEQ ID NO:2, SEQ ID NO:3 and SEQ ID NO:4, and b) a nucleic acid sequence having at least 75% sequence identity to a), and c) a nucleic acid sequence, which is a sub-sequence of a) and b).
17. The vector according to claim 16, wherein said vector is further adapted for insertion of a second heterologous nucleic acid sequence encoding a polypeptide operably linked to said first nucleic acid sequence.
18. The vector comprising a recombinant hybrid nucleic acid sequence according to any of the claims 1-9.
19. The vector according to any of claims 16-18, wherein said recombinant hybrid nucleic acid sequence is operably linked to an expression control sequence, as the araBAD promoter system such as the Ara-promoter (SEQ ID NO: 13), Ara-operator 1 (SEQ ID NO: 14), Ara-operator 2 (SEQ ID NO: 15) and the AraC repressor (SEQ ID NO: 16), the lac promoter system such as the Lac Promoter (SEQ ID NO: 17), the Lac-operator (SEQ ID NO: 18) and the Lad repressor (SEQ ID NO: 19), Tac/Trc-promoter system (SEQ ID NO: 20), Lambda PL - promoter (SEQ ID NO:21), Lambda PR promoter (SEQ ID NO:22), T7 promoter (SEQ ID NO:23), T3 promoter (SEQ ID NO:24), SP6 promoter (SEQ ID NO:25) and the AmpR promoter (SEQ ID NO:26).
20. The vector of any of claims 16-19, wherein the vector is selected from the group consisting of plasmids, cosmids, phages, bacterial artificial chromosomes
(BAC), Phagemids and Pl-derived artificial chromosomes.
21. The vector according to any of the claims 16-20, wherein the vector is an in vitro translation vector.
22. The vector according to any of claims 16-21 , wherein said plasmid is selected from the group consisting of TA cloning vectors, Gateway cloning vectors, restriction cloning vectors, Topo cloning vectors, pET vector system and pBAD vector systems.
23. A host cell comprising a vector according to any of claims 16-22.
24. The host cell according to claim 23, wherein said host cell is a bacterium or fungus.
25. The host cell according to claim 23 or 24, wherein said host cell is E.coli.
26. The host cell according to any of claims 23-25, wherein the host cell is E. coli JM109 and E. coli BL21(DE3).
5
27. An isolated antibody or antigen-binding fragment thereof which selectively binds to the first polypeptide amino acid sequence of any of the claims 10-15.
28. The antibody, isolated antibody or antigen-binding fragment there of
10 according to claim 27, wherein the antibody is a purification tag or an epitope tag.
29. Use of a recombinant hybrid nucleic acid sequence encoding a hybrid polypeptide according to the preceding claims for the expression of a said polypeptide.
15
30. The use according to claim 29, wherein said recombinant hybrid nucleic acid sequence is comprised in a vector according to any of the claims 16-22 adapted for expression of said recombinant polypeptide.
20 31. A method for expressing a recombinant polypeptide comprising a) introducing into an expression system the hybrid nucleic acid sequence according to any of the claims 1-9 encoding a recombinant hybrid polypeptide, b) expressing said nucleic acid sequence in the expression system, and 25 optionally c) purifying the expressed recombinant polypeptide .
32. A method according to claim 31, wherein the hybrid nucleic acid sequence is comprised in a vector according to any of claims 16-22 is used.
30
33. A method according to claims 31 or 32, wherein the expression system is a host according to any of claims 23-26.
34. A method according to claims 31 or 32, wherein said expression system is an 35 in vitro expression system.
35. A method according to claim 34, wherein the in vitro translation system comprises extract from E.coli.
36. A method according to any of the claims 31-35, wherein an antibody according to claims 27 or 28 is used for purification of said recombinant polypeptide.
37. A method according to any of the claims 31-36, wherein an antibody according to claims 18 or 19 is used for detecting said recombinant polypeptide.
38. A method according to any of the claims 31-37, wherein said first polypeptide amino acid sequence is subsequently removed from the expressed target nucleic acid sequence by proteolytic cleavage.
39. A kit comprising a vector according to any of the claims 16-22.
PCT/DK2010/050043 2009-02-20 2010-02-18 Expressivity tag and use thereof Ceased WO2010094288A1 (en)

Applications Claiming Priority (4)

Application Number Priority Date Filing Date Title
DKPA200900239 2009-02-20
DKPA200900239 2009-02-20
DKPA200970302 2009-12-28
DKPA200970302 2009-12-28

Publications (1)

Publication Number Publication Date
WO2010094288A1 true WO2010094288A1 (en) 2010-08-26

Family

ID=42049304

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/DK2010/050043 Ceased WO2010094288A1 (en) 2009-02-20 2010-02-18 Expressivity tag and use thereof

Country Status (1)

Country Link
WO (1) WO2010094288A1 (en)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109971768A (en) * 2019-04-18 2019-07-05 贵州大学 A kind of sorghum transcription factor SbSBP5 gene and its recombinant vector and expression method
WO2023161371A1 (en) * 2022-02-25 2023-08-31 Danmarks Tekniske Universitet Engineered bacteria secreting cytokines for use in cancer immunotherapy
EP4410982A4 (en) * 2021-09-27 2025-01-15 FUJIFILM Corporation METHOD FOR PRODUCING POLYPEPTIDE, MARKER, EXPRESSION VECTOR, METHOD FOR EVALUATING POLYPEPTIDE, METHOD FOR PRODUCING NUCLEIC ACID DISPLAY LIBRARY, AND SCREENING METHOD

Non-Patent Citations (17)

* Cited by examiner, † Cited by third party
Title
ARBABI-GHAHROUDI M; TANHA J; MACKENZIE R.: "Prokaryotic expression of antibodies", CANCER METASTASIS REV., vol. 24, no. 4, 2005, pages 501 - 19, XP019205203, DOI: doi:10.1007/s10555-005-6193-1
ARGARANA, C.E. ET AL.: "Molecular cloning and nucleotide sequence of the streptavidin gene", NUCLEIC ACIDS RESEARCH, vol. 14, no. 4, 1986, pages 1871, XP002951809
DATABASE UniProt [online] 10 February 2009 (2009-02-10), "SubName: Full=Translation initiation factor IF-2;", XP002577094, retrieved from EBI accession no. UNIPROT:B6ZXX7 Database accession no. B6ZXX7 *
GONZALES DE VALDIVIA E.: "A codon window in mRNA downstream of the initiation codon where NGG codons give strongly reduced gene expression in Escherichia coli", NUCLEIC ACIDS RES., vol. 32, no. 17, 2004, pages 5198 - 5205
JANA S ET AL: "Strategies for efficient production of heterologous proteins in Escherichia coli", APPLIED MICROBIOLOGY AND BIOTECHNOLOGY, SPRINGER, BERLIN, DE LNKD- DOI:10.1007/S00253-004-1814-0, vol. 67, no. 3, 1 May 2005 (2005-05-01), pages 289 - 298, XP019331816, ISSN: 1432-0614 *
LOOMAN A.C.: "Influence of the codon following the AUG initiation codon on the expression of a modified lacZ gene in Escherichia coli", THE EMBO JOUR., vol. 6, no. 8, 1987, pages 2489 - 2492
MAEDA H; OKAMOTO T; AKAIKE T.: "Human matrix metalloprotease activation by insults of bacterial infection involving proteases and free radicals", BIOL CHEM., vol. 379, no. 2, February 1998 (1998-02-01), pages 193 - 200
MIKSCH G; RYU S; RISSE JM; FLASCHEL E.: "Factors that influence the extracellular expression of streptavidin in Escherichia coli using a bacteriocin release protein", APPL MICROBIOL BIOTECHNOL, vol. 81, no. 2, 2008, pages 319 - 26, XP019654160, DOI: doi:10.1007/s00253-008-1673-1
PETI ET AL: "Strategies to maximize heterologous protein expression in Escherichia coli with minimal cost", PROTEIN EXPRESSION AND PURIFICATION, ACADEMIC PRESS, SAN DIEGO, CA LNKD- DOI:10.1016/J.PEP.2006.06.024, vol. 51, no. 1, 30 November 2006 (2006-11-30), pages 1 - 10, XP005727958, ISSN: 1046-5928 *
PETI W.; PAGE R.: "expression in Escherichia coli with minimal cost", PROTEIN EXPR PURIF., vol. 51, 2007, pages 1 - 10, XP005727958, DOI: doi:10.1016/j.pep.2006.06.024
SANZ, L. ET AL.: "Generation and characterization of recombinant human antibodies specific for native laminin epitopes: potential application in cancer therapy", CANCER IMMUNOLOGY, IMMUNOTHERAPY, vol. 50, no. 10, 2001, pages 557 - 565, XP002365676, DOI: doi:10.1007/s00262-001-0235-5
SANZ, L. ET AL.: "Single-chain antibody-based gene therapy: inhibition of tumor growth by in situ production of phage-derived human antibody fragments blocking functionally active sites of cell-associated matrices", GENE THERAPY, vol. 9, no. 15, 2002, pages 1049
SORENSEN H P ET AL: "A favorable solubility partner for the recombinant expression of streptavidin", PROTEIN EXPRESSION AND PURIFICATION, ACADEMIC PRESS, SAN DIEGO, CA LNKD- DOI:10.1016/J.PEP.2003.07.001, vol. 32, no. 2, 1 December 2003 (2003-12-01), pages 252 - 259, XP004475036, ISSN: 1046-5928 *
SØRENSEN: "A favorable solubility partner for the recombinant expression of streptavidin", PROTEIN EXPR PURIF., vol. 32, 2003, pages 252 - 259, XP004475036, DOI: doi:10.1016/j.pep.2003.07.001
SPECTOR S.: "Expression of N-formylated protein in Escherichia coli", PROTEIN EXPR PURIF., vol. 32, 2003, pages 317 - 322, XP004475044, DOI: doi:10.1016/j.pep.2003.08.004
STENSTROM C.M. ET AL.: "Codon bias at the 3'-side of the initiation codon is correlated with translation initiation in Escherichia coli", GENE, vol. 263, 2000, pages 273 - 284
WOLFGANG PETI A; REBECCA PAGE: "Strategies to maximize heterologous protein expression in Escherichia coli with minimal cost", PROTEIN EXPRESSION AND PURIFICATION, vol. 51, 2007, pages 1 - 10, XP005727958, DOI: doi:10.1016/j.pep.2006.06.024

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109971768A (en) * 2019-04-18 2019-07-05 贵州大学 A kind of sorghum transcription factor SbSBP5 gene and its recombinant vector and expression method
EP4410982A4 (en) * 2021-09-27 2025-01-15 FUJIFILM Corporation METHOD FOR PRODUCING POLYPEPTIDE, MARKER, EXPRESSION VECTOR, METHOD FOR EVALUATING POLYPEPTIDE, METHOD FOR PRODUCING NUCLEIC ACID DISPLAY LIBRARY, AND SCREENING METHOD
WO2023161371A1 (en) * 2022-02-25 2023-08-31 Danmarks Tekniske Universitet Engineered bacteria secreting cytokines for use in cancer immunotherapy

Similar Documents

Publication Publication Date Title
JP7619938B2 (en) Protein Purification Methods
EP2995683B1 (en) Method for producing peptide library, peptide library, and screening method
WO2011098772A1 (en) Peptide tag systems that spontaneously form an irreversible link to protein partners via isopeptide bonds
Öjemalm et al. Quantitative analysis of SecYEG-mediated insertion of transmembrane α-helices into the bacterial inner membrane
CN116670284A (en) Method for screening candidate molecules capable of forming complexes with multiple target molecules
Ojima-Kato et al. Nascent MSKIK peptide cancels ribosomal stalling by arrest peptides in Escherichia coli
Matsunaga et al. Addition of arginine hydrochloride and proline to the culture medium enhances recombinant protein expression in Brevibacillus choshinensis: The case of RBD of SARS-CoV-2 spike protein and its antibody
WO2010094288A1 (en) Expressivity tag and use thereof
JP4303112B2 (en) Methods for the generation and identification of soluble protein domains
Marino et al. Bicistronic mRNAs to enhance membrane protein overexpression
JPWO2020080490A1 (en) Method for producing peptide library
Rojano-Nisimura et al. Codon selection affects recruitment of ribosome-associating factors during translation
WO2005033286A2 (en) Compositions and methods for synthesizing, purifying and detecting biomolecules
US7435538B2 (en) High throughput screening method of drug for physiologically active protein
JP5809687B2 (en) RNF8-FHA domain modified protein and method for producing the same
JP2020502493A (en) Methods for screening and detecting proteins
Choi et al. Functional characterization of the C-terminus of YhaV in the Escherichia coli PrlF-YhaV toxin-antitoxin system
Burgos et al. Widespread ribosome stalling in a genome-reduced bacterium and the need for translational quality control
Okada et al. Contribution of the second OB fold of ribosomal protein S1 from Escherichia coli to the recognition of TmRNA
Endoh et al. Effective approaches for the production of heterologous proteins using the Thermococcus kodakaraensis-based translation system
Ito et al. Trans‐translation mediated by Bacillus subtilis tmRNA
JP2017160272A (en) Library of azoline compound and azole compound, and production method thereof
US8527211B2 (en) Method to produce recombinant MBP8298 and other polypeptides by nucleotide structure optimization
Vazulka et al. RNA-seq reveals multifaceted gene expression response to Fab production in Escherichia coli fed-batch processes with particular focus on ribosome stalling
Baijal et al. Identification of polyphosphate-binding proteins in E. coli uncovers targets involved in translation control and ribosome biogenesis

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 10705533

Country of ref document: EP

Kind code of ref document: A1

122 Ep: pct application non-entry in european phase

Ref document number: 10705533

Country of ref document: EP

Kind code of ref document: A1