EP4655405A1 - New genomic integration site for fungal protein and fine chemicals production - Google Patents

New genomic integration site for fungal protein and fine chemicals production

Info

Publication number
EP4655405A1
EP4655405A1 EP24702258.5A EP24702258A EP4655405A1 EP 4655405 A1 EP4655405 A1 EP 4655405A1 EP 24702258 A EP24702258 A EP 24702258A EP 4655405 A1 EP4655405 A1 EP 4655405A1
Authority
EP
European Patent Office
Prior art keywords
seq
sequence
nucleic acid
acid sequence
amino acid
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24702258.5A
Other languages
German (de)
French (fr)
Inventor
Katja Gemperlein
Stefan Haefner
Sebastian Christopher SPOHNER
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
BASF SE
Original Assignee
BASF SE
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by BASF SE filed Critical BASF SE
Publication of EP4655405A1 publication Critical patent/EP4655405A1/en
Pending legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/79Vectors or expression systems specially adapted for eukaryotic hosts
    • C12N15/80Vectors or expression systems specially adapted for eukaryotic hosts for fungi
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/87Introduction of foreign genetic material using processes not otherwise provided for, e.g. co-transformation
    • C12N15/90Stable introduction of foreign DNA into chromosome
    • C12N15/902Stable introduction of foreign DNA into chromosome using homologous recombination
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/24Hydrolases (3) acting on glycosyl compounds (3.2)
    • C12N9/2402Hydrolases (3) acting on glycosyl compounds (3.2) hydrolysing O- and S- glycosyl compounds (3.2.1)
    • C12N9/2405Glucanases
    • C12N9/2434Glucanases acting on beta-1,4-glucosidic bonds
    • C12N9/2437Cellulases (3.2.1.4; 3.2.1.74; 3.2.1.91; 3.2.1.150)
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12RINDEXING SCHEME ASSOCIATED WITH SUBCLASSES C12C - C12Q, RELATING TO MICROORGANISMS
    • C12R2001/00Microorganisms ; Processes using microorganisms
    • C12R2001/645Fungi ; Processes using fungi

Definitions

  • the present invention relates to genomic insertion sites allowing for strong expression in filamentous fungi.
  • Biocatalytical and fermentative production of recombinant proteins and fine chemicals has been applied in industrial scale for a long time.
  • demand for products derived from biological processes is increasing.
  • Various organisms have been used in such production including bacteria and fungi.
  • filamentous fungi although difficult to transform and modify genetically have been used more and more as they have been proven to be highly efficient in biocatalytical and fermentative production.
  • Filamentous fungi have been shown to be excellent hosts for the production of a variety of proteins.
  • Fungal strains such as Aspergillus, Trichoderma, Penicillium and Thermothelomyces have been applied in the industrial production of a wide range of enzymes, since they can secrete large amounts of protein into the fermentation broth.
  • the protein-secreting capacity of these fungi makes them preferred hosts for the targeted production of specific enzymes or enzyme mixtures.
  • T. thermophilus' Wild type Thermothelomyces thermophilus (T. thermophilus') C1 (also known as Myceliophthora thermophila, previously described as Chrysosporium lucknowense) is a thermotolerant ascomy- cetous filamentous fungus which has been described as being attractive for production of cellulases and other proteins on a commercial scale. Only recently, the strain has also been shown to be highly efficient for producing various fine chemicals, e.g. in WO20/161682 it is disclosed that T. thermophilus C1 is capable of producing cannabinoids and precursors thereof.
  • WO00/20555 and US2012/0005812 disclose transformation systems for T. thermophilus C1 and describe expressing and secreting heterologous proteins or polypeptides. Also disclosed is a process for producing large amounts of polypeptides or proteins in an economical manner.
  • WO15/004241 discloses multiple proteases deficient filamentous fungal cells and methods useful for the production of heterologous proteins.
  • a first embodiment of the invention is an expression system comprising a. a filamentous fungi host cell comprising a native genomic region comprising a sequence selected from the list consisting of i) a nucleic acid sequence comprising SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 44, and iii) a fragment of at least 500, preferably at least 750, preferably at least
  • a recombinant integrating nucleic acid molecule comprising a nucleic acid molecule flanked by distinct homology arms each at least 70% identical, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% most preferably identical over the whole length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 1000
  • the nucleic acid molecule flanked by said homology arms may for example comprise a sequence encoding a protein or an RNA molecule, landing pad or a recombinant expression construct.
  • Said landing pad may comprise a Cre/lox site, one or more sequences recognized by guide RNAs having reduced off-target effects guiding a CRISPR/Cas molecule to this site, one or more recognition sites for a homing endonuclease and the like (Bourgeois et al (2016) ACS Synth Biol 7(11)).
  • Said recombinant expression construct may comprise a promoter, a sequence encoding a protein or an RNA molecule and optionally a terminator.
  • a further embodiment of the invention is the expression system as defined above wherein said native genomic region is comprising an ORF encoding a small secreted protein flanked at one side by at least 250 nucleotides non-coding DNA and at the other side by at least 250 nucleotides non-coding DNA.
  • said ORF is flanked downstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, preferably at least 1100, more preferably at least 1200 nucleotides non-coding DNA.
  • said ORF is downstream flanked by at least 1296 nucleotides non-coding DNA.
  • said ORF is flanked upstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, preferably at least 1700, more preferably at least 1800 nucleotides non-coding DNA.
  • said ORF is upstream flanked by at least 1847 nucleotides non-coding DNA.
  • An additional embodiment of the invention is any of the expression systems as defined before wherein said small secreted protein is comprising a sequence selected from the list of i. an amino acid sequence comprising SEQ ID NO: 46 ii. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 46, iii. a fragment of at least 100, preferably at least 110, more preferably at least 120, even more preferably at least 130, most preferably at least 140 consecutive amino acids of i) or ii).
  • ORF is encoding the protein MYCTH_2314963 (GeneBank Gene ID 11512560 updated November 28, 2019.
  • Another embodiment of the invention is any of the expression systems as defined above wherein the native genomic region is flanked at one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 41 b. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 41, and c. a fragment of at least 100, preferably at least 150, preferably at least
  • flanking regions are not necessarily directly adjacent to a sequence used to define said genomic region.
  • the genomic region may be defined by a coding sequence, located within said genomic region. However, there may be spacers between such coding region and the respective flanking regions defining the respective genomic region.
  • a further embodiment of the invention is any of the expression systems as defined above wherein the filamentous fungi host cell is from a species of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysospori- um, Filibasidium, Fusarium, Humicola, Magnaporthe, Monascus, Mucor, Myceliophthora, Mor- tierella, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Piromyces, Phanerochaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordaria, Talaromyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypocladium, Trametes and Trichoderma, more preferred Aspergillus n
  • a further embodiment of the invention is a method for stable integration of a recombinant polynucleotide in a filamentous fungi host cell comprising the steps of c. providing a filamentous fungi host cell comprising a native genomic region comprising a sequence selected from the list consisting of i) a nucleic acid sequence comprising SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99%to SEQ ID NO 44, and iii) a fragment of at least 500, preferably at least 750, preferably at least
  • nucleic acid molecule comprising a nucleic acid molecule flanked by distinct homology arms each at least 70% identical, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% most preferably identical over the whole length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more
  • the recombinant nucleic acid molecule comprises a coding sequence, for example comprised in a recombinant expression construct it is preferred embodiment that the host cell having integrated said recombinant nucleic acid molecule produces the recombinant peptide, protein and/or nucleic acid of interest encoded by said coding region.
  • said native genomic region before recombination said recombinant nucleic acid molecule comprises a native gene encoding a small secreted protein flanked upstream by 250 kb non-coding DNA and downstream by 250 kb non-coding DNA.
  • the ORF of said small secreted protein is flanked downstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, preferably at least 1100, more preferably at least 1200 nucleotides non-coding DNA.
  • said ORF is downstream flanked by at least 1296 nucleotides non-coding DNA.
  • said ORF is flanked upstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, preferably at least 1700, more preferably at least 1800 nucleotides non-coding DNA.
  • said ORF is upstream flanked by at least 1847 nucleotides non-coding DNA.
  • the landing pad or recombinant expression construct introduced into the genome of the fungi host cell is replacing in part or completely the gene encoding the small secreted protein comprised in said native genomic region. More preferably it is replacing all or part of the ORF encoding the small secreted protein.
  • a further embodiment of the invention is the method as defined before wherein said small secreted protein is comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 46 b. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99%to SEQ ID NO: 46, and c.
  • a further embodiment of the invention is any of the methods as defined above wherein said native genomic region is flanked at one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 41 b.
  • a further embodiment of the invention is any of the methods as defined above wherein the filamentous fungi host cell is from a species selected from the group consisting of of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysosporium, Filibasidium, Fusarium, Humicola, Magnaporthe, Monascus, Mu- cor, Myceliophthora, Mortierella, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Piro- myces, Phanerochaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordaria, Tala- romyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypocladium, Trametes and Trichoderma, more preferred As
  • a further embodiment of the invention is a recombinant filamentous fungi host cell comprising a recombinant nucleic acid molecule flanked by genomic DNA comprising on one side of said recombinant nucleic acid molecule a sequence selected from the list consisting of i) a nucleic acid sequence comprising 1296 consecutive bases from the 5- prime end of SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least
  • 600 preferably at least 700, preferably at least 800, preferably at least
  • 900 more preferably at least 1000, more preferably at least 1100, even more preferably at least 1200, most preferably at least 1296 consecutive nucleic acids of i) or ii). and comprising on the other side of said recombinant nucleic acid molecule a sequence selected from the list consisting of iv) a nucleic acid sequence comprising 1847 consecutive bases from the 3- prime end of SEQ ID NO: 44 v) a nucleic acid sequence with an identity of at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least
  • 600 preferably at least 700, preferably at least 800, preferably at least
  • 900 preferably at least 1000, preferably at least 1100, more preferably at least 1200, preferably at least 1300, preferably at least 1400, preferably at least 1500, more preferably at least 1600, more preferably at least 1700, even more preferably at least 1800, most preferably at least 1847 consecutive nucleic acids of iv) or v).
  • a further embodiment of the invention is a recombinant filamentous fungi host cell comprising a recombinant nucleic acid molecule flanked at one side by the XP_003662454.1 gene and at the other side by the XP_003662452.1 gene, preferably wherein the XP_003662454.1 gene is encoding an amino acid molecule selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 43, b.
  • XP_003662452.1 gene is encoding an amino acid molecule selected from the list consisting of d. an amino acid sequence comprising SEQ ID NO: 41, e.
  • a further embodiment of the invention is the recombinant filamentous fungi host cell as defined above wherein the host cell is from a species selected from the group consisting of of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysosporium, Filibasidium, Fusarium, Humicola, Magnaporthe, Monascus, Mu- cor, Myceliophthora, Mortierella, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Piro- myces, Phanerochaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordaria, Tala- romyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypocladium, Trametes and Trichoderma
  • An additional embodiment of the invention is a recombinant nucleic acid molecule flanked by distinct homology arms each at least 70% identical, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% most preferably identical over the whole length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 1000, even more preferably
  • a further embodiment is a recombinant nucleic acid molecule flanked by distinct homology arms each at least 70% identical, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% most preferably identical over the whole length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 1000, even more preferably at least 10
  • a further embodiment of the invention is a vector comprising any of the recombinant nucleic acid molecules as defined above.
  • Coding region As used herein the terms "coding region” or “open reading frame’’ or “ORF” when used in reference to a structural gene refers to the nucleotide sequences which encode the amino acids found in the nascent polypeptide as a result of translation of a mRNA molecule.
  • the coding region is bounded, in eukaryotes, on the 5'-side by the nucleotide triplet "ATG” which encodes the initiator methionine and on the 3'-side by one of the three triplets which specify stop codons (i.e., TAA, TAG, TGA).
  • genomic forms of a gene may also include sequences located on both the 5'- and 3'-end of the sequences which are present on the RNA transcript. These sequences are referred to as "flanking" sequences or regions (these flanking sequences are located 5' or 3' to the non-translated sequences present on the mRNA transcript).
  • the 5'-flanking region may contain regulatory sequences such as promoters and enhancers which control or influence the transcription of the gene.
  • the 3'- flanking region may contain sequences which direct the termination of transcription, post- transcriptional cleavage and polyadenylation.
  • Complementary refers to two nucleotide sequences which comprise antiparallel nucleotide sequences capable of pairing with one another (by the base-pairing rules) upon formation of hydrogen bonds between the complementary base residues in the antiparallel nucleotide sequences.
  • sequence 5'-AGT-3' is complementary to the sequence 5'-ACT-3'.
  • Complementarity can be "partial” or “total.”
  • Partial complementarity is where one or more nucleic acid bases are not matched according to the base pairing rules.
  • Total or “complete” complementarity between nucleic acid molecules is where each and every nucleic acid base is matched with another base under the base pairing rules.
  • a "complement" of a nucleic acid sequence as used herein refers to a nucleotide sequence whose nucleic acid molecules show total complementarity to the nucleic acid molecules of the nucleic acid sequence.
  • Endogenous nucleotide sequence refers to a nucleotide sequence, which is present in the genome of the untransformed cell.
  • Enhanced expression “enhance” or “increase” the expression of a nucleic acid molecule in a cell are used equivalently herein and mean that the level of expression of the nucleic acid molecule in a cell after applying a method of the present invention is higher than its expression in the cell before applying the method, or compared to a reference cell lacking a recombinant nucleic acid molecule of the invention.
  • the term "enhanced” or “increased” as used herein are synonymous and means herein higher, preferably significantly higher expression of the nucleic acid molecule to be expressed.
  • an “enhancement” or “increase” of the level of an agent such as a protein, mRNA or RNA means that the level is increased relative to a substantially identical cell grown under substantially identical conditions, lacking a recombinant nucleic acid molecule of the invention, the recombinant construct or recombinant vector of the invention.
  • “enhancement” or “increase” of the level of an agent means that the level is increased 50% or more, for example 100% or more, preferably 200% or more, more preferably 5 fold or more, even more preferably 10 fold or more, most preferably 20 fold or more for example 50 fold relative to a cell or organism lacking a recombinant nucleic acid molecule of the invention.
  • the enhancement or increase can be determined by methods with which the skilled worker is familiar.
  • the enhancement or increase of the nucleic acid or protein quantity can be determined for example by an immunological detection of the protein.
  • techniques such as protein assay, fluorescence, Northern hybridization, nuclease protection assay, reverse transcription (quantitative RT-PCR), ELISA (enzyme-linked immunosorbent assay), Western blotting, radioimmunoassay (RIA) or other immunoassays and fluorescence-activated cell analysis (FACS) can be employed to measure a specific protein or RNA in a cell.
  • RIA radioimmunoassay
  • FACS fluorescence-activated cell analysis
  • Expression refers to the biosynthesis of a gene product, preferably to the transcription and/or translation of a nucleotide sequence, for example an endogenous gene or a heterologous gene, in a cell.
  • expression involves transcription of the structural gene into mRNA and - optionally - the subsequent translation of mRNA into one or more polypeptides. In other cases, expression may refer only to the transcription of the DNA harboring an RNA molecule.
  • Expression construct as used herein mean a DNA sequence capable of directing expression of a particular nucleotide sequence in a cell, comprising a promoter functional in said cell into which it will be introduced, operatively linked to the nucleotide sequence of interest which is - optionally - operatively linked to termination signals. If translation is required, it also typically comprises sequences required for proper translation of the nucleotide sequence.
  • the coding region may code for a protein of interest but may also code for a functional RNA of interest, for example RNAa, siRNA, snoRNA, snRNA, microRNA, ta-siRNA or any other noncoding regulatory RNA, in the sense or antisense direction.
  • the expression construct comprising the nucleotide sequence of interest may be chimeric, meaning that one or more of its components is heterologous with respect to one or more of its other components.
  • the expression construct may also be one, which is naturally occurring but has been obtained in a recombinant form useful for heterologous expression.
  • the expression construct is heterologous with respect to the host, i.e., the particular DNA sequence of the expression construct does not occur naturally in the host cell and must have been introduced into the host cell or an ancestor of the host cell by a transformation event.
  • the expression of the nucleotide sequence in the expression construct may be under the control of a constitutive promoter or of an inducible promoter, which initiates transcription only when the host cell is exposed to some particular external stimulus.
  • the promoter can also be specific to a particular stage of development.
  • Foreign refers to any nucleic acid molecule (e.g., gene sequence) which is introduced into the genome of a cell by experimental manipulations and may include sequences found in that cell so long as the introduced sequence contains some modification (e.g., a point mutation, the presence of a selectable marker gene, etc.) and is therefore distinct relative to the naturally-occurring sequence.
  • nucleic acid molecule e.g., gene sequence
  • some modification e.g., a point mutation, the presence of a selectable marker gene, etc.
  • Functional linkage or “functionally linked” are to be understood as meaning, for example, the sequential arrangement of a regulatory element (e.g. a promoter) with a nucleic acid sequence to be expressed and, if appropriate, further regulatory elements (such as e.g., a terminator) in such a way that each of the regulatory elements can fulfil its intended function to allow, modify, facilitate or otherwise influence expression of said nucleic acid sequence.
  • a regulatory element e.g. a promoter
  • further regulatory elements such as e.g., a terminator
  • operble linkage or “operably linked” may be used.
  • the expression may result depending on the arrangement of the nucleic acid sequences in relation to sense or antisense RNA. To this end, direct linkage in the chemical sense is not necessarily required.
  • Genetic control sequences such as, for example, enhancer sequences, can also exert their function on the target sequence from positions which are further away, or indeed from other DNA molecules.
  • Preferred arrangements are those in which the nucleic acid sequence to be expressed recombinantly is positioned behind the sequence acting as promoter, so that the two sequences are linked covalently to each other.
  • the distance between the promoter sequence and the nucleic acid sequence to be expressed recombinantly is preferably less than 200 base pairs, especially preferably less than 100 base pairs, very especially preferably less than 50 base pairs.
  • the nucleic acid sequence to be transcribed is located behind the promoter in such a way that the transcription start is identical with the desired beginning of the chimeric RNA of the invention.
  • Functional linkage, and an expression construct can be generated by means of customary recombination and cloning techniques as described (e.g., in Maniatis T, Fritsch EF and Sambrook J (1989) Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory, Cold Spring Harbor (NY); Silhavy et al. (1984) Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor (NY); Au- subel et al. (1987) Current Protocols in Molecular Biology, Greene Publishing Assoc, and Wiley Interscience).
  • sequences which, for example, act as a linker with specific cleavage sites for restriction enzymes, or as a signal peptide, may also be positioned between the two sequences.
  • the insertion of sequences may also lead to the expression of fusion proteins.
  • the expression construct consisting of a linkage of a regulatory region for example a promoter and nucleic acid sequence to be expressed, can exist in a vector-integrated form and be inserted into the genome, for example by transformation.
  • Gene refers to a region operably joined to appropriate regulatory sequences capable of regulating the expression of the gene product (e.g., a polypeptide or a functional RNA) in some manner.
  • a gene includes untranslated regulatory regions of DNA (e.g., promoters, enhancers, repressors, etc.) preceding (up-stream) and following (downstream) the coding region (open reading frame, ORF) as well as, where applicable, intervening sequences (i.e., introns) between individual coding regions (i.e., exons).
  • constructural gene as used herein is intended to mean a DNA sequence that is transcribed into mRNA which is then translated into a sequence of amino acids characteristic of a specific polypeptide.
  • Genome and genomic DNA The terms “genome” or “genomic DNA” is referring to the heritable genetic information of a host organism. Preferably the terms genome or genomic DNA is referring to the chromosomal DNA of the nucleus.
  • Heterologous refers to a nucleic acid molecule which is operably linked to, or is manipulated to become operably linked to, a second nucleic acid molecule, e.g. a promoter to which it is not operably linked in nature, e.g. in the genome of a WT cell, or to which it is operably linked at a different location or position in nature, e.g. in the genome of a WT cell.
  • heterologous with respect to a nucleic acid molecule or DNA, e.g. promoter refers to a nucleic acid molecule which is operably linked to, or is manipulated to become operably linked to, a second nucleic acid molecule, e.g. a ORF to which it is not operably linked in nature.
  • a heterologous expression construct comprising a nucleic acid molecule and one or more regulatory nucleic acid molecule (such as a promoter or a transcription termination signal) linked thereto for example is a constructs originating by experimental manipulations in which either a) said nucleic acid molecule, or b) said regulatory nucleic acid molecule or c) both (i.e. (a) and (b)) is not located in its natural (native) genetic environment or has been modified by experimental manipulations, an example of a modification being a substitution, addition, deletion, inversion or insertion of one or more nucleotide residues.
  • Natural genetic environment refers to the natural chromosomal locus in the organism of origin, or to the presence in a genomic library.
  • the natural genetic environment of the sequence of the nucleic acid molecule is preferably retained, at least in part.
  • the environment flanks the nucleic acid sequence at least at one side and has a sequence of at least 50 bp, preferably at least 500 bp, especially preferably at least 1 ,000 bp, very especially preferably at least 5,000 bp, in length.
  • a naturally occurring expression construct for example the naturally occurring combination of a promoter with the corresponding gene - becomes a transgenic expression construct when it is modified by non-natural, synthetic “artificial” methods such as, for example, mutagenization. Such methods have been described (US 5,565,350; WO 00/15815).
  • a protein encoding nucleic acid molecule operably linked to a promoter is considered to be heterologous with respect to the promoter.
  • heterologous DNA is not endogenous to or not naturally associated with the cell into which it is introduced, but has been obtained from another cell or has been synthesized.
  • Heterologous DNA also includes an endogenous DNA sequence, which contains some modification, non-naturally occurring, multiple copies of an endogenous DNA sequence, or a DNA sequence which is not naturally associated with another DNA sequence physically linked thereto.
  • Hybridization is a process wherein substantially complementary nucleotide sequences anneal to each other.
  • the hybridisation process can occur entirely in solution, i.e. both complementary nucleic acids are in solution.
  • the hybridisation process can also occur with one of the complementary nucleic acids immobilised to a matrix such as magnetic beads, Sepharose beads or any other resin.
  • the hybridisation process can furthermore occur with one of the complementary nucleic acids immobilised to a solid support such as a nitro-cellulose or nylon membrane or immobilised by e.g. photolithography to, for example, a siliceous glass support (the latter known as nucleic acid arrays or microarrays or as nucleic acid chips).
  • the nucleic acid molecules are generally thermally or chemically denatured to melt a double strand into two single strands and/or to remove hairpins or other secondary structures from single stranded nucleic acids.
  • stringency refers to the conditions under which a hybridisation takes place.
  • the stringency of hybridisation is influenced by conditions such as temperature, salt concentration, ionic strength and hybridisation buffer composition. Generally, low stringency conditions are selected to be about 30°C lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH. Medium stringency conditions are when the temperature is 20°C below Tm, and high stringency conditions are when the temperature is 10°C below Tm. High stringency hybridisation conditions are typically used for isolating hybridising sequences that have high sequence similarity to the target nucleic acid sequence. However, nucleic acids may deviate in sequence and still encode a substantially identical polypeptide, due to the degeneracy of the genetic code. Therefore, medium stringency hybridisation conditions may sometimes be needed to identify such nucleic acid molecules.
  • the “Tm” is the temperature under defined ionic strength and pH, at which 50% of the target sequence hybridises to a perfectly matched probe.
  • the Tm is dependent upon the solution conditions and the base composition and length of the probe. For example, longer sequences hybridise specifically at higher temperatures.
  • the maximum rate of hybridisation is obtained from about 16°C up to 32°C below Tm.
  • the presence of monovalent cations in the hybridisation solution reduce the electrostatic repulsion between the two nucleic acid strands thereby promoting hybrid formation; this effect is visible for sodium concentrations of up to 0.4M (for higher concentrations, this effect may be ignored).
  • Formamide reduces the melting temperature of DNA- DNA and DNA-RNA duplexes with 0.6 to 0.7°C for each percent formamide, and addition of 50% formamide allows hybridisation to be performed at 30 to 45°C, though the rate of hybridisation will be lowered.
  • Base pair mismatches reduce the hybridisation rate and the thermal stability of the duplexes.
  • the Tm decreases about 1°C per % base mismatch. The Tm may be calculated using the following equations, depending on the types of hybrids:
  • Tm 22 + 1.46 (In ) a or for other monovalent cation, but only accurate in the 0.01-0.4 M range.
  • b only accurate for %GC in the 30% to 75% range.
  • c L length of duplex in base pairs.
  • Non-specific binding may be controlled using any one of a number of known techniques such as, for example, blocking the membrane with protein containing solutions, additions of heterologous RNA, DNA, and SDS to the hybridisation buffer, and treatment with Rnase.
  • a series of hybridizations may be performed by varying one of (i) progressively lowering the annealing temperature (for example from 68°C to 42°C) or (ii) progressively lowering the formamide concentration (for example from 50% to 0%).
  • progressively lowering the annealing temperature for example from 68°C to 42°C
  • formamide concentration for example from 50% to 0%
  • hybridisation typically also depends on the function of post-hybridisation washes.
  • samples are washed with dilute salt solutions.
  • Critical factors of such washes include the ionic strength and temperature of the final wash solution: the lower the salt concentration and the higher the wash temperature, the higher the stringency of the wash.
  • Wash conditions are typically performed at or below hybridisation stringency. A positive hybridisation gives a signal that is at least twice of that of the background.
  • suitable stringent conditions for nucleic acid hybridisation assays or gene amplification detection procedures are as set forth above. More or less stringent conditions may also be selected. The skilled artisan is aware of various parameters which may be altered during washing and which will either maintain or change the stringency conditions.
  • typical high stringency hybridisation conditions for DNA hybrids longer than 50 nucleotides encompass hybridisation at 65°C in 1x SSC or at 42°C in 1x SSC and 50% formamide, followed by washing at 65°C in 0.3x SSC.
  • Examples of medium stringency hybridisation conditions for DNA hybrids longer than 50 nucleotides encompass hybridisation at 50°C in 4x SSC or at 40°C in 6x SSC and 50% formamide, followed by washing at 50°C in 2x SSC.
  • the length of the hybrid is the anticipated length for the hybridising nucleic acid. When nucleic acids of known sequence are hybridised, the hybrid length may be determined by aligning the sequences and identifying the conserved regions described herein.
  • 1 xSSC is 0.15M NaCI and 15mM sodium citrate; the hybridisation solution and wash solutions may additionally include 5x Denhardt's reagent, 0.5-1.0% SDS, 100 pg/ml denatured, fragmented salmon sperm DNA, 0.5% sodium pyrophosphate.
  • 5x Denhardt's reagent 0.5-1.0% SDS
  • 100 pg/ml denatured, fragmented salmon sperm DNA 0.5% sodium pyrophosphate.
  • Another example of high stringency conditions is hybridisation at 65°C in 0.1 x SSC comprising 0.1 SDS and optionally 5x Denhardt's reagent, 100 pg/ml denatured, fragmented salmon sperm DNA, 0.5% sodium pyrophosphate, followed by the washing at 65°C in 0.3x SSC.
  • Identity when used in respect to the comparison of two or more nucleic acid or amino acid molecules means that the sequences of said molecules share a certain degree of sequence similarity, the sequences being partially identical.
  • Enzyme variants may be defined by their sequence identity when compared to a parent enzyme or nucleic acid molecule. Sequence identity usually is provided as “% sequence identity” or “% identity”. To determine the percent-identity between two amino acid sequences in a first step a pairwise sequence alignment is generated between those two sequences, wherein the two sequences are aligned over their complete length (i.e., a pairwise global alignment). The alignment is generated with a program implementing the Needleman and Wunsch algorithm (J. Mol. Biol. (1979) 48, p.
  • sequence B is sequence B.
  • Seq B — GAT-CTGA
  • the “I” symbol in the alignment indicates identical residues (which means bases for DNA or amino acids for proteins). The number of identical residues is 6.
  • the symbol in the alignment indicates gaps.
  • the number of gaps introduced by alignment within the Seq B is 1 .
  • the number of gaps introduced by alignment at borders of Seq B is 2, and at borders of Seq A is 1.
  • the alignment length showing the aligned sequences over their complete length is 10.
  • the alignment length showing the shorter sequence over its complete length is 8 (one gap is present which is factored in the alignment length of the shorter sequence).
  • the alignment length showing Seq A over its complete length would be 9 (meaning Seq A is the sequence of the invention).
  • the alignment length showing Seq B over its complete length would be 8 (meaning Seq B is the sequence of the invention).
  • an identity value is determined from the alignment produced.
  • sequence identity in relation to comparison of two amino acid sequences according to this embodiment is calculated by dividing the number of identical residues by the length of the alignment region which is showing the respective sequence of this invention over its complete length. This value is multiplied with 100 to give “%- identity”.
  • introduction means any introduction of the sequence of the donor DNA molecule into the target region for example by the physical integration of the donor DNA molecule or a part thereof into the target region or the introduction of the sequence of the donor DNA molecule or a part thereof into the target region wherein the donor DNA is used as template for a polymerase.
  • Isogenic organisms (e.g., fungi), which are genetically identical, except that they may differ by the presence or absence of a heterologous DNA sequence.
  • Isolated means that a material has been removed by the hand of man and exists apart from its original, native environment and is therefore not a product of nature.
  • An isolated material or molecule (such as a DNA molecule or enzyme) may exist in a purified form or may exist in a non-native environment such as, for example, in a transgenic host cell.
  • a naturally occurring polynucleotide or polypeptide present in a living cell is not isolated, but the same polynucleotide or polypeptide, separated from some or all of the coexisting materials in the natural system, is isolated.
  • Such polynucleotides can be part of a vector and/or such polynucleotides or polypeptides could be part of a composition and would be isolated in that such a vector or composition is not part of its original environment.
  • isolated when used in relation to a nucleic acid molecule, as in "an isolated nucleic acid sequence” refers to a nucleic acid sequence that is identified and separated from at least one contaminant nucleic acid molecule with which it is ordinarily associated in its natural source. Isolated nucleic acid molecule is nucleic acid molecule present in a form or setting that is different from that in which it is found in nature.
  • non-isolated nucleic acid molecules are nucleic acid molecules such as DNA and RNA, which are found in the state they exist in nature.
  • a given DNA sequence e.g., a gene
  • RNA sequences such as a specific mRNA sequence encoding a specific protein, are found in the cell as a mixture with numerous other mRNAs, which encode a multitude of proteins.
  • an isolated nucleic acid sequence comprising for example SEQ ID NO: 1 includes, by way of example, such nucleic acid sequences in cells which ordinarily contain SEQ ID NO:1 where the nucleic acid sequence is in a chromosomal or ex- trachromosomal location different from that of natural cells or is otherwise flanked by a different nucleic acid sequence than that found in nature.
  • the isolated nucleic acid sequence may be present in single-stranded or double-stranded form.
  • the nucleic acid sequence will contain at a minimum at least a portion of the sense or coding strand (i.e., the nucleic acid sequence may be single-stranded). Alternatively, it may contain both the sense and anti-sense strands (i.e., the nucleic acid sequence may be double-stranded).
  • Minimal Promoter promoter elements, particularly a TATA element, that are inactive or that have greatly reduced promoter activity in the absence of upstream activation. In the presence of a suitable transcription factor, the minimal promoter functions to permit transcription.
  • Non-coding refers to sequences of nucleic acid molecules that do not encode part or all of an expressed protein. Non-coding sequences include but are not limited to functional RNAs, introns, enhancers, promoter regions, 3' untranslated regions, and 5' untranslated regions.
  • Nucleic acids and nucleotides refer to naturally occurring or synthetic or artificial nucleic acid or nucleotides.
  • nucleic acids and “nucleotides” comprise deoxyribonucleotides or ribonucleotides or any nucleotide analogue and polymers or hybrids thereof in either single- or double-stranded, sense or antisense form.
  • a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated.
  • nucleic acid is used interchangeably herein with “gene”, “cDNA, “mRNA”, “oligonucleotide,” and “polynucleotide”.
  • Nucleotide analogues include nucleotides having modifications in the chemical structure of the base, sugar and/or phosphate, including, but not limited to, 5-position pyrimidine modifications, 8- position purine modifications, modifications at cytosine exocyclic amines, substitution of 5- bromo-uracil, and the like; and 2'-position sugar modifications, including but not limited to, sug- ar-modified ribonucleotides in which the 2'-OH is replaced by a group selected from H, OR, R, halo, SH, SR, NH2, NHR, NR2, or CN.
  • Short hairpin RNAs also can comprise nonnatural elements such as non-natural bases, e.g., ionosin and xanthine, non-natural sugars, e.g., 2'-methoxy ribose, or non-natural phosphodiester linkages, e.g., methylphosphonates, phosphorothioates and peptides.
  • non-natural bases e.g., ionosin and xanthine
  • non-natural sugars e.g., 2'-methoxy ribose
  • non-natural phosphodiester linkages e.g., methylphosphonates, phosphorothioates and peptides.
  • nucleic acid sequence refers to a single or doublestranded polymer of deoxyribonucleotide or ribonucleotide bases read from the 5'- to the 3'-end. It includes chromosomal DNA, self-replicating plasmids, infectious polymers of DNA or RNA and DNA or RNA that performs a primarily structural role. "Nucleic acid sequence” also refers to a consecutive list of abbreviations, letters, characters or words, which represent nucleotides.
  • a nucleic acid can be a "probe” which is a relatively short nucleic acid, usually less than 100 nucleotides in length.
  • nucleic acid probe is from about 50 nucleotides in length to about 10 nucleotides in length.
  • a "target region” of a nucleic acid is a portion of a nucleic acid that is identified to be of interest.
  • a “coding region” of a nucleic acid is the portion of the nucleic acid, which is transcribed and translated in a sequence-specific manner to produce into a particular polypeptide or protein when placed under the control of appropriate regulatory sequences. The coding region is said to encode such a polypeptide or protein.
  • Oligonucleotide refers to an oligomer or polymer of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA) or mimetics thereof, as well as oligonucleotides having non-naturally-occurring portions which function similarly. Such modified or substituted oligonucleotides are often preferred over native forms because of desirable properties such as, for example, enhanced cellular uptake, enhanced affinity for nucleic acid target and increased stability in the presence of nucleases.
  • An oligonucleotide preferably includes two or more nucleomono- mers covalently coupled to each other by linkages (e.g., phosphodiesters) or substitute linkages.
  • Overhang is a relatively short single-stranded nucleotide sequence on the 5'- or 3'-hydroxyl end of a double-stranded oligonucleotide molecule (also referred to as an "extension,” “protruding end,” or “sticky end”).
  • Polypeptide The terms “polypeptide”, “peptide”, “oligopeptide”, “polypeptide”, “gene product”, “expression product” and “protein” are used interchangeably herein to refer to a polymer or oligomer of consecutive amino acid residues.
  • “Precise” with respect to the introduction of a donor DNA molecule in target region means that the sequence of the donor DNA molecule is introduced into the target region without any InDeis, duplications or other mutations as compared to the unaltered DNA sequence of the target region that are not comprised in the donor DNA molecule sequence.
  • promoter refers to a DNA sequence which when operably linked to a nucleotide sequence of interest is capable of controlling the transcription of the nucleotide sequence of interest into RNA.
  • a promoter is located 5' (i.e. , upstream), proximal to the transcriptional start site of a nucleotide sequence of interest whose transcription into RNA it controls, and provides a site for specific binding by RNA polymerase and other transcription factors for initiation of transcription.
  • Said pro- moter comprises for example the at least 10 kb, for example 5 kb or 2 kb proximal to the transcription start site.
  • the promoter may also comprise the at least 1500 bp proximal to the transcriptional start site, preferably the at least 1000 bp.
  • the promoter comprises the at least 50 bp proximal to the transcription start site, for example, at least 25 bp.
  • the promoter does not comprise exon and/or intron regions or 5' untranslated regions.
  • the promoter may for example be heterologous or homologous to the respective cell.
  • a polynucleotide sequence is "heterologous to" an organism or a second polynucleotide sequence if it originates from a foreign species, or, if from the same species, is modified from its original form.
  • a promoter operably linked to a heterologous coding sequence refers to a coding sequence from a species different from that from which the promoter was derived, or, if from the same species, a coding sequence which is not naturally associated with the promoter (e.g. a genetically engineered coding sequence or an allele from a different ecotype or variety).
  • Suitable promoters can be derived from genes of the host cells where expression should occur or from pathogens for this host cells.
  • Activity of a promoter may be evaluated by, for example, operably linking a reporter gene to the promoter sequence to generate a reporter construct, introducing the reporter construct into the genome of a cell and detecting the expression of the reporter gene (e.g., detecting mRNA, protein, or the activity of a protein encoded by the reporter gene).
  • purified refers to molecules, either nucleic or amino acid sequences that are removed from their natural environment, isolated or separated. “Substantially purified” molecules are at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which they are naturally associated.
  • a purified nucleic acid sequence may be an isolated nucleic acid sequence.
  • Recombinant refers to nucleic acid molecules produced by recombinant DNA techniques.
  • Recombinant nucleic acid molecules may also comprise molecules, which as such does not exist in nature but are modified, changed, mutated or otherwise manipulated by man.
  • a "recombinant nucleic acid molecule” is a non-naturally occurring nucleic acid molecule that differs in sequence from a naturally occurring nucleic acid molecule by at least one nucleic acid.
  • a “recombinant nucleic acid molecule” may also comprise a “recombinant construct” which comprises, preferably operably linked, a sequence of nucleic acid molecules not naturally occurring in that order.
  • Preferred methods for producing said recombinant nucleic acid molecule may comprise cloning techniques, directed or non-directed mutagenesis, synthesis or recombination techniques.
  • Sense is understood to mean a nucleic acid molecule having a sequence which is complementary or identical to a target sequence, for example a sequence which binds to a protein transcription factor and which is involved in the expression of a given gene.
  • the nucleic acid molecule comprises a gene of interest and elements allowing the expression of the said gene of interest.
  • an increase or decrease for example in enzymatic activity or in gene expression, that is larger than the margin of error inherent in the measurement technique, preferably an increase or decrease by about 2-fold or greater of the activity of the control enzyme or expression in the control cell, more preferably an increase or decrease by about 5- fold or greater, and most preferably an increase or decrease by about 10-fold or greater.
  • Small nucleic acid molecules are understood as molecules consisting of nucleic acids or derivatives thereof such as RNA or DNA. They may be doublestranded or single-stranded and are between about 15 and about 30 bp, for example between 15 and 30 bp, more preferred between about 19 and about 26 bp, for example between 19 and 26 bp, even more preferred between about 20 and about 25 bp for example between 20 and 25 bp.
  • the oligonucleotides are between about 21 and about 24 bp, for example between 21 and 24 bp.
  • the small nucleic acid molecules are about 21 bp and about 24 bp, for example 21 bp and 24 bp.
  • substantially complementary when used herein with respect to a nucleotide sequence in relation to a reference or target nucleotide sequence, means a nucleotide sequence having a percentage of identity between the substantially complementary nucleotide sequence and the exact complementary sequence of said reference or target nucleotide sequence of at least 60%, more desirably at least 70%, more desirably at least 80% or 85%, preferably at least 90%, more preferably at least 93%, still more preferably at least 95% or 96%, yet still more preferably at least 97% or 98%, yet still more preferably at least 99% or most preferably 100% (the latter being equivalent to the term “identical” in this context).
  • identity is assessed over a length of at least 19 nucleotides, preferably at least 50 nucleotides, more preferably the entire length of the nucleic acid sequence to said reference sequence (if not specified otherwise below). Sequence comparisons are carried out using default GAP analysis with the University of Wisconsin GCG, SEQWEB application of GAP, based on the algorithm of Needleman and Wunsch (Needleman and Wun- sch (1970) J Mol. Biol. 48: 443-453; as defined above).
  • a nucleotide sequence "substantially complementary " to a reference nucleotide sequence hybridizes to the reference nucleotide sequence under low stringency conditions, preferably medium stringency conditions, most preferably high stringency conditions (as defined above).
  • Target site means the position in the genome at which a double strand break or one or a pair of single strand breaks (nicks) are induced using recombinant technologies such as Zn-finger, TALEN, restriction enzymes, homing endonucleases, RNA-guided nucleases, RNA-guided nickases such as CRISPR/Cas nucleases or nickases and the like.
  • transgene refers to any nucleic acid sequence, which is introduced into the genome of a cell by experimental manipulations.
  • a transgene may be an "endogenous DNA sequence," or a “heterologous DNA sequence” (i.e., “foreign DNA”).
  • endogenous DNA sequence refers to a nucleotide sequence, which is naturally found in the cell into which it is introduced so long as it does not contain some modification (e.g., a point mutation, the presence of a selectable marker gene, etc.) relative to the naturally-occurring sequence.
  • transgenic when referring to an organism means transformed, preferably stably transformed, with a recombinant DNA molecule that preferably comprises a suitable promoter operatively linked to a DNA sequence of interest.
  • Vector refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it has been linked.
  • a genomic integrated vector or "integrated vector” which can become integrated into the chromosomal DNA of the host cell.
  • Another type of vector is an episomal vector, i.e., a nucleic acid molecule capable of extra-chromosomal replication.
  • Vectors capable of directing the expression of genes to which they are operatively linked are referred to herein as "expression vectors”.
  • expression vectors vectors capable of directing the expression of genes to which they are operatively linked.
  • Expression vectors designed to produce RNAs as described herein in vitro or in vivo may contain sequences recognized by any RNA polymerase, including mitochondrial RNA polymerase, RNA pol I, RNA pol II, and RNA pol III. These vectors can be used to transcribe the desired RNA molecule in the cell according to this invention.
  • Wild-type The term “wild-type”, “natural” or “natural origin” means with respect to an organism, polypeptide, or nucleic acid sequence, that said organism is naturally occurring or available in at least one naturally occurring organism which is not changed, mutated, or otherwise manipulated by man.
  • Figure 1 shows the total specific phytase activity in U/g protein of the different phytase production strains.
  • Figure 2 shows the total specific glucanase activity in U/g protein of the different glucanase production strains.
  • Figure 3 shows the total specific phytase activity in U/g protein of the different phytase production strains.
  • Figure 5 shows the genomic organization of the genomic integration site of the invention in T. thermophilus) C1.
  • cloning procedures carried out for the purposes of the present invention including restriction digest, agarose gel electrophoresis, purification of nucleic acids, Ligation of nucleic acids, transformation, selection and cultivation of bacterial cells were performed as described (Sambrook et al., 1989). Sequence analyses of recombinant DNA were performed with a laser fluorescence DNA sequencer (Applied Biosystems, Foster City, CA, USA) using the Sanger technology (Sanger et al., 1977). Unless described otherwise, chemicals and reagents were obtained from Sigma Aldrich (Sigma Aldrich, St.
  • T. thermophilus protoplasts were prepared by inoculating a 25 ml preculture of a standard fungal growth media with 0.7-1x 10 5 spores/ml in a 100 ml shake flask for 24 h at 37°C and 250 rpm.
  • the main culture was prepared by inoculating 100 ml of a standard fungal growth media with 20 ml of the preculture in a 500 ml shake flask for 24 h at 37°C and 250 rpm.
  • the mycelium was harvested by filtration through a sterile Cell Strainer (VWR) and washed with 100 ml 2000 mosmol/L NaCI/CaCI 2 (0.6 M NaCI and 0.27 M CaCI 2 *H 2 O). 1 g of the washed mycelium was transferred into a 100 ml flask. The mycelium was mixed with 150 mg VinoTaste Pro solution (2.5 mg/ml in 2000 mosmol/L NaCI/CaCI 2 ) and 10 mg Yatalase solution (0.625 mg/ ml in 2000 mosmol/L NaCI/CaCI 2 ) and 10 ml of 2000 mosmol/L NaCI/CaCI 2 .
  • VWR sterile Cell Strainer
  • the mycelium suspension was incubated at 30°C and 70 rpm for 50-70 min until protoplasts are visible under the microscope.
  • Harvesting of protoplasts was done by filtration through a sterile Cell Strainer into a sterile 50 ml tube.
  • 25 ml ice-cold STC solution 1.2 M sorbitol, 50 mM CaCI 2 , 35 mM NaCI, 10 mM Tris/HCI pH 7.5
  • the protoplasts were harvested by centrifugation (1200 x g, 10 min, 4°C).
  • the protoplasts were washed again in 50 ml STC and resuspended in 0.5-1.2 ml STC.
  • amdS gene is used as selection marker, enriched minimal medium with uridine and uracil and with acetamide is used to select positive transformants (sucrose is only added in case protoplasts are plated):
  • nat1 gene is used as additional selection marker, enriched minimal medium without uridine and uracil and with nourseothricin (clonNAT) is used to select positive transformants (sucrose is only added in case protoplasts are plated):
  • Agar 16 g/L set pH to 6.5
  • a synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from bacterial origin (disclosed in WO 2012/143862; as phytase PhV-99; SEQ ID NO. 2) was used for the construction of a phytase expression plasmids.
  • a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the phytase.
  • a promotor sequence amplified from the upstream region of the Chi1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the phytase.
  • the expression plasmid MT2701 (SEQ ID NO: 7) was constructed based on the E. coli standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5).
  • Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, pyr5 gene, upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination, and lacZ for blue/white screening.
  • Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the expression plasmid contains the upstream region of cbh1 from bases 1913 - 3345, the chi1 promotor sequence from bases 3350 - 5162, the phytase including a signal sequence from bases 5165 - 6478, the cbh1 terminator sequence from bases 6483 - 6689, and the downstream region of cbh1 from bases 7981 - 9610.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • Expression construct with phytase and TEF1 promoter was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from bacterial origin (disclosed in WO 2012/143862; as phytase PhV-99; SEQ ID NO. 2) was used for the construction of a phytase expression plasmids.
  • a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the phytase.
  • a promotor sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH:XP_003660173.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the phytase.
  • the expression plasmid MT2703 (SEQ ID NO: 8) was constructed based on the E. coli standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5).
  • Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, pyr5 gene, upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination, and lacZ for blue/white screening.
  • Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the expression plasmid contains the upstream region of cbh1 from bases 1913 - 3345, the TEF1 promotor sequence from bases 3350 - 5847, the phytase including a signal sequence from bases 5850 - 7163, the cbh1 terminator sequence from bases 7168 - 7374, and the downstream region of cbh1 from bases 8666 - 10295.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) encoding a glucanase from Rasamsonia emersonii (SEQ ID NO. 3) was used for the construction of a glucanase expression plasmids.
  • a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the glucanase.
  • a promotor sequence amplified from the upstream region of the Chi 1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the glucanase.
  • the expression plasmid MT2705 (SEQ ID NO: 9) was constructed based on the E. coli standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5).
  • Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, pyr5 gene, upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination, and lacZ for blue/white screening.
  • Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the expression plasmid contains the upstream region of cbh1 from bases 1913 - 3345, the chi1 promotor sequence from bases 3350 - 5162, the glu- canase including a signal sequence from bases 5165 - 6151 , the cbh1 terminator sequence from bases 6156 - 6362, and the downstream region of cbh1 from bases 7654 - 9283.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) encoding a glucanase from Rasamsonia emersonii (SEQ ID NO. 3) was used for the construction of a glucanase expression plasmids.
  • a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the glucanase.
  • a promotor sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the glucanase.
  • the expression plasmid MT2707 (SEQ ID NO: 10) was constructed based on the E. coll standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5).
  • Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, pyr5 gene, upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination, and lacZ for blue/white screening.
  • Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the expression plasmid contains the upstream region of cbh1 from bases 1913 - 3345, the TEF1 promotor sequence from bases 3350 - 5847, the glucanase including a signal sequence from bases 5850 - 6836, the cbh1 terminator sequence from bases 6841 - 7047, and the downstream region of cbh1 from bases 8339 - 9968.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from bacterial origin (disclosed in WO 2012/143862; as phytase PhV-99; SEQ ID NO. 2) was used for the construction of a phytase expression plasmids.
  • a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the phytase.
  • a promotor sequence amplified from the upstream region of the Chi1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the phytase.
  • the expression plasmid MT2709 (SEQ ID NO: 11) was constructed based on the E. coli standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6).
  • Both plasmids consist of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the expression plasmid contains the upstream region of MYCTH:XP_003662453.1 from bases 1913 - 3412, the chi1 promoter sequence from bases 3417 - 5229, the phytase including a signal sequence from bases 5232 - 6545, the cbh1 terminator sequence from bases 6550 - 6756, and the downstream region of MYCTH:XP_003662453.1 from bases 8048 - 9518.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from bacterial origin (disclosed in WO 2012/143862; as phytase PhV-99; SEQ ID NO. 2) was used for the construction of a phytase expression plasmids.
  • a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the phytase.
  • a promotor sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the phytase.
  • the expression plasmid MT2711 (SEQ ID NO: 12) was constructed based on the E. coli standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6).
  • Both plasmids consist of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the expression plasmid contains the upstream region of MYCTH:XP_003662453.1 from bases 1913 - 3412, the TEF1 promotor sequence from bases 3417 - 5914, the phytase including a signal sequence from bases 5917 - 7230, the cbh1 terminator sequence from bases 7235 - 7441 , and the downstream region of MYCTH:XP_003662453.1 from bases 8733 - 10203.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) encoding a glucanase from Rasamsonia emersonii (SEQ ID NO. 3) was used for the construction of a glucanase expression plasmids.
  • a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the glucanase.
  • a promotor sequence amplified from the upstream region of the Chi 1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the glucanase.
  • the expression plasmid MT2713 (SEQ ID NO: 13) was constructed based on the E. coll standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6).
  • Both plasmids consist of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the expression plasmid contains the upstream region of MYCTH:XP_003662453.1 from bases 1913 — 3412, the chi1 promotor sequence from bases 3417 - 5229, the glucanase including a signal sequence from bases 5232 - 6218, the cbh1 terminator sequence from bases 6223 - 6429, and the downstream region of MYCTH:XP_003662453.1 from bases 7721 - 9191 .
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) encoding a glucanase from Rasamsonia emersonii (SEQ ID NO. 3) was used for the construction of a glucanase expression plasmids.
  • a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the glucanase.
  • a promotor sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the glucanase.
  • the expression plasmid MT2715 (SEQ ID NO: 14) was constructed based on the E. coli stand- ard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6).
  • Both plasmids consist of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the expression plasmid contains the upstream region of MYCTH:XP_003662453.1 from bases 1913 — 3412, the TEF1 promotor sequence from bases 3417 - 5914, the glucanase including a signal sequence from bases 5917 - 6903, the cbh1 terminator sequence from bases 6908 - 7114, and the downstream region of MYCTH:XP_003662453.1 from bases 8406 - 9876.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic nourseothricin acetyltransferase gene (nat1) selection marker cassette (SEQ ID NO: 15) was fused with the upstream and downstream regions of the MYCTH:XP_003662453.1 locus from T. thermophilus for homologous recombination.
  • the KO plasmid MT2599 (SEQ ID NO: 16) was constructed based on the E. coli standard cloning vector MT944 (SEQ ID NO: 5). It consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the nat1 selection marker cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • a synthetic nourseothricin acetyltransferase gene (nat1) selection marker cassette (SEQ ID NO: 15) was fused with the upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination.
  • the KO plasmid MT2699 (SEQ ID NO: 17) was constructed based on the E. coli standard cloning vector MT944 (SEQ ID NO: 5). It consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
  • the plasmid was digested with Not to remove the vector backbone and the fragment containing the nat1 selection marker cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
  • thermophilus strains UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6
  • thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 was co-transformed as described in example 1 with the two isolated deletion fragments from plasmids MT185 (SEQ ID No. 18) and MT186 (SEQ ID NO. 19) to delete gene pep4.
  • Enriched minimal medium for amdS selection was used for incubation.
  • After re-streaking on enriched minimal medium with fluoroacetamide for counterselection the transformants were analyzed by PCR for the correct integration of the deletion cassettes in the targeted locus, for the disappearance of the intact gene, and for removal of the amdS marker.
  • the resulting strain UV18#100f Apyr5 Aalpl Aku70 Apep4 was modified in consecutive transformation rounds, and the following genes were deleted in the identical manner; gene prt1 with the two isolated deletion fragments from plasmids MT111 (SEQ ID No. 20) and MT112 (SEQ ID NO. 21), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl was obtained; gene prt2 with the two isolated deletion fragments from plasmids MT368 (SEQ ID No. 22) and MT453 (SEQ ID NO.
  • the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 (construction described in detail in WO 2017/093450) from the C1 lineage, a strain with uracil auxotrophy, reduced protease activity, and impaired non-homologous end joining (NHEJ) repair system, was transformed as described in example 1 with the /Vofl-digested and isolated phytase expression constructs (s. example 2) from plasmids MT2701 (SEQ ID NO: 7) and MT2703 (SEQ ID NO: 8).
  • thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 (construction described in detail in WO 2017/093450) was transformed as described in example 1 with the /Votl-digested and isolated glucanase expression constructs (s. example 2) from plasmids MT2705 (SEQ ID NO: 9) and MT2707 (SEQ ID NO: 10).
  • the transformants were incubated for 3-6 days at 37°C on enriched minimal medium for pyr5 selection to select for restored uracil prototrophy by complementing the pyr5 deletion with the pyr5 marker as known in the art.
  • Colonies were re-streaked and checked for the integration of the phytase or glucanase expression cassette using PCR with primer pairs specific for the phytase or glucanase expression cassette and the cbh1 locus as known in the art.
  • a transformant tested positive for the phytase or glucanase expression construct at the cbh1 locus was selected for further characterization.
  • the phytase production strain UV18#100f Apyr5 Aalpl Aku70 cbh1 ::Pchi-phytase-Tcbh1 was transformed as described in example 1 with the Not ⁇ -digested and isolated nat1 selection marker cassette (s.
  • example 4 from plasmid MT2599 (SEQ ID NO: 16).
  • the transformants were incubated for 3-6 days at 37°C on enriched minimal medium for nat1 selection to select for resistance to nourseothricin (clonNAT) by integrating the nat1 marker as known in the art.
  • Colonies were re-streaked and checked for the integration of the nat1 selection marker and the deletion of the MYCTH:XP_003662453.1 locus using PCR with primer pairs specific for nat1 and the MYCTH:XP_003662453.1 locus as known in the art.
  • the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 (construction described in detail in WO 2017/093450) from the C1 lineage, a strain with uracil auxotrophy, reduced protease activity, and impaired non-homologous end joining (NHEJ) repair system, was transformed as described in example 1 with the /Vofl-digested and isolated phytase expression constructs (s. example 3) from plasmids MT2709 (SEQ ID NO: 11) and MT2711 (SEQ ID NO: 12).
  • thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 (construction described in detail in WO 2017/093450) was transformed as described in example 1 with the A/otl-digested and isolated glucanase expression constructs (s. example 3) from plasmids MT2713 (SEQ ID NO: 13) and MT2715 (SEQ ID NO: 14).
  • the transformants were incubated for 3-6 days at 37°C on enriched minimal medium for pyr5 selection to select for restored uracil prototrophy by complementing the pyr5 deletion with the pyr5 marker as known in the art.
  • Colonies were re-streaked and checked for the integration of the phytase or glucanase expression cassette using PCR with primer pairs specific for the phytase or glucanase expression cassette and the MYCTH:XP_003662453.1 locus as known in the art.
  • a transformant tested positive for the phytase or glucanase expression construct at the MYCTH:XP_003662453.1 locus was selected for further characterization.
  • the phytase production strain UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 was transformed as described in example 1 with the A/otl-digested and isolated nat1 selection marker cassette (s. example 4) from plasmid MT2699 (SEQ ID NO: 17).
  • the transformants were incubated for 3-6 days at 37°C on enriched minimal medium for nat1 selection to select for resistance to nourseothricin (clonNAT) by integrating the nat1 marker as known in the art. Colonies were re-streaked and checked for the integration of the nat1 selection marker and the deletion of the cbh1 locus using PCR with primer pairs specific for nat1 and the cbh1 locus as known in the art.
  • the phytase activity is determined in microtiter plates.
  • the phytase containing supernatant is diluted in reaction buffer (250 mM Na-acetate, 1 mM CaCh, 0.01 % Tween 20, pH 5.5).
  • 10 pl of the enzyme solution are incubated with 140 pl substrate solution (6 mM Na-phytate (Sigma P3168) in reaction buffer) for 1 h at 37°C.
  • the reaction is quenched by adding 150 pl of trichloroacetic acid solution (15% w/w).
  • 20 pl of the quenched reaction solution are treated with 280 pl of freshly made-up color reagent (60 mM L-ascorbic acid (Sigma A7506), 2.2 mM ammonium molybdate tetrahydrate, 325 mM H2SO4), and incubated for 20 min at 37°C, and the absorption at 820 nm is subsequently determined.
  • the substrate buffer on its own is incubated at 37°C and the 10 pl of enzyme sample are only added after quenching with trichloroacetic acid.
  • the color reaction is performed analogously to the remaining measurements.
  • the amount of liberated phosphate is determined via a calibration curve of the color reaction with a phosphate solution of known concentration.
  • the glucanase activity is determined in microtiter plates using the DNS method.
  • the glucanase containing supernatant is diluted in reaction buffer (100 mM Na-acetate, 0.005 % Tween 20, pH 4.5).
  • reaction buffer 100 mM Na-acetate, 0.005 % Tween 20, pH 4.5.
  • CMC carboxymethyl cellulose
  • 35 pl of the diluted enzyme solution is mixed with 35 pL of the substrate solution 30min at 40°C.
  • the reaction is quenched by the addition of 105 pl of DNS solution (16 g/L NaOH, 10 g/L dinitrosalicylic acid, 300 g/L Na-K-tartrate tetrahydrate).
  • T. thermophilus strains grown on agar were inoculated in 1 ml cultivation medium as shown in Table 1 in a 96-deepwell microtiter plate.
  • the strains were fermented at 37°C on a microtiter plate shaker at 900 rpm and 80% humidity for 72 hours.
  • 300 pl of the 72-hour- preculture were transferred in 700 pl cultivation medium as shown in Table 2 in a 96-deepwell microtiter plate.
  • the strains were fermented at 37°C on a microtiter plate shaker at 900 rpm and 80% humidity for 96 hours.
  • production of phytase upon integration of the expression cassette at cbh1 locus is independent of the presence of MYCTH:XP_003662453.1 locus or the replace- ment of MYCTH:XP_003662453.1 locus by nat1 selection marker cassette; analogously, production of phytase upon integration of the expression cassette at MYCTH:XP_003662453.1 locus is independent of the presence of cbh1 locus or the replacement of cbh1 locus by nat1 selection marker cassette. As it can be seen in Figure 4, production of phytase upon integration of the expression cassette at cbh1 locus is higher when the T.
  • thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6 was used compared to usage of the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70; analogously, production of phytase upon integration of the expression cassette at MYCTH:XP_003662453.1 locus is higher when the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6 was used compared to usage of the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70.

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Organic Chemistry (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Engineering & Computer Science (AREA)
  • Biotechnology (AREA)
  • Biomedical Technology (AREA)
  • Molecular Biology (AREA)
  • Microbiology (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Plant Pathology (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Mycology (AREA)
  • Medicinal Chemistry (AREA)
  • Micro-Organisms Or Cultivation Processes Thereof (AREA)

Abstract

The present invention relates to genomic insertion sites allowing for strong expression in filamentous fungi.

Description

New genomic integration site for fungal protein and fine chemicals production
Description of the Invention
The present invention relates to genomic insertion sites allowing for strong expression in filamentous fungi.
Introduction
Biocatalytical and fermentative production of recombinant proteins and fine chemicals has been applied in industrial scale for a long time. In recent years, demand for products derived from biological processes is increasing. Various organisms have been used in such production including bacteria and fungi. Especially filamentous fungi, although difficult to transform and modify genetically have been used more and more as they have been proven to be highly efficient in biocatalytical and fermentative production.
Filamentous fungi have been shown to be excellent hosts for the production of a variety of proteins. Fungal strains such as Aspergillus, Trichoderma, Penicillium and Thermothelomyces have been applied in the industrial production of a wide range of enzymes, since they can secrete large amounts of protein into the fermentation broth. The protein-secreting capacity of these fungi makes them preferred hosts for the targeted production of specific enzymes or enzyme mixtures.
Wild type Thermothelomyces thermophilus (T. thermophilus') C1 (also known as Myceliophthora thermophila, previously described as Chrysosporium lucknowense) is a thermotolerant ascomy- cetous filamentous fungus which has been described as being attractive for production of cellulases and other proteins on a commercial scale. Only recently, the strain has also been shown to be highly efficient for producing various fine chemicals, e.g. in WO20/161682 it is disclosed that T. thermophilus C1 is capable of producing cannabinoids and precursors thereof.
WO00/20555 and US2012/0005812 disclose transformation systems for T. thermophilus C1 and describe expressing and secreting heterologous proteins or polypeptides. Also disclosed is a process for producing large amounts of polypeptides or proteins in an economical manner. WO15/004241 discloses multiple proteases deficient filamentous fungal cells and methods useful for the production of heterologous proteins.
For the expression of recombinant proteins in filamentous fungi, especially in T. thermophilus 01 , a few promoters conferring strong expression of recombinant proteins are known in the art (e.g. Visser et al 2011, Industrial Biotechnology 7 (3), US2014/127788).
However, strength and stability of expression are not only depending on the regulatory elements functionally linked to the recombinant protein to be expressed. The position of the recombinant construct in the genome is also determining both strength and stability of expression, a phenomenon called position effect (Chen and Zhang (2016), Cell Systems 2; Bilyk et al (2017), Microb Cell Fact 16(5)). The number of genomic regions beneficial for the strong and/or stable expression known in the art in T. thermophilus C1 is limited and there is a need to provide fur- ther genomic loci for integration of recombinant expression constructs. It is an object of the present invention to identify such beneficial genomic loci having a positive position effect on recombinant expression.
Detailed description of the Invention
A first embodiment of the invention is an expression system comprising a. a filamentous fungi host cell comprising a native genomic region comprising a sequence selected from the list consisting of i) a nucleic acid sequence comprising SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 44, and iii) a fragment of at least 500, preferably at least 750, preferably at least
1000, more preferably at least 1250, more preferably at least 1500, even more preferably at least 1750, even more preferably at least 2000, even more preferably at least 2250, even more preferably at least 2500, even more preferably at least 2750, even more preferably at least 3000, even more preferably at least 3250 consecutive nucleic acids of a nucleic acid sequence of i) or ii) and b. a recombinant integrating nucleic acid molecule comprising a nucleic acid molecule flanked by distinct homology arms each at least 70% identical, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% most preferably identical over the whole length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 1000, even more preferably at least 1050, even more preferably at least 1100, even more preferably at least 1150, even more preferably at least 1200, even more preferably at least 1250, even more preferably at least 1300, even more preferably at least 1350, even more preferably at least 1400, even more preferably at least 1450, even more preferably at least 1500 contiguous bases of said native genomic region.
The nucleic acid molecule flanked by said homology arms may for example comprise a sequence encoding a protein or an RNA molecule, landing pad or a recombinant expression construct. Said landing pad may comprise a Cre/lox site, one or more sequences recognized by guide RNAs having reduced off-target effects guiding a CRISPR/Cas molecule to this site, one or more recognition sites for a homing endonuclease and the like (Bourgeois et al (2018) ACS Synth Biol 7(11)). Said recombinant expression construct may comprise a promoter, a sequence encoding a protein or an RNA molecule and optionally a terminator.
A further embodiment of the invention is the expression system as defined above wherein said native genomic region is comprising an ORF encoding a small secreted protein flanked at one side by at least 250 nucleotides non-coding DNA and at the other side by at least 250 nucleotides non-coding DNA. Preferably, said ORF is flanked downstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, preferably at least 1100, more preferably at least 1200 nucleotides non-coding DNA. In a most preferred embodiment said ORF is downstream flanked by at least 1296 nucleotides non-coding DNA. Preferably, said ORF is flanked upstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, preferably at least 1700, more preferably at least 1800 nucleotides non-coding DNA. In a most preferred embodiment said ORF is upstream flanked by at least 1847 nucleotides non-coding DNA.
An additional embodiment of the invention is any of the expression systems as defined before wherein said small secreted protein is comprising a sequence selected from the list of i. an amino acid sequence comprising SEQ ID NO: 46 ii. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 46, iii. a fragment of at least 100, preferably at least 110, more preferably at least 120, even more preferably at least 130, most preferably at least 140 consecutive amino acids of i) or ii).
Preferably said ORF is encoding the protein MYCTH_2314963 (GeneBank Gene ID 11512560 updated November 28, 2019.
Another embodiment of the invention is any of the expression systems as defined above wherein the native genomic region is flanked at one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 41 b. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 41, and c. a fragment of at least 100, preferably at least 150, preferably at least
200, preferably at least 250, more preferably at least 300, even more preferably at least 320, most preferably at least 340 consecutive amino acids of a) or b), and at the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of d. an amino acid sequence comprising SEQ ID NO: 43 e. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 43, and f. a fragment of at least 100, preferably at least 150, preferably at least
200, preferably at least 250, more preferably at least 300, even more preferably at least 325, even more preferably 350, most preferably at least 370 consecutive amino acids of d) or e).
In the context of this invention, the term “flanked” means that the respective re- gions/ORFs/genes “flanking” a certain genomic region are defining the borders of said genomic region. The flanking regions are not necessarily directly adjacent to a sequence used to define said genomic region. For example, the genomic region may be defined by a coding sequence, located within said genomic region. However, there may be spacers between such coding region and the respective flanking regions defining the respective genomic region.
A further embodiment of the invention is any of the expression systems as defined above wherein the filamentous fungi host cell is from a species of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysospori- um, Filibasidium, Fusarium, Humicola, Magnaporthe, Monascus, Mucor, Myceliophthora, Mor- tierella, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Piromyces, Phanerochaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordaria, Talaromyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypocladium, Trametes and Trichoderma, more preferred Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rasamsonia emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Ther- mothelomyces heterothallica and Thermothelomyces thermophilus (formerly known as Myceli- ophthora thermophila and as Chrysosporium lucknowense), most preferably Thermothelomyces therm ophilus.
A further embodiment of the invention is a method for stable integration of a recombinant polynucleotide in a filamentous fungi host cell comprising the steps of c. providing a filamentous fungi host cell comprising a native genomic region comprising a sequence selected from the list consisting of i) a nucleic acid sequence comprising SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99%to SEQ ID NO 44, and iii) a fragment of at least 500, preferably at least 750, preferably at least
1000, more preferably at least 1250, more preferably at least 1500, even more preferably at least 1750, even more preferably at least 2000, even more preferably at least 2250, even more preferably at least 2500, even more preferably at least 2750, even more preferably at least 3000, even more preferably at least 3250 consecutive nucleic acids of a nucleic acid sequence of i) or ii), d. introducing into said host cell a recombinant integrating nucleic acid molecule comprising a nucleic acid molecule flanked by distinct homology arms each at least 70% identical, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% most preferably identical over the whole length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 1000, even more preferably at least 1050, even more preferably at least 1100, even more preferably at least 1150, even more preferably at least 1200, even more preferably at least 1250, even more preferably at least 1300, even more preferably at least 1350, even more preferably at least 1400, even more prefera- bly at least 1450, even more preferably at least 1500 contiguous bases of said native genomic region, e. allowing for homologous recombination of said recombinant nucleic acid molecule into said native genomic region.
In the event, the recombinant nucleic acid molecule comprises a coding sequence, for example comprised in a recombinant expression construct it is preferred embodiment that the host cell having integrated said recombinant nucleic acid molecule produces the recombinant peptide, protein and/or nucleic acid of interest encoded by said coding region.
In a further preferred embodiment, said native genomic region before recombination said recombinant nucleic acid molecule comprises a native gene encoding a small secreted protein flanked upstream by 250 kb non-coding DNA and downstream by 250 kb non-coding DNA. Preferably, the ORF of said small secreted protein is flanked downstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, preferably at least 1100, more preferably at least 1200 nucleotides non-coding DNA. In a most preferred embodiment said ORF is downstream flanked by at least 1296 nucleotides non-coding DNA. Preferably, said ORF is flanked upstream by at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1100, at least 1200, at least 1300, at least 1400, at least 1500, at least 1600, preferably at least 1700, more preferably at least 1800 nucleotides non-coding DNA. In a most preferred embodiment said ORF is upstream flanked by at least 1847 nucleotides non-coding DNA.
In a further preferred embodiment, the landing pad or recombinant expression construct introduced into the genome of the fungi host cell is replacing in part or completely the gene encoding the small secreted protein comprised in said native genomic region. More preferably it is replacing all or part of the ORF encoding the small secreted protein.
A further embodiment of the invention is the method as defined before wherein said small secreted protein is comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 46 b. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99%to SEQ ID NO: 46, and c. a fragment of at least 100, preferably at least 110, more preferably at least 120, even more preferably at least 130, most preferably at least 140 consecutive amino acids consecutive amino acids of a) or b). A further embodiment of the invention is any of the methods as defined above wherein said native genomic region is flanked at one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 41 b. an amino acid sequence with a homology of at least 50% preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 41, and c. a fragment of at least 100, preferably at least 150, preferably at least
200, preferably at least 250, more preferably at least 300, even more preferably at least 320, most preferably at least 340 consecutive amino acids of a) or b), and at the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of d. an amino acid sequence comprising SEQ ID NO: 43 e. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 43, and f. a fragment of at least 100, preferably at least 150, preferably at least
200, preferably at least 250, more preferably at least 300, even more preferably at least 325, even more preferably 350, most preferably at least 370 consecutive amino acids of d) or e).
A further embodiment of the invention is any of the methods as defined above wherein the filamentous fungi host cell is from a species selected from the group consisting of of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysosporium, Filibasidium, Fusarium, Humicola, Magnaporthe, Monascus, Mu- cor, Myceliophthora, Mortierella, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Piro- myces, Phanerochaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordaria, Tala- romyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypocladium, Trametes and Trichoderma, more preferred Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rasamsonia emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermothelomyces heterothallica and Thermothelomyces thermophilus (former- ly known as Myceliophthora thermophila and as Chrysosporium lucknowense), most preferably Thermothelomyces thermophilus.
A further embodiment of the invention is a recombinant filamentous fungi host cell comprising a recombinant nucleic acid molecule flanked by genomic DNA comprising on one side of said recombinant nucleic acid molecule a sequence selected from the list consisting of i) a nucleic acid sequence comprising 1296 consecutive bases from the 5- prime end of SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least
97%, at least 98% or at least 99% to i), and iii) a fragment of at least 100, preferably at least 200, preferably at least
300, preferably at least 400, preferably at least 500, preferably at least
600, preferably at least 700, preferably at least 800, preferably at least
900, more preferably at least 1000, more preferably at least 1100, even more preferably at least 1200, most preferably at least 1296 consecutive nucleic acids of i) or ii). and comprising on the other side of said recombinant nucleic acid molecule a sequence selected from the list consisting of iv) a nucleic acid sequence comprising 1847 consecutive bases from the 3- prime end of SEQ ID NO: 44 v) a nucleic acid sequence with an identity of at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least
97%, at least 98% or at least 99% to iv), and vi) a fragment of at least 100, preferably at least 200, preferably at least
300, preferably at least 400, preferably at least 500, preferably at least
600, preferably at least 700, preferably at least 800, preferably at least
900, preferably at least 1000, preferably at least 1100, more preferably at least 1200, preferably at least 1300, preferably at least 1400, preferably at least 1500, more preferably at least 1600, more preferably at least 1700, even more preferably at least 1800, most preferably at least 1847 consecutive nucleic acids of iv) or v).
A further embodiment of the invention is a recombinant filamentous fungi host cell comprising a recombinant nucleic acid molecule flanked at one side by the XP_003662454.1 gene and at the other side by the XP_003662452.1 gene, preferably wherein the XP_003662454.1 gene is encoding an amino acid molecule selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 43, b. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 43, and c. a fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 325, even more preferably 350, most preferably at least 370 consecutive amino acids consecutive amino acids of a) or b), and the XP_003662452.1 gene is encoding an amino acid molecule selected from the list consisting of d. an amino acid sequence comprising SEQ ID NO: 41, e. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 41, and f. a fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 320, most preferably at least 340 consecutive amino acids of d) or e).
A further embodiment of the invention is the recombinant filamentous fungi host cell as defined above wherein the host cell is from a species selected from the group consisting of of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysosporium, Filibasidium, Fusarium, Humicola, Magnaporthe, Monascus, Mu- cor, Myceliophthora, Mortierella, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Piro- myces, Phanerochaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordaria, Tala- romyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypocladium, Trametes and Trichoderma, more preferred Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rasamsonia emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermothelomyces heterothallica and Thermothelomyces thermophilus (formerly known as Myceliophthora thermophila and as Chrysosporium lucknowense), most preferably Thermothelomyces thermophilus..
An additional embodiment of the invention is a recombinant nucleic acid molecule flanked by distinct homology arms each at least 70% identical, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% most preferably identical over the whole length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 1000, even more preferably at least 1050, even more preferably at least 1100, even more preferably at least 1150, even more preferably at least 1200, even more preferably at least 1250, even more preferably at least 1300, even more preferably at least 1350, even more preferably at least 1400, even more preferably at least 1450, even more preferably at least 1500 contiguous bases of a native genomic region of a filamentous fungus comprising a sequence selected from the list consisting of i) a nucleic acid comprising SEQ ID NO: 44, ii)a nucleic acid sequence with an identity of at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to SEQ ID NO: 44, and iii) a fragment of at least 500, preferably at least 750, preferably at least
1000, more preferably at least 1250, more preferably at least 1500, even more preferably at least 1750, even more preferably at least 2000, even more preferably at least 2250, even more preferably at least 2500, even more preferably at least 2750, even more preferably at least 3000, even more preferably at least 3250 consecutive nucleic acids of i) or ii).
A further embodiment is a recombinant nucleic acid molecule flanked by distinct homology arms each at least 70% identical, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% most preferably identical over the whole length to at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, more preferably at least 350, even more preferably at least 400, even more preferably at least 450, even more preferably at least 500, even more preferably at least 550, even more preferably at least 600, even more preferably at least 650, even more preferably at least 700, even more preferably at least 750, even more preferably at least 800, even more preferably at least 850, even more preferably at least 900, even more preferably at least 950, even more preferably at least 1000, even more preferably at least 1050, even more preferably at least 1100, even more preferably at least 1150, even more preferably at least 1200, even more preferably at least 1250, even more preferably at least 1300, even more preferably at least 1350, even more preferably at least 1400, even more preferably at least 1450, even more preferably at least 1500 contiguous bases of a native genomic region of a filamentous fungus which is flanked at one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of
1. an amino acid sequence comprising SEQ ID NO: 41,
2. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to to SEQ ID NO: 41, and
3. a fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 320, most preferably at least 340 consecutive amino acids of 1) or 2), and at the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of
4. an amino acid sequence comprising SEQ ID NO: 43,
5. an amino acid sequence with a homology of at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, more preferably at least 85%, more preferably at least 90%, more preferably at least 95%, even more preferably at least 96%, at least 97%, at least 98% or at least 99% to to SEQ ID NO: 43, and
6. a fragment of at least 100, preferably at least 150, preferably at least 200, preferably at least 250, more preferably at least 300, even more preferably at least 325, even more preferably 350, most preferably at least 370 consecutive amino acids of 4) or 5). A further embodiment of the invention is a vector comprising any of the recombinant nucleic acid molecules as defined above.
DEFINITIONS
It is to be understood that this invention is not limited to the particular methodology or protocols. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the present invention which will be limited only by the appended claims. It must be noted that as used herein and in the appended claims, the singular forms "a," "and," and "the" include plural reference unless the context clearly dictates otherwise. Thus, for example, reference to "a vector" is a reference to one or more vectors and includes equivalents thereof known to those skilled in the art, and so forth. The term "about" is used herein to mean approximately, roughly, around, or in the region of. When the term "about" is used in conjunction with a numerical range, it modifies that range by extending the boundaries above and below the numerical values set forth. In general, the term "about" is used herein to modify a numerical value above and below the stated value by a variance of 20 percent, preferably 10 percent up or down (higher or lower). As used herein, the word "or" means any one member of a particular list and also includes any combination of members of that list. The words "comprise," "comprising," "include," "including," and "includes" when used in this specification and in the following claims are intended to specify the presence of one or more stated features, integers, components, or steps, but they do not preclude the presence or addition of one or more other features, integers, components, steps, or groups thereof. For clarity, certain terms used in the specification are defined and used as follows:
Coding region: As used herein the terms "coding region" or “open reading frame’’ or “ORF” when used in reference to a structural gene refers to the nucleotide sequences which encode the amino acids found in the nascent polypeptide as a result of translation of a mRNA molecule. The coding region is bounded, in eukaryotes, on the 5'-side by the nucleotide triplet "ATG" which encodes the initiator methionine and on the 3'-side by one of the three triplets which specify stop codons (i.e., TAA, TAG, TGA). In addition to containing introns, genomic forms of a gene may also include sequences located on both the 5'- and 3'-end of the sequences which are present on the RNA transcript. These sequences are referred to as "flanking" sequences or regions (these flanking sequences are located 5' or 3' to the non-translated sequences present on the mRNA transcript). The 5'-flanking region may contain regulatory sequences such as promoters and enhancers which control or influence the transcription of the gene. The 3'- flanking region may contain sequences which direct the termination of transcription, post- transcriptional cleavage and polyadenylation. Complementary: "Complementary" or "complementarity" refers to two nucleotide sequences which comprise antiparallel nucleotide sequences capable of pairing with one another (by the base-pairing rules) upon formation of hydrogen bonds between the complementary base residues in the antiparallel nucleotide sequences. For example, the sequence 5'-AGT-3' is complementary to the sequence 5'-ACT-3'. Complementarity can be "partial" or "total." "Partial" complementarity is where one or more nucleic acid bases are not matched according to the base pairing rules. "Total" or "complete" complementarity between nucleic acid molecules is where each and every nucleic acid base is matched with another base under the base pairing rules. The degree of complementarity between nucleic acid molecule strands has significant effects on the efficiency and strength of hybridization between nucleic acid molecule strands. A "complement" of a nucleic acid sequence as used herein refers to a nucleotide sequence whose nucleic acid molecules show total complementarity to the nucleic acid molecules of the nucleic acid sequence.
Endogenous: An "endogenous" nucleotide sequence refers to a nucleotide sequence, which is present in the genome of the untransformed cell.
Enhanced expression: “enhance” or “increase” the expression of a nucleic acid molecule in a cell are used equivalently herein and mean that the level of expression of the nucleic acid molecule in a cell after applying a method of the present invention is higher than its expression in the cell before applying the method, or compared to a reference cell lacking a recombinant nucleic acid molecule of the invention. The term "enhanced” or “increased" as used herein are synonymous and means herein higher, preferably significantly higher expression of the nucleic acid molecule to be expressed. As used herein, an “enhancement” or “increase” of the level of an agent such as a protein, mRNA or RNA means that the level is increased relative to a substantially identical cell grown under substantially identical conditions, lacking a recombinant nucleic acid molecule of the invention, the recombinant construct or recombinant vector of the invention. As used herein, “enhancement” or “increase” of the level of an agent, such as for example a preRNA, mRNA, rRNA, tRNA, snoRNA, snRNA expressed by the target gene and/or of the protein product encoded by it, means that the level is increased 50% or more, for example 100% or more, preferably 200% or more, more preferably 5 fold or more, even more preferably 10 fold or more, most preferably 20 fold or more for example 50 fold relative to a cell or organism lacking a recombinant nucleic acid molecule of the invention. The enhancement or increase can be determined by methods with which the skilled worker is familiar. Thus, the enhancement or increase of the nucleic acid or protein quantity can be determined for example by an immunological detection of the protein. Moreover, techniques such as protein assay, fluorescence, Northern hybridization, nuclease protection assay, reverse transcription (quantitative RT-PCR), ELISA (enzyme-linked immunosorbent assay), Western blotting, radioimmunoassay (RIA) or other immunoassays and fluorescence-activated cell analysis (FACS) can be employed to measure a specific protein or RNA in a cell. Depending on the type of the induced protein product, its activity or the effect on the phenotype of the cell may also be determined. Methods for determining the protein quantity are known to the skilled worker. Examples, which may be mentioned, are: the micro-Biuret method (Goa J (1953) Scand J Clin Lab Invest 5:218-222), the Fo- lin-Ciocalteau method (Lowry OH et al. (1951) J Biol Chem 193:265-275) or measuring the absorption of CBB G-250 (Bradford MM (1976) Analyt Biochem 72:248-254).
Expression: "Expression" refers to the biosynthesis of a gene product, preferably to the transcription and/or translation of a nucleotide sequence, for example an endogenous gene or a heterologous gene, in a cell. For example, in the case of a structural gene, expression involves transcription of the structural gene into mRNA and - optionally - the subsequent translation of mRNA into one or more polypeptides. In other cases, expression may refer only to the transcription of the DNA harboring an RNA molecule.
Expression construct: "Expression construct" as used herein mean a DNA sequence capable of directing expression of a particular nucleotide sequence in a cell, comprising a promoter functional in said cell into which it will be introduced, operatively linked to the nucleotide sequence of interest which is - optionally - operatively linked to termination signals. If translation is required, it also typically comprises sequences required for proper translation of the nucleotide sequence. The coding region may code for a protein of interest but may also code for a functional RNA of interest, for example RNAa, siRNA, snoRNA, snRNA, microRNA, ta-siRNA or any other noncoding regulatory RNA, in the sense or antisense direction. The expression construct comprising the nucleotide sequence of interest may be chimeric, meaning that one or more of its components is heterologous with respect to one or more of its other components. The expression construct may also be one, which is naturally occurring but has been obtained in a recombinant form useful for heterologous expression. Typically, however, the expression construct is heterologous with respect to the host, i.e., the particular DNA sequence of the expression construct does not occur naturally in the host cell and must have been introduced into the host cell or an ancestor of the host cell by a transformation event. The expression of the nucleotide sequence in the expression construct may be under the control of a constitutive promoter or of an inducible promoter, which initiates transcription only when the host cell is exposed to some particular external stimulus. The promoter can also be specific to a particular stage of development.
Foreign: The term "foreign" refers to any nucleic acid molecule (e.g., gene sequence) which is introduced into the genome of a cell by experimental manipulations and may include sequences found in that cell so long as the introduced sequence contains some modification (e.g., a point mutation, the presence of a selectable marker gene, etc.) and is therefore distinct relative to the naturally-occurring sequence.
Functional linkage: The term "functional linkage" or "functionally linked" are to be understood as meaning, for example, the sequential arrangement of a regulatory element (e.g. a promoter) with a nucleic acid sequence to be expressed and, if appropriate, further regulatory elements (such as e.g., a terminator) in such a way that each of the regulatory elements can fulfil its intended function to allow, modify, facilitate or otherwise influence expression of said nucleic acid sequence. As a synonym the wording “operable linkage” or “operably linked” may be used. The expression may result depending on the arrangement of the nucleic acid sequences in relation to sense or antisense RNA. To this end, direct linkage in the chemical sense is not necessarily required. Genetic control sequences such as, for example, enhancer sequences, can also exert their function on the target sequence from positions which are further away, or indeed from other DNA molecules. Preferred arrangements are those in which the nucleic acid sequence to be expressed recombinantly is positioned behind the sequence acting as promoter, so that the two sequences are linked covalently to each other. The distance between the promoter sequence and the nucleic acid sequence to be expressed recombinantly is preferably less than 200 base pairs, especially preferably less than 100 base pairs, very especially preferably less than 50 base pairs. In a preferred embodiment, the nucleic acid sequence to be transcribed is located behind the promoter in such a way that the transcription start is identical with the desired beginning of the chimeric RNA of the invention. Functional linkage, and an expression construct, can be generated by means of customary recombination and cloning techniques as described (e.g., in Maniatis T, Fritsch EF and Sambrook J (1989) Molecular Cloning: A Laboratory Manual, 2nd Ed., Cold Spring Harbor Laboratory, Cold Spring Harbor (NY); Silhavy et al. (1984) Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor (NY); Au- subel et al. (1987) Current Protocols in Molecular Biology, Greene Publishing Assoc, and Wiley Interscience). However, further sequences, which, for example, act as a linker with specific cleavage sites for restriction enzymes, or as a signal peptide, may also be positioned between the two sequences. The insertion of sequences may also lead to the expression of fusion proteins. Preferably, the expression construct, consisting of a linkage of a regulatory region for example a promoter and nucleic acid sequence to be expressed, can exist in a vector-integrated form and be inserted into the genome, for example by transformation.
Gene: The term "gene" refers to a region operably joined to appropriate regulatory sequences capable of regulating the expression of the gene product (e.g., a polypeptide or a functional RNA) in some manner. A gene includes untranslated regulatory regions of DNA (e.g., promoters, enhancers, repressors, etc.) preceding (up-stream) and following (downstream) the coding region (open reading frame, ORF) as well as, where applicable, intervening sequences (i.e., introns) between individual coding regions (i.e., exons). The term "structural gene" as used herein is intended to mean a DNA sequence that is transcribed into mRNA which is then translated into a sequence of amino acids characteristic of a specific polypeptide.
Genome and genomic DNA: The terms “genome” or “genomic DNA” is referring to the heritable genetic information of a host organism. Preferably the terms genome or genomic DNA is referring to the chromosomal DNA of the nucleus.
Heterologous: The term "heterologous” with respect to a nucleic acid molecule or DNA refers to a nucleic acid molecule which is operably linked to, or is manipulated to become operably linked to, a second nucleic acid molecule, e.g. a promoter to which it is not operably linked in nature, e.g. in the genome of a WT cell, or to which it is operably linked at a different location or position in nature, e.g. in the genome of a WT cell.
Preferably the term "heterologous” with respect to a nucleic acid molecule or DNA, e.g. promoter refers to a nucleic acid molecule which is operably linked to, or is manipulated to become operably linked to, a second nucleic acid molecule, e.g. a ORF to which it is not operably linked in nature.
A heterologous expression construct comprising a nucleic acid molecule and one or more regulatory nucleic acid molecule (such as a promoter or a transcription termination signal) linked thereto for example is a constructs originating by experimental manipulations in which either a) said nucleic acid molecule, or b) said regulatory nucleic acid molecule or c) both (i.e. (a) and (b)) is not located in its natural (native) genetic environment or has been modified by experimental manipulations, an example of a modification being a substitution, addition, deletion, inversion or insertion of one or more nucleotide residues. Natural genetic environment refers to the natural chromosomal locus in the organism of origin, or to the presence in a genomic library. In the case of a genomic library, the natural genetic environment of the sequence of the nucleic acid molecule is preferably retained, at least in part. The environment flanks the nucleic acid sequence at least at one side and has a sequence of at least 50 bp, preferably at least 500 bp, especially preferably at least 1 ,000 bp, very especially preferably at least 5,000 bp, in length. A naturally occurring expression construct - for example the naturally occurring combination of a promoter with the corresponding gene - becomes a transgenic expression construct when it is modified by non-natural, synthetic “artificial” methods such as, for example, mutagenization. Such methods have been described (US 5,565,350; WO 00/15815). For example, a protein encoding nucleic acid molecule operably linked to a promoter, which is not the native promoter of this molecule, is considered to be heterologous with respect to the promoter. Preferably, heterologous DNA is not endogenous to or not naturally associated with the cell into which it is introduced, but has been obtained from another cell or has been synthesized. Heterologous DNA also includes an endogenous DNA sequence, which contains some modification, non-naturally occurring, multiple copies of an endogenous DNA sequence, or a DNA sequence which is not naturally associated with another DNA sequence physically linked thereto.
Hybridization: The term "hybridization" as defined herein is a process wherein substantially complementary nucleotide sequences anneal to each other. The hybridisation process can occur entirely in solution, i.e. both complementary nucleic acids are in solution. The hybridisation process can also occur with one of the complementary nucleic acids immobilised to a matrix such as magnetic beads, Sepharose beads or any other resin. The hybridisation process can furthermore occur with one of the complementary nucleic acids immobilised to a solid support such as a nitro-cellulose or nylon membrane or immobilised by e.g. photolithography to, for example, a siliceous glass support (the latter known as nucleic acid arrays or microarrays or as nucleic acid chips). In order to allow hybridisation to occur, the nucleic acid molecules are generally thermally or chemically denatured to melt a double strand into two single strands and/or to remove hairpins or other secondary structures from single stranded nucleic acids.
The term “stringency” refers to the conditions under which a hybridisation takes place. The stringency of hybridisation is influenced by conditions such as temperature, salt concentration, ionic strength and hybridisation buffer composition. Generally, low stringency conditions are selected to be about 30°C lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH. Medium stringency conditions are when the temperature is 20°C below Tm, and high stringency conditions are when the temperature is 10°C below Tm. High stringency hybridisation conditions are typically used for isolating hybridising sequences that have high sequence similarity to the target nucleic acid sequence. However, nucleic acids may deviate in sequence and still encode a substantially identical polypeptide, due to the degeneracy of the genetic code. Therefore, medium stringency hybridisation conditions may sometimes be needed to identify such nucleic acid molecules.
The “Tm” is the temperature under defined ionic strength and pH, at which 50% of the target sequence hybridises to a perfectly matched probe. The Tm is dependent upon the solution conditions and the base composition and length of the probe. For example, longer sequences hybridise specifically at higher temperatures. The maximum rate of hybridisation is obtained from about 16°C up to 32°C below Tm. The presence of monovalent cations in the hybridisation solution reduce the electrostatic repulsion between the two nucleic acid strands thereby promoting hybrid formation; this effect is visible for sodium concentrations of up to 0.4M (for higher concentrations, this effect may be ignored). Formamide reduces the melting temperature of DNA- DNA and DNA-RNA duplexes with 0.6 to 0.7°C for each percent formamide, and addition of 50% formamide allows hybridisation to be performed at 30 to 45°C, though the rate of hybridisation will be lowered. Base pair mismatches reduce the hybridisation rate and the thermal stability of the duplexes. On average and for large probes, the Tm decreases about 1°C per % base mismatch. The Tm may be calculated using the following equations, depending on the types of hybrids:
DNA-DNA hybrids (Meinkoth and Wahl, Anal. Biochem., 138: 267-284, 1984):
Tm= 81.5°C + 16.6xlog[Na+]a + 0.41x%[G/Cb] - 500x[Lc]-1 - 0.61x% formamide DNA-RNA or RNA-RNA hybrids:
Tm= 79.8 + 18.5 (log10[Na+]a) + 0.58 (%G/Cb) + 11.8 (%G/Cb)2 - 820/Lc oligo-DNA or oligo-RNAd hybrids: For <20 nucleotides: Tm= 2 (In)
For 20-35 nucleotides: Tm= 22 + 1.46 (In ) a or for other monovalent cation, but only accurate in the 0.01-0.4 M range. b only accurate for %GC in the 30% to 75% range. c L = length of duplex in base pairs. d Oligo, oligonucleotide; In, effective length of primer = 2*(no. of G/C)+(no. of A/T). Non-specific binding may be controlled using any one of a number of known techniques such as, for example, blocking the membrane with protein containing solutions, additions of heterologous RNA, DNA, and SDS to the hybridisation buffer, and treatment with Rnase. For nonrelated probes, a series of hybridizations may be performed by varying one of (i) progressively lowering the annealing temperature (for example from 68°C to 42°C) or (ii) progressively lowering the formamide concentration (for example from 50% to 0%). The skilled artisan is aware of various parameters which may be altered during hybridisation and which will either maintain or change the stringency conditions.
Besides the hybridisation conditions, specificity of hybridisation typically also depends on the function of post-hybridisation washes. To remove background resulting from non-specific hybridisation, samples are washed with dilute salt solutions. Critical factors of such washes include the ionic strength and temperature of the final wash solution: the lower the salt concentration and the higher the wash temperature, the higher the stringency of the wash. Wash conditions are typically performed at or below hybridisation stringency. A positive hybridisation gives a signal that is at least twice of that of the background. Generally, suitable stringent conditions for nucleic acid hybridisation assays or gene amplification detection procedures are as set forth above. More or less stringent conditions may also be selected. The skilled artisan is aware of various parameters which may be altered during washing and which will either maintain or change the stringency conditions.
For example, typical high stringency hybridisation conditions for DNA hybrids longer than 50 nucleotides encompass hybridisation at 65°C in 1x SSC or at 42°C in 1x SSC and 50% formamide, followed by washing at 65°C in 0.3x SSC. Examples of medium stringency hybridisation conditions for DNA hybrids longer than 50 nucleotides encompass hybridisation at 50°C in 4x SSC or at 40°C in 6x SSC and 50% formamide, followed by washing at 50°C in 2x SSC. The length of the hybrid is the anticipated length for the hybridising nucleic acid. When nucleic acids of known sequence are hybridised, the hybrid length may be determined by aligning the sequences and identifying the conserved regions described herein. 1 xSSC is 0.15M NaCI and 15mM sodium citrate; the hybridisation solution and wash solutions may additionally include 5x Denhardt's reagent, 0.5-1.0% SDS, 100 pg/ml denatured, fragmented salmon sperm DNA, 0.5% sodium pyrophosphate. Another example of high stringency conditions is hybridisation at 65°C in 0.1 x SSC comprising 0.1 SDS and optionally 5x Denhardt's reagent, 100 pg/ml denatured, fragmented salmon sperm DNA, 0.5% sodium pyrophosphate, followed by the washing at 65°C in 0.3x SSC.
For the purposes of defining the level of stringency, reference can be made to Sambrook et al. (2001) Molecular Cloning: a laboratory manual, 3rd Edition, Cold Spring Harbor Laboratory Press, CSH, New York or to Current Protocols in Molecular Biology, John Wiley & Sons, N.Y. (1989 and yearly updates).
"Identity”: “Identity” when used in respect to the comparison of two or more nucleic acid or amino acid molecules means that the sequences of said molecules share a certain degree of sequence similarity, the sequences being partially identical.
Enzyme variants may be defined by their sequence identity when compared to a parent enzyme or nucleic acid molecule. Sequence identity usually is provided as “% sequence identity” or “% identity”. To determine the percent-identity between two amino acid sequences in a first step a pairwise sequence alignment is generated between those two sequences, wherein the two sequences are aligned over their complete length (i.e., a pairwise global alignment). The alignment is generated with a program implementing the Needleman and Wunsch algorithm (J. Mol. Biol. (1979) 48, p. 443-453), preferably by using the program “NEEDLE” (The European Molecular Biology Open Software Suite (EMBOSS)) with the programs default parameters (gapo- pen=10.0, gapextend=0.5 and matrix=EBLOSUM62). The preferred alignment for the purpose of this invention is that alignment, from which the highest sequence identity can be determined. The following example is meant to illustrate two nucleotide sequences, but the same calculations apply to protein sequences: Seq A: AAGATACTG length: 9 bases Seq B: GATCTGA length: 7 bases
Hence, the shorter sequence is sequence B.
Producing a pairwise global alignment which is showing both sequences over their complete lengths results in
Seq A : AAGATACTG-
I I I I I I
Seq B : — GAT-CTGA The “I” symbol in the alignment indicates identical residues (which means bases for DNA or amino acids for proteins). The number of identical residues is 6.
The symbol in the alignment indicates gaps. The number of gaps introduced by alignment within the Seq B is 1 . The number of gaps introduced by alignment at borders of Seq B is 2, and at borders of Seq A is 1.
The alignment length showing the aligned sequences over their complete length is 10.
Producing a pairwise alignment which is showing the shorter sequence over its complete length according to the invention consequently results in:
Seq A : GATACTG-
I I I I I I
Seq B : GAT-CTGA
Producing a pairwise alignment which is showing sequence A over its complete length according to the invention consequently results in:
Seq A : AAGATACTG
I I I I I I
Seq B : — GAT-CTG
Producing a pairwise alignment which is showing sequence B over its complete length according to the invention consequently results in:
Seq A : GATACTG-
I I I I I I
Seq B : GAT-CTGA
The alignment length showing the shorter sequence over its complete length is 8 (one gap is present which is factored in the alignment length of the shorter sequence).
Accordingly, the alignment length showing Seq A over its complete length would be 9 (meaning Seq A is the sequence of the invention).
Accordingly, the alignment length showing Seq B over its complete length would be 8 (meaning Seq B is the sequence of the invention).
After aligning two sequences, in a second step, an identity value is determined from the alignment produced. For purposes of this description, percent identity is calculated by %-identity = (identical residues I length of the alignment region which is showing the respective sequence of this invention over its complete length) *100. Thus, sequence identity in relation to comparison of two amino acid sequences according to this embodiment is calculated by dividing the number of identical residues by the length of the alignment region which is showing the respective sequence of this invention over its complete length. This value is multiplied with 100 to give “%- identity”. According to the example provided above, %-identity is: for Seq A being the sequence of the invention (6 / 9) * 100 = 66.7 %; for Seq B being the sequence of the invention (6 / 8) * 100 =75%.
The term “Introducing”, “introduction” and the like with respect to the introduction of a donor DNA molecule in the target site of a target DNA means any introduction of the sequence of the donor DNA molecule into the target region for example by the physical integration of the donor DNA molecule or a part thereof into the target region or the introduction of the sequence of the donor DNA molecule or a part thereof into the target region wherein the donor DNA is used as template for a polymerase.
Isogenic: organisms (e.g., fungi), which are genetically identical, except that they may differ by the presence or absence of a heterologous DNA sequence.
Isolated: The term "isolated" as used herein means that a material has been removed by the hand of man and exists apart from its original, native environment and is therefore not a product of nature. An isolated material or molecule (such as a DNA molecule or enzyme) may exist in a purified form or may exist in a non-native environment such as, for example, in a transgenic host cell. For example, a naturally occurring polynucleotide or polypeptide present in a living cell is not isolated, but the same polynucleotide or polypeptide, separated from some or all of the coexisting materials in the natural system, is isolated. Such polynucleotides can be part of a vector and/or such polynucleotides or polypeptides could be part of a composition and would be isolated in that such a vector or composition is not part of its original environment. Preferably, the term "isolated" when used in relation to a nucleic acid molecule, as in "an isolated nucleic acid sequence" refers to a nucleic acid sequence that is identified and separated from at least one contaminant nucleic acid molecule with which it is ordinarily associated in its natural source. Isolated nucleic acid molecule is nucleic acid molecule present in a form or setting that is different from that in which it is found in nature. In contrast, non-isolated nucleic acid molecules are nucleic acid molecules such as DNA and RNA, which are found in the state they exist in nature. For example, a given DNA sequence (e.g., a gene) is found on the host cell chromosome in proximity to neighbouring genes; RNA sequences, such as a specific mRNA sequence encoding a specific protein, are found in the cell as a mixture with numerous other mRNAs, which encode a multitude of proteins. However, an isolated nucleic acid sequence comprising for example SEQ ID NO: 1 includes, by way of example, such nucleic acid sequences in cells which ordinarily contain SEQ ID NO:1 where the nucleic acid sequence is in a chromosomal or ex- trachromosomal location different from that of natural cells or is otherwise flanked by a different nucleic acid sequence than that found in nature. The isolated nucleic acid sequence may be present in single-stranded or double-stranded form. When an isolated nucleic acid sequence is to be utilized to express a protein, the nucleic acid sequence will contain at a minimum at least a portion of the sense or coding strand (i.e., the nucleic acid sequence may be single-stranded). Alternatively, it may contain both the sense and anti-sense strands (i.e., the nucleic acid sequence may be double-stranded).
Minimal Promoter: promoter elements, particularly a TATA element, that are inactive or that have greatly reduced promoter activity in the absence of upstream activation. In the presence of a suitable transcription factor, the minimal promoter functions to permit transcription.
Non-coding: The term "non-coding" refers to sequences of nucleic acid molecules that do not encode part or all of an expressed protein. Non-coding sequences include but are not limited to functional RNAs, introns, enhancers, promoter regions, 3' untranslated regions, and 5' untranslated regions.
Nucleic acids and nucleotides: The terms "Nucleic Acids" and "Nucleotides" refer to naturally occurring or synthetic or artificial nucleic acid or nucleotides. The terms “nucleic acids” and "nucleotides” comprise deoxyribonucleotides or ribonucleotides or any nucleotide analogue and polymers or hybrids thereof in either single- or double-stranded, sense or antisense form. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated. The term "nucleic acid" is used interchangeably herein with "gene", "cDNA, "mRNA", "oligonucleotide," and "polynucleotide". Nucleotide analogues include nucleotides having modifications in the chemical structure of the base, sugar and/or phosphate, including, but not limited to, 5-position pyrimidine modifications, 8- position purine modifications, modifications at cytosine exocyclic amines, substitution of 5- bromo-uracil, and the like; and 2'-position sugar modifications, including but not limited to, sug- ar-modified ribonucleotides in which the 2'-OH is replaced by a group selected from H, OR, R, halo, SH, SR, NH2, NHR, NR2, or CN. Short hairpin RNAs (shRNAs) also can comprise nonnatural elements such as non-natural bases, e.g., ionosin and xanthine, non-natural sugars, e.g., 2'-methoxy ribose, or non-natural phosphodiester linkages, e.g., methylphosphonates, phosphorothioates and peptides.
Nucleic acid sequence: The phrase "nucleic acid sequence" refers to a single or doublestranded polymer of deoxyribonucleotide or ribonucleotide bases read from the 5'- to the 3'-end. It includes chromosomal DNA, self-replicating plasmids, infectious polymers of DNA or RNA and DNA or RNA that performs a primarily structural role. "Nucleic acid sequence" also refers to a consecutive list of abbreviations, letters, characters or words, which represent nucleotides. In one embodiment, a nucleic acid can be a "probe" which is a relatively short nucleic acid, usually less than 100 nucleotides in length. Often a nucleic acid probe is from about 50 nucleotides in length to about 10 nucleotides in length. A "target region" of a nucleic acid is a portion of a nucleic acid that is identified to be of interest. A "coding region" of a nucleic acid is the portion of the nucleic acid, which is transcribed and translated in a sequence-specific manner to produce into a particular polypeptide or protein when placed under the control of appropriate regulatory sequences. The coding region is said to encode such a polypeptide or protein.
Oligonucleotide: The term "oligonucleotide" refers to an oligomer or polymer of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA) or mimetics thereof, as well as oligonucleotides having non-naturally-occurring portions which function similarly. Such modified or substituted oligonucleotides are often preferred over native forms because of desirable properties such as, for example, enhanced cellular uptake, enhanced affinity for nucleic acid target and increased stability in the presence of nucleases. An oligonucleotide preferably includes two or more nucleomono- mers covalently coupled to each other by linkages (e.g., phosphodiesters) or substitute linkages.
Overhang: An "overhang" is a relatively short single-stranded nucleotide sequence on the 5'- or 3'-hydroxyl end of a double-stranded oligonucleotide molecule (also referred to as an "extension," "protruding end," or "sticky end").
Polypeptide: The terms "polypeptide", "peptide", "oligopeptide", "polypeptide", "gene product", "expression product" and "protein" are used interchangeably herein to refer to a polymer or oligomer of consecutive amino acid residues.
“Precise” with respect to the introduction of a donor DNA molecule in target region means that the sequence of the donor DNA molecule is introduced into the target region without any InDeis, duplications or other mutations as compared to the unaltered DNA sequence of the target region that are not comprised in the donor DNA molecule sequence.
Promoter: The terms "promoter", or "promoter sequence" are equivalents and as used herein, refer to a DNA sequence which when operably linked to a nucleotide sequence of interest is capable of controlling the transcription of the nucleotide sequence of interest into RNA. A promoter is located 5' (i.e. , upstream), proximal to the transcriptional start site of a nucleotide sequence of interest whose transcription into RNA it controls, and provides a site for specific binding by RNA polymerase and other transcription factors for initiation of transcription. Said pro- moter comprises for example the at least 10 kb, for example 5 kb or 2 kb proximal to the transcription start site. It may also comprise the at least 1500 bp proximal to the transcriptional start site, preferably the at least 1000 bp. In a further preferred embodiment, the promoter comprises the at least 50 bp proximal to the transcription start site, for example, at least 25 bp. The promoter does not comprise exon and/or intron regions or 5' untranslated regions. The promoter may for example be heterologous or homologous to the respective cell. A polynucleotide sequence is "heterologous to" an organism or a second polynucleotide sequence if it originates from a foreign species, or, if from the same species, is modified from its original form. For example, a promoter operably linked to a heterologous coding sequence refers to a coding sequence from a species different from that from which the promoter was derived, or, if from the same species, a coding sequence which is not naturally associated with the promoter (e.g. a genetically engineered coding sequence or an allele from a different ecotype or variety). Suitable promoters can be derived from genes of the host cells where expression should occur or from pathogens for this host cells. Activity of a promoter may be evaluated by, for example, operably linking a reporter gene to the promoter sequence to generate a reporter construct, introducing the reporter construct into the genome of a cell and detecting the expression of the reporter gene (e.g., detecting mRNA, protein, or the activity of a protein encoded by the reporter gene).
Purified: As used herein, the term "purified" refers to molecules, either nucleic or amino acid sequences that are removed from their natural environment, isolated or separated. "Substantially purified" molecules are at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which they are naturally associated. A purified nucleic acid sequence may be an isolated nucleic acid sequence.
Recombinant: The term "recombinant" with respect to nucleic acid molecules refers to nucleic acid molecules produced by recombinant DNA techniques. Recombinant nucleic acid molecules may also comprise molecules, which as such does not exist in nature but are modified, changed, mutated or otherwise manipulated by man. Preferably, a "recombinant nucleic acid molecule" is a non-naturally occurring nucleic acid molecule that differs in sequence from a naturally occurring nucleic acid molecule by at least one nucleic acid. A “recombinant nucleic acid molecule” may also comprise a “recombinant construct” which comprises, preferably operably linked, a sequence of nucleic acid molecules not naturally occurring in that order. Preferred methods for producing said recombinant nucleic acid molecule may comprise cloning techniques, directed or non-directed mutagenesis, synthesis or recombination techniques.
Sense: The term "sense" is understood to mean a nucleic acid molecule having a sequence which is complementary or identical to a target sequence, for example a sequence which binds to a protein transcription factor and which is involved in the expression of a given gene. According to a preferred embodiment, the nucleic acid molecule comprises a gene of interest and elements allowing the expression of the said gene of interest.
Significant increase or decrease: An increase or decrease, for example in enzymatic activity or in gene expression, that is larger than the margin of error inherent in the measurement technique, preferably an increase or decrease by about 2-fold or greater of the activity of the control enzyme or expression in the control cell, more preferably an increase or decrease by about 5- fold or greater, and most preferably an increase or decrease by about 10-fold or greater.
Small nucleic acid molecules: “small nucleic acid molecules” are understood as molecules consisting of nucleic acids or derivatives thereof such as RNA or DNA. They may be doublestranded or single-stranded and are between about 15 and about 30 bp, for example between 15 and 30 bp, more preferred between about 19 and about 26 bp, for example between 19 and 26 bp, even more preferred between about 20 and about 25 bp for example between 20 and 25 bp. In an especially preferred embodiment, the oligonucleotides are between about 21 and about 24 bp, for example between 21 and 24 bp. In a most preferred embodiment, the small nucleic acid molecules are about 21 bp and about 24 bp, for example 21 bp and 24 bp.
Substantially complementary: In its broadest sense, the term "substantially complementary", when used herein with respect to a nucleotide sequence in relation to a reference or target nucleotide sequence, means a nucleotide sequence having a percentage of identity between the substantially complementary nucleotide sequence and the exact complementary sequence of said reference or target nucleotide sequence of at least 60%, more desirably at least 70%, more desirably at least 80% or 85%, preferably at least 90%, more preferably at least 93%, still more preferably at least 95% or 96%, yet still more preferably at least 97% or 98%, yet still more preferably at least 99% or most preferably 100% (the latter being equivalent to the term “identical” in this context). Preferably identity is assessed over a length of at least 19 nucleotides, preferably at least 50 nucleotides, more preferably the entire length of the nucleic acid sequence to said reference sequence (if not specified otherwise below). Sequence comparisons are carried out using default GAP analysis with the University of Wisconsin GCG, SEQWEB application of GAP, based on the algorithm of Needleman and Wunsch (Needleman and Wun- sch (1970) J Mol. Biol. 48: 443-453; as defined above). A nucleotide sequence "substantially complementary " to a reference nucleotide sequence hybridizes to the reference nucleotide sequence under low stringency conditions, preferably medium stringency conditions, most preferably high stringency conditions (as defined above). “Target site” as used herein means the position in the genome at which a double strand break or one or a pair of single strand breaks (nicks) are induced using recombinant technologies such as Zn-finger, TALEN, restriction enzymes, homing endonucleases, RNA-guided nucleases, RNA-guided nickases such as CRISPR/Cas nucleases or nickases and the like.
Transgene: The term "transgene" as used herein refers to any nucleic acid sequence, which is introduced into the genome of a cell by experimental manipulations. A transgene may be an "endogenous DNA sequence," or a "heterologous DNA sequence" (i.e., "foreign DNA"). The term "endogenous DNA sequence" refers to a nucleotide sequence, which is naturally found in the cell into which it is introduced so long as it does not contain some modification (e.g., a point mutation, the presence of a selectable marker gene, etc.) relative to the naturally-occurring sequence.
Transgenic: The term transgenic when referring to an organism means transformed, preferably stably transformed, with a recombinant DNA molecule that preferably comprises a suitable promoter operatively linked to a DNA sequence of interest.
Vector: As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it has been linked. One type of vector is a genomic integrated vector, or "integrated vector", which can become integrated into the chromosomal DNA of the host cell. Another type of vector is an episomal vector, i.e., a nucleic acid molecule capable of extra-chromosomal replication. Vectors capable of directing the expression of genes to which they are operatively linked are referred to herein as "expression vectors". In the present specification, "plasmid" and "vector" are used interchangeably unless otherwise clear from the context. Expression vectors designed to produce RNAs as described herein in vitro or in vivo may contain sequences recognized by any RNA polymerase, including mitochondrial RNA polymerase, RNA pol I, RNA pol II, and RNA pol III. These vectors can be used to transcribe the desired RNA molecule in the cell according to this invention.
Wild-type: The term "wild-type", "natural" or "natural origin" means with respect to an organism, polypeptide, or nucleic acid sequence, that said organism is naturally occurring or available in at least one naturally occurring organism which is not changed, mutated, or otherwise manipulated by man.
Figures:
Figure 1 shows the total specific phytase activity in U/g protein of the different phytase production strains. A: UV18#100f Apyr5 Aalpl Aku70 cbh1 ::Pchi-phytase-Tcbh1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at cbh1 locus)
B: UV18#100f Apyr5 Aalpl Aku70 cbh1 ::PTEF-phytase-Tcbh1 (overexpression of phytase gene driven by TEF promoter; expression cassette integrated at cbh1 locus)
C: UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at MYCTH:XP_003662453.1 locus)
D: UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::PTEF-phytase-Tcbh1 (overexpression of phytase gene driven by TEF promoter; expression cassette integrated at MYCTH:XP_003662453.1 locus).
Figure 2 shows the total specific glucanase activity in U/g protein of the different glucanase production strains.
A: UV18#100f Apyr5 Aalpl Aku70 cbh1 ::Pchi-glucanase-Tcbh1 (overexpression of glucanase gene driven by chi1 promoter; expression cassette integrated at cbh1 locus)
B: UV18#100f Apyr5 Aalpl Aku70 cbh1 ::PTEF-glucanase-Tcbh1 (overexpression of glucanase gene driven by TEF promoter; expression cassette integrated at cbh1 locus)
C: UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::Pchi-glucanase-Tcbh1 (overexpression of glucanase gene driven by chi1 promoter; expression cassette integrated at MYCTH:XP_003662453.1 locus)
D: UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::PTEF-glucanase-Tcbh1 (overexpression of glucanase gene driven by TEF promoter; expression cassette integrated at MYCTH:XP_003662453.1 locus).
Figure 3 shows the total specific phytase activity in U/g protein of the different phytase production strains.
A: UV18#100f Apyr5 Aalpl Aku70 cbh1 ::Pchi-phytase-Tcbh1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at cbh1 locus)
B: UV18#100f Apyr5 Aalpl Aku70 cbh1::Pchi-phytase-Tcbh1 MYCTH:XP_003662453.1::nat1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at cbh1 locus; MYCTH:XP_003662453.1 locus replaced by nat1 selection marker cassette)
C: UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at MYCTH:XP_003662453.1 locus)
D: UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 cbh1::nat1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at MYCTH:XP_003662453.1 locus; cbh1 locus replaced by nat1 selection marker cassette). Figure 4 shows the total specific phytase activity in U/g protein of the different phytase production strains.
A: UV18#100f Apyr5 Aalpl Aku70 cbh1 ::Pchi-phytase-Tcbh1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at cbh1 locus)
B: UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at MYCTH:XP_003662453.1 locus)
C: UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal
Aprt5 Aprt6 cbh1 ::Pchi-phytase-Tcbh1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at cbh1 locus)
D: UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 AprtB MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 (overexpression of phytase gene driven by chi1 promoter; expression cassette integrated at MYCTH:XP_003662453.1 locus).
Figure 5 shows the genomic organization of the genomic integration site of the invention in T. thermophilus) C1.
EXAMPLES
Chemicals and common methods
Unless indicated otherwise, cloning procedures carried out for the purposes of the present invention including restriction digest, agarose gel electrophoresis, purification of nucleic acids, Ligation of nucleic acids, transformation, selection and cultivation of bacterial cells were performed as described (Sambrook et al., 1989). Sequence analyses of recombinant DNA were performed with a laser fluorescence DNA sequencer (Applied Biosystems, Foster City, CA, USA) using the Sanger technology (Sanger et al., 1977). Unless described otherwise, chemicals and reagents were obtained from Sigma Aldrich (Sigma Aldrich, St. Louis, USA), from Promega (Madison, Wl, USA), Duchefa (Haarlem, The Netherlands) or Invitrogen (Carlsbad, CA, USA). Restriction endonucleases were from New England Biolabs (Ipswich, MA, USA). Oligonucleotides were synthesized by Integrated DNA Technologies (Coralville, IA, USA).
Example 1
Transformation of Thermothelomyces thermophilus
Several methods for the transformation of T. thermophilus protoplasts are described in the literature (WO 00/20555, US 2012/0005812, Verdoes et al. (2007) Industrial Biotechnology 3(1): 48-57). Protoplasts of T. thermophilus strains were prepared by inoculating a 25 ml preculture of a standard fungal growth media with 0.7-1x 105 spores/ml in a 100 ml shake flask for 24 h at 37°C and 250 rpm. The main culture was prepared by inoculating 100 ml of a standard fungal growth media with 20 ml of the preculture in a 500 ml shake flask for 24 h at 37°C and 250 rpm. The mycelium was harvested by filtration through a sterile Cell Strainer (VWR) and washed with 100 ml 2000 mosmol/L NaCI/CaCI2 (0.6 M NaCI and 0.27 M CaCI2*H2O). 1 g of the washed mycelium was transferred into a 100 ml flask. The mycelium was mixed with 150 mg VinoTaste Pro solution (2.5 mg/ml in 2000 mosmol/L NaCI/CaCI2) and 10 mg Yatalase solution (0.625 mg/ ml in 2000 mosmol/L NaCI/CaCI2) and 10 ml of 2000 mosmol/L NaCI/CaCI2. The mycelium suspension was incubated at 30°C and 70 rpm for 50-70 min until protoplasts are visible under the microscope. Harvesting of protoplasts was done by filtration through a sterile Cell Strainer into a sterile 50 ml tube. After the addition of 25 ml ice-cold STC solution (1.2 M sorbitol, 50 mM CaCI2, 35 mM NaCI, 10 mM Tris/HCI pH 7.5) to the flow through, the protoplasts were harvested by centrifugation (1200 x g, 10 min, 4°C). The protoplasts were washed again in 50 ml STC and resuspended in 0.5-1.2 ml STC.
For transformation, 5-10 pg of linearized DNA, 1 pl 0.5 M aurintricarboxylic acid (ATA) and 100 pl of protoplast suspension were mixed and incubated for 30-40 min on ice. Then 1.7 ml of PEG solution (60% PEG4000 [polyethylenglycol], 50 mM CaCI2, 35 mM NaCI, 10 mM Tris/HCI pH 7.5) was added and mixed gently. After incubation for 30 min at 4°C, the tube was filled with 11 ml STC solution, centrifuged (900 x g, 10 min, 4°C), and the supernatant discarded. The pellet was resuspended in the remaining STC and plated on selective media plates as known in the art. After incubation of the plates for 3-6 days at 37°C, transformants were picked and restreaked on selective media.
Selective media plates
If the amdS gene is used as selection marker, enriched minimal medium with uridine and uracil and with acetamide is used to select positive transformants (sucrose is only added in case protoplasts are plated):
Glucose 10 g/L
Sucrose 230 g/L
Mg2SO4*7H2O 0.49 g/L
KCI 0.52 g/L
KH2PO4 1.52 g/L
CUSO4*5H2O 1 .6 mg/L
FeSO4*7H2O 5 mg/L
ZnSO4*7H2O 22 mg/L MnSO4*H2O 4.3 mg/L
CoCI2*6H2O 1 .6 mg/L
Na2MoO4*2H2O 1.5 mg/L
H3BO3 11 mg/L
CsCI 2.52 g/L
EDTA 50 mg/L
Penicillin 50000 U/L
Streptomycin 50 mg/L
Uracil 1.12 g/L
Uridine 2.44 g/L
Acetamide 0.6 g/L
Agar 16 g/L
If selection of clones with lost acetamidase functionality is carried out, enriched minimal medium with uridine and uracil and with fluoroacetamide is used:
Glucose 10 g/L
Sucrose 230 g/L
Mg2SO4*7H2O 0.49 g/L KCI 0.52 g/L
KH2PO4 1.52 g/L
CUSO4*5H2O 1 .6 mg/L
FeSO4*7H2O 5 mg/L
ZnSO4*7H2O 22 mg/L
MnSO4*H2O 4.3 mg/L
COCI2*6H2O 1 .6 mg/L Na2MoO4*2H2O 1.5 mg/L H3BO3 11 mg/L
CsCI 2.52 g/L
EDTA 50 mg/L
Penicillin 50000 U/L
Streptomycin 50 mg/L
Uracil 1.12 g/L
Uridine 2.44 g/L
Fluoroacetamide 5 g/L
Urea 0.3 g/L
Agar 16 g/L If the pyr5 gene is used as selection marker, enriched minimal medium without uridine and uracil is used to select positive transformants (sucrose is only added in case protoplasts are plated):
Glucose 10 g/L
Sucrose 230 g/L
Mg2SO4*7H2O 0.49 g/L
KCI 0.52 g/L
KH2PO4 1.52 g/L
NaNO3 6 g/L
CUSO4*5H2O 1 .6 mg/L
FeSO4*7H2O 5 mg/L
ZnSO4*7H2O 22 mg/L
MnSO4*H2O 4.3 mg/L
COCI2*6H2O 1 .6 mg/L Na2MoO4*2H2O 1.5 mg/L H3BO3 1 1 mg/L
EDTA 50 mg/L
Penicillin 50000 U/L
Streptomycin 50 mg/L Casaminoacids 1 g/L Agar 16 g/L
If the nat1 gene is used as additional selection marker, enriched minimal medium without uridine and uracil and with nourseothricin (clonNAT) is used to select positive transformants (sucrose is only added in case protoplasts are plated):
Glucose 10 g/L
Sucrose 230 g/L
Mg2SO4*7H2O 0.49 g/L
KCI 0.52 g/L
KH2PO4 1.52 g/L
NaNO3 6 g/L
CUSO4*5H2O 1 .6 mg/L
FeSO4*7H2O 5 mg/L
ZnSO4*7H2O 22 mg/L
MnSO4*H2O 4.3 mg/L
COCI2*6H2O 1 .6 mg/L Na2MoO4*2H2O 1.5 mg/L
H3BO3 11 mg/L
EDTA 50 mg/L
Penicillin 50000 U/L
Streptomycin 50 mg/L
Nourseoth ricin 35 mg/L
Casaminoacids 1 g/L
Agar 16 g/L set pH to 6.5
Example 2
Generation of expression plasmids for integration at cbh1 locus
Expression construct with phytase and chi1 promoter
A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from bacterial origin (disclosed in WO 2012/143862; as phytase PhV-99; SEQ ID NO. 2) was used for the construction of a phytase expression plasmids. For the secretion of the phytase, a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the phytase. A promotor sequence amplified from the upstream region of the Chi1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the phytase. Using standard cloning techniques, the expression plasmid MT2701 (SEQ ID NO: 7) was constructed based on the E. coli standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5). Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, pyr5 gene, upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination, and lacZ for blue/white screening. Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening. The expression plasmid contains the upstream region of cbh1 from bases 1913 - 3345, the chi1 promotor sequence from bases 3350 - 5162, the phytase including a signal sequence from bases 5165 - 6478, the cbh1 terminator sequence from bases 6483 - 6689, and the downstream region of cbh1 from bases 7981 - 9610.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation. Expression construct with phytase and TEF1 promoter
A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from bacterial origin (disclosed in WO 2012/143862; as phytase PhV-99; SEQ ID NO. 2) was used for the construction of a phytase expression plasmids. For the secretion of the phytase, a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the phytase. A promotor sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH:XP_003660173.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the phytase. Using standard cloning techniques, the expression plasmid MT2703 (SEQ ID NO: 8) was constructed based on the E. coli standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5). Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, pyr5 gene, upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination, and lacZ for blue/white screening. Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening. The expression plasmid contains the upstream region of cbh1 from bases 1913 - 3345, the TEF1 promotor sequence from bases 3350 - 5847, the phytase including a signal sequence from bases 5850 - 7163, the cbh1 terminator sequence from bases 7168 - 7374, and the downstream region of cbh1 from bases 8666 - 10295.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
Expression construct with glucanase and chi1 promoter
A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) encoding a glucanase from Rasamsonia emersonii (SEQ ID NO. 3) was used for the construction of a glucanase expression plasmids. For the secretion of the glucanase, a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the glucanase. A promotor sequence amplified from the upstream region of the Chi 1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the glucanase. Using standard cloning techniques, the expression plasmid MT2705 (SEQ ID NO: 9) was constructed based on the E. coli standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5). Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, pyr5 gene, upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination, and lacZ for blue/white screening. Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening. The expression plasmid contains the upstream region of cbh1 from bases 1913 - 3345, the chi1 promotor sequence from bases 3350 - 5162, the glu- canase including a signal sequence from bases 5165 - 6151 , the cbh1 terminator sequence from bases 6156 - 6362, and the downstream region of cbh1 from bases 7654 - 9283.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
Expression construct with glucanase and TEF1 promoter
A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) encoding a glucanase from Rasamsonia emersonii (SEQ ID NO. 3) was used for the construction of a glucanase expression plasmids. For the secretion of the glucanase, a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the glucanase. A promotor sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the glucanase. Using standard cloning techniques, the expression plasmid MT2707 (SEQ ID NO: 10) was constructed based on the E. coll standard cloning vectors MT940 (SEQ ID NO: 4) and MT944 (SEQ ID NO: 5). Plasmid MT940 consists of the pMB1 origin of replication, kan resistance, pyr5 gene, upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination, and lacZ for blue/white screening. Plasmid MT944 consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening. The expression plasmid contains the upstream region of cbh1 from bases 1913 - 3345, the TEF1 promotor sequence from bases 3350 - 5847, the glucanase including a signal sequence from bases 5850 - 6836, the cbh1 terminator sequence from bases 6841 - 7047, and the downstream region of cbh1 from bases 8339 - 9968.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
Example 3
Generation of expression plasmids for integration at MYCTH:XP_003662453.1 locus Expression construct with phytase and chi1 promoter
A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from bacterial origin (disclosed in WO 2012/143862; as phytase PhV-99; SEQ ID NO. 2) was used for the construction of a phytase expression plasmids. For the secretion of the phytase, a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the phytase. A promotor sequence amplified from the upstream region of the Chi1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the phytase. Using standard cloning techniques, the expression plasmid MT2709 (SEQ ID NO: 11) was constructed based on the E. coli standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6). Both plasmids consist of the pUC origin of replication, kan resistance, and lacZ for blue/white screening. The expression plasmid contains the upstream region of MYCTH:XP_003662453.1 from bases 1913 - 3412, the chi1 promoter sequence from bases 3417 - 5229, the phytase including a signal sequence from bases 5232 - 6545, the cbh1 terminator sequence from bases 6550 - 6756, and the downstream region of MYCTH:XP_003662453.1 from bases 8048 - 9518.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
Expression construct with phytase and TEF1 promoter
A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) (SEQ ID NO: 1) encoding a synthetic phytase from bacterial origin (disclosed in WO 2012/143862; as phytase PhV-99; SEQ ID NO. 2) was used for the construction of a phytase expression plasmids. For the secretion of the phytase, a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the phytase. A promotor sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the phytase. Using standard cloning techniques, the expression plasmid MT2711 (SEQ ID NO: 12) was constructed based on the E. coli standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6). Both plasmids consist of the pUC origin of replication, kan resistance, and lacZ for blue/white screening. The expression plasmid contains the upstream region of MYCTH:XP_003662453.1 from bases 1913 - 3412, the TEF1 promotor sequence from bases 3417 - 5914, the phytase including a signal sequence from bases 5917 - 7230, the cbh1 terminator sequence from bases 7235 - 7441 , and the downstream region of MYCTH:XP_003662453.1 from bases 8733 - 10203.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the phytase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
Expression construct with glucanase and chi1 promoter
A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) encoding a glucanase from Rasamsonia emersonii (SEQ ID NO. 3) was used for the construction of a glucanase expression plasmids. For the secretion of the glucanase, a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the glucanase. A promotor sequence amplified from the upstream region of the Chi 1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the glucanase. Using standard cloning techniques, the expression plasmid MT2713 (SEQ ID NO: 13) was constructed based on the E. coll standard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6). Both plasmids consist of the pUC origin of replication, kan resistance, and lacZ for blue/white screening. The expression plasmid contains the upstream region of MYCTH:XP_003662453.1 from bases 1913 — 3412, the chi1 promotor sequence from bases 3417 - 5229, the glucanase including a signal sequence from bases 5232 - 6218, the cbh1 terminator sequence from bases 6223 - 6429, and the downstream region of MYCTH:XP_003662453.1 from bases 7721 - 9191 .
The plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
Expression construct with glucanase and TEF1 promoter
A synthetic gene (GeneArt, ThermoFisher Scientific Inc., USA) encoding a glucanase from Rasamsonia emersonii (SEQ ID NO. 3) was used for the construction of a glucanase expression plasmids. For the secretion of the glucanase, a signal sequence encoding for a signal peptide derived from T. thermophilus was added to the mature sequence of the glucanase. A promotor sequence amplified from the upstream region of the TEF1 encoding gene (MYCTH:XP_003663544.1) and a terminator sequence amplified from the downstream region of the Cbh1 encoding gene (MYCTH:XP_003660789.1) from T. thermophilus were used as regulatory elements to drive the expression of the glucanase. Using standard cloning techniques, the expression plasmid MT2715 (SEQ ID NO: 14) was constructed based on the E. coli stand- ard cloning vectors MT944 (SEQ ID NO: 5) and MT1266 (SEQ ID NO: 6). Both plasmids consist of the pUC origin of replication, kan resistance, and lacZ for blue/white screening. The expression plasmid contains the upstream region of MYCTH:XP_003662453.1 from bases 1913 — 3412, the TEF1 promotor sequence from bases 3417 - 5914, the glucanase including a signal sequence from bases 5917 - 6903, the cbh1 terminator sequence from bases 6908 - 7114, and the downstream region of MYCTH:XP_003662453.1 from bases 8406 - 9876.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the glucanase expression cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
Example 4
Generation of a plasmids for the deletion of cbh1 locus and of MYCTH:XP_003662453.1 locus
Deletion construct for MYCTH:XP_003662453.1 locus
A synthetic nourseothricin acetyltransferase gene (nat1) selection marker cassette (SEQ ID NO: 15) was fused with the upstream and downstream regions of the MYCTH:XP_003662453.1 locus from T. thermophilus for homologous recombination. Using standard cloning techniques, the KO plasmid MT2599 (SEQ ID NO: 16) was constructed based on the E. coli standard cloning vector MT944 (SEQ ID NO: 5). It consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the nat1 selection marker cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation.
Deletion construct for cbh1 locus
A synthetic nourseothricin acetyltransferase gene (nat1) selection marker cassette (SEQ ID NO: 15) was fused with the upstream and downstream regions of the Cbh1 encoding gene from T. thermophilus for homologous recombination. Using standard cloning techniques, the KO plasmid MT2699 (SEQ ID NO: 17) was constructed based on the E. coli standard cloning vector MT944 (SEQ ID NO: 5). It consists of the pUC origin of replication, kan resistance, and lacZ for blue/white screening.
The plasmid was digested with Not to remove the vector backbone and the fragment containing the nat1 selection marker cassette was isolated from an agarose gel. Only the isolated DNA fragment was later used for transformation. Example 5
Generation of T. thermophilus strains UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6
T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 was co-transformed as described in example 1 with the two isolated deletion fragments from plasmids MT185 (SEQ ID No. 18) and MT186 (SEQ ID NO. 19) to delete gene pep4. Enriched minimal medium for amdS selection was used for incubation. After re-streaking on enriched minimal medium with fluoroacetamide for counterselection, the transformants were analyzed by PCR for the correct integration of the deletion cassettes in the targeted locus, for the disappearance of the intact gene, and for removal of the amdS marker. The resulting strain UV18#100f Apyr5 Aalpl Aku70 Apep4 was modified in consecutive transformation rounds, and the following genes were deleted in the identical manner; gene prt1 with the two isolated deletion fragments from plasmids MT111 (SEQ ID No. 20) and MT112 (SEQ ID NO. 21), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl was obtained; gene prt2 with the two isolated deletion fragments from plasmids MT368 (SEQ ID No. 22) and MT453 (SEQ ID NO. 23), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 was obtained; gene prt3 with the two isolated deletion fragments from plasmids MT392 (SEQ ID No. 24) and MT393 (SEQ ID NO. 25), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 was obtained; gene prt4 with the two isolated deletion fragments from plasmids MT386 (SEQ ID No. 26) and MT649 (SEQ ID NO. 27), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 was obtained; gene Iam2 with the two isolated deletion fragments from plasmids MT706 (SEQ ID No. 28) and MT707 (SEQ ID NO. 29), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 was obtained; gene tre1 with the two isolated deletion fragments from plasmids MT710 (SEQ ID No. 30) and MT711 (SEQ ID NO. 31), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel was obtained; gene mycA with the two isolated deletion fragments from plasmids MT738 (SEQ ID No. 32) and MT739 (SEQ ID NO. 33), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA was obtained; gene gla1 with the two isolated deletion fragments from plasmids MT605 (SEQ ID No. 34) and MT606 (SEQ ID NO. 35), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal was obtained; gene prt5 with the two isolated deletion fragments from plasmids MT396 (SEQ ID No. 36) and MT397 (SEQ ID NO. 37), after removal of the amdS marker strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 was obtained; gene prtGwith the two isolated deletion fragments from plasmids MT388 (SEQ ID No. 38) and MT389 (SEQ ID NO. 39). Finally, positive tested clones were denoted as UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6. Example 6
Generation of T. thermophilus strains with phytase or glucanase expression constructs integrated at cbh1 locus
For the expression of a phytase, the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 (construction described in detail in WO 2017/093450) from the C1 lineage, a strain with uracil auxotrophy, reduced protease activity, and impaired non-homologous end joining (NHEJ) repair system, was transformed as described in example 1 with the /Vofl-digested and isolated phytase expression constructs (s. example 2) from plasmids MT2701 (SEQ ID NO: 7) and MT2703 (SEQ ID NO: 8). The T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6 (s. example 5) from the C1 lineage, a strain with uracil auxotrophy, reduced protease activity, and impaired non-homologous end joining (NHEJ) repair system, was transformed as described in example 1 with the /Vofl-digested and isolated phytase expression construct (s. example 2) from plasmid MT2701 (SEQ ID NO: 7). For the expression of a glucanase, the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 (construction described in detail in WO 2017/093450) was transformed as described in example 1 with the /Votl-digested and isolated glucanase expression constructs (s. example 2) from plasmids MT2705 (SEQ ID NO: 9) and MT2707 (SEQ ID NO: 10). The transformants were incubated for 3-6 days at 37°C on enriched minimal medium for pyr5 selection to select for restored uracil prototrophy by complementing the pyr5 deletion with the pyr5 marker as known in the art. Colonies were re-streaked and checked for the integration of the phytase or glucanase expression cassette using PCR with primer pairs specific for the phytase or glucanase expression cassette and the cbh1 locus as known in the art. A transformant tested positive for the phytase or glucanase expression construct at the cbh1 locus was selected for further characterization. The phytase production strain UV18#100f Apyr5 Aalpl Aku70 cbh1 ::Pchi-phytase-Tcbh1 was transformed as described in example 1 with the Not\ -digested and isolated nat1 selection marker cassette (s. example 4) from plasmid MT2599 (SEQ ID NO: 16). The transformants were incubated for 3-6 days at 37°C on enriched minimal medium for nat1 selection to select for resistance to nourseothricin (clonNAT) by integrating the nat1 marker as known in the art. Colonies were re-streaked and checked for the integration of the nat1 selection marker and the deletion of the MYCTH:XP_003662453.1 locus using PCR with primer pairs specific for nat1 and the MYCTH:XP_003662453.1 locus as known in the art.
Example 7
Generation of T. thermophilus strains with phytase or glucanase expression constructs integrated at MYCTH:XP_003662453.1 locus
For the expression of a phytase, the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 (construction described in detail in WO 2017/093450) from the C1 lineage, a strain with uracil auxotrophy, reduced protease activity, and impaired non-homologous end joining (NHEJ) repair system, was transformed as described in example 1 with the /Vofl-digested and isolated phytase expression constructs (s. example 3) from plasmids MT2709 (SEQ ID NO: 11) and MT2711 (SEQ ID NO: 12). The T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6 (s. example 5) from the C1 lineage, a strain with uracil auxotrophy, reduced protease activity, and impaired non-homologous end joining (NHEJ) repair system, was transformed as described in example 1 with the A/otl-digested and isolated phytase expression construct (s. example 3) from plasmid MT2709 (SEQ ID NO: 11). For the expression of a glucanase, the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 (construction described in detail in WO 2017/093450) was transformed as described in example 1 with the A/otl-digested and isolated glucanase expression constructs (s. example 3) from plasmids MT2713 (SEQ ID NO: 13) and MT2715 (SEQ ID NO: 14). The transformants were incubated for 3-6 days at 37°C on enriched minimal medium for pyr5 selection to select for restored uracil prototrophy by complementing the pyr5 deletion with the pyr5 marker as known in the art. Colonies were re-streaked and checked for the integration of the phytase or glucanase expression cassette using PCR with primer pairs specific for the phytase or glucanase expression cassette and the MYCTH:XP_003662453.1 locus as known in the art. A transformant tested positive for the phytase or glucanase expression construct at the MYCTH:XP_003662453.1 locus was selected for further characterization. The phytase production strain UV18#100f Apyr5 Aalpl Aku70 MYCTH:XP_003662453.1::Pchi-phytase-Tcbh1 was transformed as described in example 1 with the A/otl-digested and isolated nat1 selection marker cassette (s. example 4) from plasmid MT2699 (SEQ ID NO: 17). The transformants were incubated for 3-6 days at 37°C on enriched minimal medium for nat1 selection to select for resistance to nourseothricin (clonNAT) by integrating the nat1 marker as known in the art. Colonies were re-streaked and checked for the integration of the nat1 selection marker and the deletion of the cbh1 locus using PCR with primer pairs specific for nat1 and the cbh1 locus as known in the art.
Example 8
Phytase activity assay
The phytase activity is determined in microtiter plates. The phytase containing supernatant is diluted in reaction buffer (250 mM Na-acetate, 1 mM CaCh, 0.01 % Tween 20, pH 5.5). 10 pl of the enzyme solution are incubated with 140 pl substrate solution (6 mM Na-phytate (Sigma P3168) in reaction buffer) for 1 h at 37°C. The reaction is quenched by adding 150 pl of trichloroacetic acid solution (15% w/w). To detect the liberated phosphate, 20 pl of the quenched reaction solution are treated with 280 pl of freshly made-up color reagent (60 mM L-ascorbic acid (Sigma A7506), 2.2 mM ammonium molybdate tetrahydrate, 325 mM H2SO4), and incubated for 20 min at 37°C, and the absorption at 820 nm is subsequently determined. For the blank value, the substrate buffer on its own is incubated at 37°C and the 10 pl of enzyme sample are only added after quenching with trichloroacetic acid. The color reaction is performed analogously to the remaining measurements. The amount of liberated phosphate is determined via a calibration curve of the color reaction with a phosphate solution of known concentration.
Example 9
Glucanase activity assay
The glucanase activity is determined in microtiter plates using the DNS method. The glucanase containing supernatant is diluted in reaction buffer (100 mM Na-acetate, 0.005 % Tween 20, pH 4.5). As substrate, low viscosity carboxymethyl cellulose (CMC) was used which was dissolved at a concentration of 4% in reaction buffer with stirring while being heated up. Using a PCR Cycler 35 pl of the diluted enzyme solution is mixed with 35 pL of the substrate solution 30min at 40°C. The reaction is quenched by the addition of 105 pl of DNS solution (16 g/L NaOH, 10 g/L dinitrosalicylic acid, 300 g/L Na-K-tartrate tetrahydrate). After 20 min of heating using 80°C, the samples were cooled to 4°C and 150 pL of the solution is transferred to flat-bottom microtiter plates. The resulting dark orange color was detected at 540 nm. By subtracting a suitable blank and using a series of glucose solutions with known concentration the enzyme activity was determined and is reported as U/mL of enzyme solution with U is defined as pmol released reducing ends per min of reaction time.
Example 10
Analysis of phytase or glucanase production in small-scale cultivation
Generated mutant strains were fermented in small-scale cultivation and the supernatants were analyzed. T. thermophilus strains grown on agar were inoculated in 1 ml cultivation medium as shown in Table 1 in a 96-deepwell microtiter plate. The strains were fermented at 37°C on a microtiter plate shaker at 900 rpm and 80% humidity for 72 hours. 300 pl of the 72-hour- preculture were transferred in 700 pl cultivation medium as shown in Table 2 in a 96-deepwell microtiter plate. The strains were fermented at 37°C on a microtiter plate shaker at 900 rpm and 80% humidity for 96 hours.
Cell-free supernatants were harvested at the end of cultivation and subjected to a phytase activity assay (s. example 8) or a glucanase activity assay (s. example 9). As it can be seen in Figure 1, production of phytase is higher when the expression cassette was integrated at MYCTH:XP_003662453.1 locus compared to integration of the expression cassette at cbh1 locus, regardless whether expression was driven by chi1 promoter or TEF1 promoter. As it can be seen in Figure 2, production of glucanase is higher when the expression cassette was integrated at MYCTH:XP_003662453.1 locus compared to integration of the expression cassette at cbh1 locus, regardless whether expression was driven by chi1 promoter or TEF1 promoter. As it can be seen in Figure 3, production of phytase upon integration of the expression cassette at cbh1 locus is independent of the presence of MYCTH:XP_003662453.1 locus or the replace- ment of MYCTH:XP_003662453.1 locus by nat1 selection marker cassette; analogously, production of phytase upon integration of the expression cassette at MYCTH:XP_003662453.1 locus is independent of the presence of cbh1 locus or the replacement of cbh1 locus by nat1 selection marker cassette. As it can be seen in Figure 4, production of phytase upon integration of the expression cassette at cbh1 locus is higher when the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6 was used compared to usage of the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70; analogously, production of phytase upon integration of the expression cassette at MYCTH:XP_003662453.1 locus is higher when the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70 Apep4 Aprtl Aprt2 Aprt3 Aprt4 Alam2 Atrel AmycA Aglal Aprt5 Aprt6 was used compared to usage of the T. thermophilus host strain UV18#100f Apyr5 Aalpl Aku70.
Table 1 : Cultivation medium
Glucose 10 g/L
Mg2SO4*7H2O 0.49 g/L
K2SO4 1.21 g/L
CaSO4*2H2O 0.47 g/L
KH2PO4 1.75 g/L
(NH4)2SO4 4.63 g/L
CUSO4*5H2O 1 .46 mg/L
FeSO4*7H2O 4.56 mg/L
ZnSO4*7H2O 20.05 mg/L
MnSO4*H2O 3.92 mg/L
Na2MoO4*2H2O 1.37 mg/L
EDTA 50 mg/L
Biotin 0.006 mg/L
Penicillin 50000 U/L
Streptomycin 50 mg/L
Casamino acic i 1 g/L
MES 21.33 g/L set pH to 6.0
Table 2: Cultivation medium
Glucose 14.29 g/L
Mg2SO4*7H2O 0.65 g/L
K2SO4 1 .63 g/L
CaSO4*2H2O 0.63 g/L
KH2PO4 2.35 g/L
(NH4)2SO4 9.31 g/L CUSO4*5H2O 2.09 mg/L
FeSO4*7H2O 6.51 mg/L
ZnSO4*7H2O 28.64 mg/L
MnSO4*H2O 5.60 mg/L
Na2MoO4*2H2O 1.96 mg/L
EDTA 71.43 mg/L
Biotin 0.009 mg/L
Penicillin 71429 U/L
Streptomycin 71.43 mg/L
MES 60.93 g/L set pH to 6.75

Claims

What is claimed is:
1. An expression system comprising a. a filamentous fungi host cell comprising a native genomic region comprising a sequence selected from the list consisting of i) a nucleic acid sequence comprising SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60% to SEQ ID NO: 44, and iii) a fragment of at least 500 consecutive nucleic acids of a nucleic acid sequence of i) or ii). and b. a recombinant integrating nucleic acid molecule comprising a nucleic acid molecule flanked by distinct homology arms each at least 70% identical to at least 100 contiguous bases of said native genomic region.
2. The expression system of claim 1 wherein said native genomic region comprises a native gene encoding a small secreted protein flanked at one side by at least 250 nucleotides non-coding DNA and at the other side by at least 250 nucleotides non-coding DNA.
3. The expression system of claims 1 or 2 wherein said small secreted protein is comprising a sequence selected from the list of i. an amino acid sequence comprising SEQ ID NO: 46 ii. an amino acid sequence with a homology of at least 50% to SEQ ID NO: 46, iii. a fragment of at least 100 consecutive amino acids of i) or ii).
4. The expression system of claim 1 to 3 wherein the native genomic region is at one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 41 b. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
41 , and c. a fragment of at least 100 consecutive amino acids of a) or b), and at the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of d. an amino acid sequence comprising SEQ ID NO: 43 e. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
43, and f. a fragment of at least 100 consecutive amino acids of d) or e).
5. The expression system of claim 1 to 4 wherein the filamentous fungi host cell is from a species of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Au- reobasidium, Cryptococcus, Corynascus, Chrysosporium, Filibasidium, Fusarium, Humi- cola, Magnaporthe, Monascus, Mucor, Myceliophthora, Mortierella, Neocallimastix, Neu- rospora, Paecilomyces, Penicillium, Piromyces, Phanerochaete, Podospora, Pycnopo- rus, Rhizopus, Schizophyllum, Sordaria, Talaromyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypocladium, Trametes and Trichoderma, more preferred Aspergillus niger, Aspergillus oryzae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rasamsonia emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermothelomyces heterothallica and Thermothelomyces thermophilus (formerly known as Myceliophthora thermophila and as Chrysosporium lucknowense), most preferably Thermothelomyces thermophilus.
6. A method for stable integration of a recombinant polynucleotide in a filamentous fungi host cell comprising the steps of a. providing a filamentous fungi host cell comprising a native genomic region comprising a sequence selected from the list consisting of i) a nucleic acid sequence comprising SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60% to SEQ ID NO 44, and iii) a fragment of at least 500 consecutive nucleic acids of a nucleic acid sequence of i) or ii), b. introducing into said host cell a recombinant integrating nucleic acid molecule comprising a nucleic acid molecule flanked by distinct homology arms each at least 70% identical to at least 100 contiguous bases of said native genomic region, c. allowing for homologous recombination of said recombinant nucleic acid molecule into said native genomic region.
7. The method of claim 6 wherein said small secreted protein is comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 46 b. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
46, and c. a fragment of at least 100 consecutive amino acids of a) or b).
8. The method of claim 6 or 7 wherein said native genomic region is flanked at one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 41 b. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
41 , and c. a fragment of at least 100 consecutive amino acids of a) or b), and at the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of d. an amino acid sequence comprising SEQ ID NO: 43 e. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
43, and f. a fragment of at least 100 consecutive amino acids of d) or e).
9. The method of claim 6 to 8 wherein the filamentous fungi host cell is from a species selected from the group consisting of of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysosporium, Fili- basidium, Fusarium, Humicola, Magnaporthe, Monascus, Mucor, Myceliophthora, Mor- tierella, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Piromyces, Phanero- chaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordaria, Talaromyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypocladium, Trametes and Trichoderma, more preferred Aspergillus niger, Aspergillus oryzae, Aspergillus fu- migatus, Neurospora crassa, Penicillium chrysogenum, Penicillium citrinum, Acremonium chrysogenum, Trichoderma reesei, Rasamsonia emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermothelomyces heterothallica and Thermothelomyces thermophilus (formerly known as Myceliophthora thermophila and as Chrysosporium lucknowense), most preferably Thermothelomyces thermophilus.
10. A recombinant filamentous fungi host cell comprising a recombinant nucleic acid molecule flanked by genomic DNA comprising on one side of said recombinant nucleic acid molecule a sequence selected from the list consisting of i) a nucleic acid sequence comprising 1296 consecutive bases from the 5- prime end of SEQ ID NO: 44 ii)a nucleic acid sequence with an identity of at least 60% i), and iii) a fragment of at least 100 consecutive nucleic acids of a) or b). and comprising on the other side of said recombinant nucleic acid molecule a sequence selected from the list consisting of iv) a nucleic acid sequence comprising 1847 consecutive bases from the 3- prime end of SEQ ID NO: 44 v) a nucleic acid sequence with an identity of at least 60% to iv), and vi) a fragment of at least 100 consecutive nucleic acids of a) or b).
11. A recombinant filamentous fungi host cell comprising a recombinant nucleic acid molecule flanked at one side by the XP_003662454.1 gene and at the other side by the XP_003662452.1 gene.
12. The recombinant filamentous fungi host cell of claim 11 wherein the XP_003662454.1 gene is encoding an amino acid molecule selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 43, b. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
43, and c. a fragment of at least 100 consecutive amino acids of a) or b), and the XP_003662452.1 gene is encoding an amino acid molecule selected from the list consisting of d. an amino acid sequence comprising SEQ ID NO: 41 , e. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
41 , and f. a fragment of at least 100 consecutive amino acids of d) or e).
13. The recombinant filamentous fungi host cell of claims 10 to 12 wherein the host cell is from a species selected from the group consisting of of the phylum Ascomycota, preferably Acremonium, Aspergillus, Agaricus, Aureobasidium, Cryptococcus, Corynascus, Chrysosporium, Filibasidium, Fusarium, Humicola, Magnaporthe, Monascus, Mucor, Myceliophthora, Mortierella, Neocallimastix, Neurospora, Paecilomyces, Penicillium, Piromyces, Phanerochaete, Podospora, Pycnoporus, Rhizopus, Schizophyllum, Sordar- ia, Talaromyces, Rasamsonia, Thermoascus, Thermothelomyces, Thielavia, Tolypo- cladium, Trametes and Trichoderma, more preferred Aspergillus niger, Aspergillus ory- zae, Aspergillus fumigatus, Neurospora crassa, Penicillium chrysogenum, Penicillium cit- rinum, Acremonium chrysogenum, Trichoderma reesei, Rasamsonia emersonii (formerly known as Talaromyces emersonii), Aspergillus sojae, Thermothelomyces heterothallica and Thermothelomyces thermophilus (formerly known as Myceliophthora thermophila and as Chrysosporium lucknowense), most preferably Thermothelomyces thermophilus..
14. A recombinant nucleic acid molecule flanked by distinct homology arms each at least 70% identical to at least 100 contiguous bases of a native genomic region of a filamentous fungus comprising a sequence selected from the list consisting of i) a nucleic acid comprising SEQ ID NO: 44, ii)a nucleic acid sequence with an identity of at least 60% to SEQ ID NO: 44, and iii) a fragment of at least 500 consecutive nucleic acids of i) or ii).
15. A recombinant nucleic acid molecule flanked by distinct homology arms each at least 70% identical to at least 100 contiguous bases of a native genomic region of a filamentous fungus which is flanked at one side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of a. an amino acid sequence comprising SEQ ID NO: 41 , b. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
41 , and c. a fragment of at least 100 consecutive amino acids of a) or b), and at the other side by a sequence encoding an amino acid molecule comprising a sequence selected from the list consisting of d. an amino acid sequence comprising SEQ ID NO: 43, e. an amino acid sequence with a homology of at least 50% to SEQ ID NO:
43, and f. a fragment of at least 100 consecutive amino acids of d) or e).
16. A vector comprising the recombinant nucleic acid molecule of claim 14 or 15.
EP24702258.5A 2023-01-27 2024-01-18 New genomic integration site for fungal protein and fine chemicals production Pending EP4655405A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP23153687 2023-01-27
PCT/EP2024/051187 WO2024156593A1 (en) 2023-01-27 2024-01-18 New genomic integration site for fungal protein and fine chemicals production

Publications (1)

Publication Number Publication Date
EP4655405A1 true EP4655405A1 (en) 2025-12-03

Family

ID=85132880

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24702258.5A Pending EP4655405A1 (en) 2023-01-27 2024-01-18 New genomic integration site for fungal protein and fine chemicals production

Country Status (3)

Country Link
EP (1) EP4655405A1 (en)
CN (1) CN120641568A (en)
WO (1) WO2024156593A1 (en)

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
DE733059T1 (en) 1993-12-09 1997-08-28 Univ Jefferson CONNECTIONS AND METHOD FOR LOCATION-SPECIFIC MUTATION IN EUKARYOTIC CELLS
US6555732B1 (en) 1998-09-14 2003-04-29 Pioneer Hi-Bred International, Inc. Rac-like genes and methods of use
KR100618495B1 (en) 1998-10-06 2006-08-31 마크 아론 에말파브 Transformation Systems in Fibrous Fungal Host Regions:
US20120005812A1 (en) 2010-07-06 2012-01-12 Paul Corzatt Insect protective garment
ES2553458T3 (en) 2011-04-21 2015-12-09 Basf Se Synthetic variants of phytase
US20150337279A1 (en) * 2012-11-20 2015-11-26 Codexis, Inc. Recombinant fungal polypeptides
SG11201600115SA (en) 2013-07-10 2016-02-26 Novartis Ag Multiple proteases deficient filamentous fungal cells and methods of use thereof
CN105992821A (en) * 2013-12-03 2016-10-05 诺维信公司 Fungal gene library by double split-marker integration
BR112018009729A8 (en) 2015-12-02 2019-02-26 Basf Se method for producing a recombinant polypeptide, filamentous fungus myceliophthora thermophila and use of a nucleic acid construct
EP3921434A4 (en) 2019-02-10 2022-11-30 Dyadic International (USA), Inc. Production of cannabinoids in filamentous fungi

Also Published As

Publication number Publication date
CN120641568A (en) 2025-09-12
WO2024156593A1 (en) 2024-08-02

Similar Documents

Publication Publication Date Title
KR102400332B1 (en) Recombinant microorganism for improved production of fine chemicals
KR102321146B1 (en) Recombinant microorganism for improved production of fine chemicals
US12492406B2 (en) Regulatory nucleic acid molecules for enhancing gene expression in plants
KR20180011324A (en) Recombinant microorganism for improved production of alanine
KR20180011313A (en) Recombinant microorganism for improved production of fine chemicals
US20230075913A1 (en) Codon-optimized cas9 endonuclease encoding polynucleotide
EP3436574B1 (en) Fermentative production of n-butylacrylate using alcohol acyl transferase enzymes
JP5729702B2 (en) A method for improving the production amount of kojic acid using a gene essential for the production of kojic acid
Singleton et al. Primary structure and regulation of vegetative specific genes of Dictyostelium discoideum
EP2739737A1 (en) Method for identification and isolation of terminator sequences causing enhanced transcription
Fernandez et al. Vector-initiated transitive RNA interference in the filamentous fungus Aspergillus oryzae
CN113825838A (en) Regulatory nucleic acid molecules for enhancing gene expression in plants
US20230383321A1 (en) Improved method for the production of isoprenoids
EP4655405A1 (en) New genomic integration site for fungal protein and fine chemicals production
US20220017931A1 (en) Method for production of 4-cyano benzoic acid or salts thereof
WO2024132947A1 (en) New cellulase promoters for fungal protein production
WO2021069387A1 (en) Regulatory nucleic acid molecules for enhancing gene expression in plants
US12325861B2 (en) Regulatory nucleic acid molecules for enhancing gene expression in plants
WO2025262103A1 (en) Superior promoters for fungal protein production
US20260125694A1 (en) Regulatory nucleic acid molecules for enhancing gene expression in plants
WO2024083579A1 (en) Regulatory nucleic acid molecules for enhancing gene expression in plants
WO2025042825A9 (en) Methods and means for influencing expression of heteroallelic genes or alleles in plants by modification of untranslated regions

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250827

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR