WO2025101486A1 - Methods and compositions for enhanced protein production in bacillus cells - Google Patents
Methods and compositions for enhanced protein production in bacillus cells Download PDFInfo
- Publication number
- WO2025101486A1 WO2025101486A1 PCT/US2024/054516 US2024054516W WO2025101486A1 WO 2025101486 A1 WO2025101486 A1 WO 2025101486A1 US 2024054516 W US2024054516 W US 2024054516W WO 2025101486 A1 WO2025101486 A1 WO 2025101486A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- seq
- operably linked
- nucleic acid
- poi
- downstream
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C07—ORGANIC CHEMISTRY
- C07K—PEPTIDES
- C07K14/00—Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof
- C07K14/195—Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof from bacteria
- C07K14/32—Peptides having more than 20 amino acids; Gastrins; Somatostatins; Melanotropins; Derivatives thereof from bacteria from Bacillus (G)
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
- C12N9/14—Hydrolases (3)
- C12N9/24—Hydrolases (3) acting on glycosyl compounds (3.2)
- C12N9/2402—Hydrolases (3) acting on glycosyl compounds (3.2) hydrolysing O- and S- glycosyl compounds (3.2.1)
- C12N9/2405—Glucanases
- C12N9/2408—Glucanases acting on alpha -1,4-glucosidic bonds
- C12N9/2411—Amylases
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12P—FERMENTATION OR ENZYME-USING PROCESSES TO SYNTHESISE A DESIRED CHEMICAL COMPOUND OR COMPOSITION OR TO SEPARATE OPTICAL ISOMERS FROM A RACEMIC MIXTURE
- C12P21/00—Preparation of peptides or proteins
- C12P21/02—Preparation of peptides or proteins having a known sequence of two or more amino acids, e.g. glutathione
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12R—INDEXING SCHEME ASSOCIATED WITH SUBCLASSES C12C - C12Q, RELATING TO MICROORGANISMS
- C12R2001/00—Microorganisms ; Processes using microorganisms
- C12R2001/01—Bacteria or Actinomycetales ; using bacteria or Actinomycetales
- C12R2001/07—Bacillus
- C12R2001/10—Bacillus licheniformis
Definitions
- RESULTS AND COMPOSITIONS FOR ENHANCED PROTEIN PRODUCTION IN BACILLUS CELLS FIELD The present disclosure is generally related to the fields of microbial cells, molecular biology, fermentation, protein production, and the like. Certain aspects of the disclosure are related to, inter alia, recombinant Bacillus cells having enhanced protein production capabilities. CROSS REFERENCE TO RELATED APPLICATIONS [0002] This application claims benefit to U.S. Provisional Patent Application No. 63/597,527, filed November 9, 2023, which is incorporated herein by referenced in its entirety.
- the production of proteins e.g., enzymes, antibodies, receptors, etc.
- Gram-positive bacterial cells are highly desirable and unmet needs for obtaining, constructing, producing and the like, Gram-positive host cells having increased protein production capabilities. More specifically, certain aspects of the instant disclosure are related to, among other things, compositions and methods for constructing recombinant Bacillus cells having enhanced protein production capabilities, which recombinant cells are particularly useful in the production of heterologous proteins.
- certain embodiments of the disclosure are related to recombinant (modified) Bacillus cells (strains) capable of expressing/producing increased amounts of proteins of interest. Certain embodiments of the disclosure therefore provide, inter alia, nucleic acids, polynucleotides, vectors, expression cassettes, signal (peptide) sequences, proteins of interest, microbial cells, methods for constructing Bacillus cells producing proteins of interest, methods for cultivating Bacillus cells for the expression/production of proteins of interest and the like. [0007] In particular embodiments, the disclosure provides an isolated nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) set for in SEQ ID NO: 2.
- the disclosure provides a polynucleotide comprising an upstream (5 ⁇ ) nucleic acid encoding a native ypuA signal sequence (ypuAss) set forth in SEQ ID NO: 1 operably linked to a downstream (3 ⁇ ) nucleic acid encoding a heterologous protein of interest (POI).
- ypuAss native ypuA signal sequence
- POI heterologous protein of interest
- the disclosure provides a polynucleotide comprising an upstream (5 ⁇ ) nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to a downstream (3 ⁇ ) nucleic acid encoding a heterologous POI.
- the disclosure is related to expression cassettes comprising an upstream (heterologous) promoter operably linked to a downstream 5′-UTR operably linked to a downstream polynucleotide comprising an upstream nucleic acid encoding the ypuAss (SEQ ID NO: 1) operably linked to the downstream nucleic acid encoding the POI, wherein the promoter and 5′-UTR sequences are functional in Bacillus sp. cells.
- the disclosure is related to methods for producing a heterologous POI in a Bacillus sp. cell comprising (a) introducing into a Bacillus sp. cell an expression cassette comprising an upstream (5′) promoter operably linked to a downstream 5′- UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding a heterologous POI, and (b) fermenting the modified cell under suitable conditions for the production of the POI.
- an expression cassette comprising an upstream (5′) promoter operably linked to a downstream 5′- UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1,
- the modified cell produces an increased amount of the POI relative to a control cell producing the same POI
- the control cell comprises an introduced expression cassette comprising the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a modified AmyL signal sequence (mod-AmyLss) of SEQ ID NO: 4 operably linked to a downstream nucleic acid encoding the heterologous POI.
- modified AmyL signal sequence (mod-AmyLss) of SEQ ID NO: 4 operably linked to a downstream nucleic acid encoding the heterologous POI.
- the native ypuA signal sequence (SEQ ID NO: 1) comprises twenty-five (25) amino acid residues with an aspartic acid (Asp; D24) residue at position 24 of SEQ ID NO: 1
- the modified ypuA signal sequence (SEQ ID NO: 2) comprises a substitution of the aspartic acid (D24) to a serine (Ser; S24) residue at position 24 of SEQ ID NO: 2.
- FIG. 1B shows an alignment of the native B.
- the native B. licheniformis AmyL signal sequence (AmyLss; SEQ ID NO: 3) comprises twenty-nine (29) amino acid residues with an alanine (Ala; A28) residue at position 28 of SEQ ID NO: 3, whereas the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) comprises a substitution of the alanine (A28) to a serine (Ser; S28) residue at position 28 of SEQ ID NO: 4.
- FIG.2A presents schematics showing the general design of the amylase-1 expression cassettes suitable for integration into Gram-positive bacterial cells of the disclosure.
- a first copy of the amylase-1 integration cassette (1 st copy amy-1 cassette with modified ypuA signal sequence; SEQ ID NO: 7) comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream (DNA) encoding a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to a downstream DNA encoding a B.
- SEQ ID NO: 7 comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream (DNA) encoding a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to
- subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to a downstream DNA encoding the Amylase-1 (Amy-1; SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13).
- modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to a downstream DNA encoding the Amylase-1 (Amy-1; SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO: 12)
- FIG. 2B presents the sequence identification numbers (SEQ) of the 1st copy amylase-1 cassette (SEQ ID NO: 7) shown in FIG. 2A.
- FIG. 2C shows the design of the second copy of the amylase-1 integration cassette (2 nd copy amy-1 cassette with modified ypuA signal sequence; SEQ ID NO: 17) comprising an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B.
- SEQ sequence identification numbers
- subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to a downstream DNA encoding the Amylase-1 (Amy-1; SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21).
- modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to a downstream DNA encoding the Amylase-1 (Amy-1; SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO:
- FIG.2D presents the sequence identification numbers (SEQ) of the 2 nd copy amylase-1 integration cassette (SEQ ID NO: 17) shown in FIG.2C.
- SEQ sequence identification numbers
- FIG. 2E-FIG. 2H a first copy amylase-1 integration cassette (1 st copy amy-1 cassette with modified AmyL signal sequence; SEQ ID NO: 30) and a second copy amylase-1 integration cassette (2 nd copy amy-1 cassette with modified AmyL signal sequence; SEQ ID NO: 31) were constructed as controls.
- the 1 st copy amy-1 cassette (SEQ ID NO: 30) comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream (DNA) encoding a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B.
- FIG.2F presents the sequence identification numbers (SEQ) of the 1st copy amylase-1 cassette (SEQ ID NO: 30) shown in FIG.2E.
- the 2 nd copy amy-1 cassette (SEQ ID NO: 31) comprises an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B.
- FIG.2H presents the sequence identification numbers (SEQ) of the 2 nd copy amylase-1 integration cassette (SEQ ID NO: 31) shown in FIG.2G.
- FIG.3A presents schematics showing the general design of certain amylase-2 expression cassettes suitable for integration into Gram-positive bacterial cells of the disclosure.
- a first copy of the amylase-2 cassette (1 st copy amy-2 cassette with mod-AmyLss; SEQ ID NO: 39) comprises a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B.
- FIG. 3B presents the sequence identification numbers (SEQ) of the 1st copy amylase-2 cassette (SEQ ID NO: 39) shown in FIG. 3A. As shown in FIG.
- the 2 nd copy amy-2 cassette with ypuAss comprises an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the native B.
- up for the lysA locus
- lysA.up operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF
- SEQ ID NO: 19 operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B. subtilis apr
- FIG.3D presents the sequence identification numbers (SEQ) of the 2 nd copy amylase-2 integration cassette (SEQ ID NO: 38) shown in FIG. 3C. As presented in FIG.
- a 2 nd copy amy-2 cassette with mod-AmyLss comprises an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B.
- FIG.3F presents the sequence identification numbers (SEQ) of the 2 nd copy amylase-2 integration cassette (SEQ ID NO: 40) shown in FIG.3E.
- SEQ ID NO: 1 is the amino acid sequence of the native B. licheniformis ypuA signal sequence (ypuAss).
- SEQ ID NO: 2 is the amino acid sequence of a modified ypuA signal sequence (mod-ypuAss).
- SEQ ID NO: 3 is the amino acid sequence the native B. licheniformis AmyL signal sequence (AmyLss).
- SEQ ID NO: 4 is the amino acid sequence of a modified AmyL signal sequence (mod-AmyLss).
- SEQ ID NO: 5 is a DNA template for the first (1 st ) copy amylase-1 expression cassette.
- SEQ ID NO: 6 is the mature amino acid sequence of an engineered (variant) B. licheniformis Amylase-1 (Amy-1) reporter protein.
- SEQ ID NO: 7 is an integration (expression) cassette encoding the 1 st copy Amyl-1 reporter protein with mod-ypuAss.
- SEQ ID NO: 8 is a B. licheniformis serA upstream (serA.up) homology arm (DNA) sequence.
- SEQ ID NO: 9 is the open reading frame (ORF) of the B. licheniformis serA gene.
- SEQ ID NO: 10 is a synthetic p3 promoter sequence.
- SEQ ID NO: 11 is a wild-type B.
- SEQ ID NO: 12 is a wild-type B. licheniformis AmyL terminator sequence.
- SEQ ID NO: 13 is a B. licheniformis serA downstream (serA.down) homology arm sequence.
- SEQ ID NO: 14 is a synthetic primer sequence named “oAL106”.
- SEQ ID NO: 15 is a synthetic primer sequence named “oAL108”.
- SEQ ID NO: 16 is a DNA template for the second (2 nd ) copy amylase-1 expression cassette.
- SEQ ID NO: 17 is an integration (expression) cassette encoding the 2 nd copy Amy-1 reporter protein with mod-ypuAss.
- SEQ ID NO: 18 is a B. licheniformis lysA upstream (lysA.up) homology arm sequence.
- SEQ ID NO: 19 is the open reading frame (ORF) of the B. licheniformis lysA gene.
- SEQ ID NO: 20 is a synthetic p2 promoter sequence.
- SEQ ID NO: 21 is a B. licheniformis lysA downstream (lysA.down) homology arm sequence.
- SEQ ID NO: 22 is a synthetic primer sequence named “oAL079”.
- SEQ ID NO: 23 is a synthetic primer sequence named “oAL080”.
- SEQ ID NO: 24 is a synthetic primer sequence named “oAL069”.
- SEQ ID NO: 25 is a synthetic primer sequence named “oAL076”.
- SEQ ID NO: 26 is a 3,709 bp DNA fragment for screening integration of the 1 st copy amy-1 cassette (SEQ ID NO: 7) with the mod-ypuAss.
- SEQ ID NO: 27 is a synthetic primer sequence named “oAL082”.
- SEQ ID NO: 28 is a synthetic primer sequence named “oAL083”.
- SEQ ID NO: 29 is a 2,133 bp DNA fragment for screening integration of the 2 nd copy amy-1 cassette (SEQ ID NO: 17) with the mod-ypuAss.
- SEQ ID NO: 30 is a control integration (expression) cassette encoding the 1 st copy Amyl-1 reporter protein with mod-AmyLss.
- SEQ ID NO: 31 is a control integration (expression) cassette encoding the 2 nd copy Amyl-1 reporter protein with mod-AmyLss.
- SEQ ID NO: 32 is a synthetic primer sequence named “oAL081”.
- SEQ ID NO: 33 is a synthetic primer sequence named “seq1”.
- SEQ ID NO: 34 is a synthetic primer sequence named “seq2”.
- SEQ ID NO: 35 is a synthetic primer sequence named “oAL086”.
- SEQ ID NO: 36 is a plasmid (DNA) template (pWS704) for a second (2 nd ) copy amylase-2 integration (expression) cassette encoding a 2 nd copy of the Amy-2 protein (SEQ ID NO: 37) with the native ypuAss (SEQ ID NO: 1)
- SEQ ID NO: 37 is the mature amino acid sequence of an engineered (variant) Cytophaga sp. Amylase-2 (Amy-2) reporter protein.
- SEQ ID NO: 38 is an integration (expression) cassette encoding the 2 nd copy Amy-2 reporter protein with ypuAss.
- SEQ ID NO: 39 is an integration (expression) cassette encoding the 1 st copy Amy-2 reporter protein with mod-AmyLss.
- SEQ ID NO: 40 is an integration (expression) cassette encoding the 2 nd copy Amy-2 reporter protein with mod-AmyLss.
- SEQ ID NO: 41 is a synthetic primer sequence named “ws683”.
- SEQ ID NO: 42 is a synthetic primer sequence named “ws688”.
- SEQ ID NO: 43 is a synthetic primer sequence named “ws775”.
- SEQ ID NO: 44 is a synthetic primer sequence named “ws776”.
- SEQ ID NO: 45 is a 1,892 bp DNA fragment for screening integration of the 2 nd copy amy-2 cassette (SEQ ID NO: 38).
- SEQ ID NO: 46 is an expression cassette named “(p3 pro)-(aprE 5′-UTR)-(mod-ypuAss)-(amy-1)”.
- SEQ ID NO: 47 is an expression cassette named “(p2 pro)-(aprE 5′-UTR)-(mod-ypuAss)-(amy-1)”.
- SEQ ID NO: 48 is an expression cassette named “(p3 pro)-(aprE 5′-UTR)-(ypuAss)-(amy-1)”.
- SEQ ID NO: 49 is an expression cassette named “(p2 pro)-(aprE 5′-UTR)-(ypuAss)-(amy-1)”.
- SEQ ID NO: 50 is an expression cassette named “(p3 pro)-(aprE 5′-UTR)-(mod-ypuAss)-(amy-2)”.
- SEQ ID NO: 51 is an expression cassette named “(p2 pro)-(aprE 5′-UTR)-(mod-ypuAss)-(amy-2)”.
- SEQ ID NO: 52 is an expression cassette named “(p3 pro)-(aprE 5′-UTR)-(ypuAss)-(amy-2)”.
- SEQ ID NO: 53 is an expression cassette named “(p2 pro)-(aprE 5′-UTR)-(ypuAss)-(amy-2)”.
- Certain embodiments of the instant disclosure provide, inter alia, recombinant Bacillus cells capable expressing increased amounts of proteins of interest. Certain aspects of the disclosure therefore provide, among other things, novel (recombinant) Bacillus cells comprising introduced nucleic acids (e.g., vectors, expression cassettes) encoding proteins of interest, polynucleotide constructs encoding modified (protein) signal sequences operably linked to a downstream nucleic acid a encoding protein of interest, recombinant Bacillus cells comprising one or more introduced polynucleotide constructs, and related methods for cultivating and expressing heterologous proteins of interest in a recombinant Bacillus cell of the disclosure and the like.
- nucleic acids e.g., vectors, expression cassettes
- polynucleotide constructs encoding modified (protein) signal sequences operably linked to a downstream nucleic acid a encoding protein of interest
- recombinant Bacillus cells comprising one or more
- the terms “recombinant” or “non-natural” refer to an organism, microorganism, cell, nucleic acid molecule, or vector that has at least one engineered genetic alteration, or has been modified by the introduction of a heterologous nucleic acid molecule, or refer to a cell (e.g., a Gram-positive cell) that has been altered such that the expression of a heterologous nucleic acid molecule or an endogenous nucleic acid molecule or a gene can be controlled.
- Recombinant also refers to a cell that is derived from a non-natural cell, or is progeny of a non-natural cell having one or more such modifications.
- Genetic alterations include, for example, modifications introducing expressible nucleic acid molecules encoding proteins, or other nucleic acid molecule additions, deletions, substitutions or other functional alteration of a cell’s genetic material.
- recombinant cells may express genes or other nucleic acid molecules (e.g., polynucleotide expression constructs) that are not found in identical or homologous form within a native (wild-type) cell, or may provide an altered expression pattern of endogenous genes, such as being over-expressed, under-expressed, minimally expressed, or not expressed at all.
- “Recombination”, “recombining” or generating a “recombined” nucleic acid is generally the assembly of two or more nucleic acid fragments wherein the assembly gives rise to a chimeric DNA sequence that would not otherwise be found in the genome.
- the term “derived” encompasses the terms “originated”, “obtained”, “obtainable”, and “created” and generally indicates that one specified material or composition finds its origin in another specified material or composition, or has features that can be described with reference to the other specified material or composition.
- nucleic acid refers to a nucleotide or polynucleotide sequence, and fragments or portions thereof, as well as to DNA, cDNA, and RNA of genomic or synthetic origin, which may be double- stranded or single-stranded, whether representing the sense or antisense strand. It will be understood that as a result of the degeneracy of the genetic code, a multitude of nucleotide sequences may encode a given protein. [0074] It is understood that the polynucleotides (or nucleic acid molecules) described herein include “genes”, “vectors” and “plasmids”.
- the term “gene”, refers to a polynucleotide that codes for a particular sequence of amino acids, which comprise all, or part of a protein coding sequence, and may include regulatory (non- transcribed) DNA sequences, such as promoter sequences, which determine for example the conditions under which the gene is expressed.
- the transcribed region of the gene may include untranslated regions (UTRs), including introns, 5′-untranslated regions (UTRs), and 3′-UTRs, as well as the coding sequence.
- an “endogenous gene” refers to a gene in its natural location in the genome of an organism.
- a “heterologous” gene, a “non-endogenous” gene, or a “foreign” gene refer to a gene not normally found in the host organism, but that is introduced into the host organism by gene transfer.
- the term “foreign” gene(s) comprises native genes inserted into a non-native organism and/or chimeric genes inserted into a native or non-native organism.
- a “heterologous control sequence” refers to a gene expression control sequence (e.g., promoters, enhancers, terminators, etc.) which does not function in nature to regulate (control) the expression of the gene of interest.
- heterologous nucleic acids are not endogenous (native) to the cell, or a part of the genome in which they are present, and have been added to the cell, by infection, transfection, transduction, transformation, microinjection, electroporation, and the like.
- a “heterologous” nucleic acid construct may contain a control sequence/DNA coding (ORF) sequence combination that is the same as, or different, from a control sequence/DNA coding sequence combination found in the native host cell.
- ORF control sequence/DNA coding
- signal sequence and “signal peptide” refer to a sequence of amino acid residues that may participate in the secretion or direct transport of a mature protein or precursor form of a protein.
- the signal sequence is typically located N-terminal to the precursor or mature protein sequence.
- the signal sequence may be endogenous or exogenous.
- a signal sequence is normally absent from the mature protein.
- a signal sequence is typically cleaved from the protein by a signal peptidase during translocation.
- the term “expression” refers to the transcription and stable accumulation of sense (mRNA) or anti-sense RNA, derived from a nucleic acid molecule of the disclosure. Expression may also refer to translation of mRNA into a polypeptide. Thus, the term “expression” includes any steps involved in the production of the polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, secretion and the like.
- coding sequence refers to a nucleotide sequence, which directly specifies the amino acid sequence of its (encoded) protein product.
- the boundaries of the coding sequence are generally determined by an open reading frame (hereinafter, “ORF”), which usually begins with an ATG start codon.
- ORF open reading frame
- the coding sequence typically includes DNA, cDNA, and recombinant nucleotide sequences.
- promoter refers to a nucleic acid sequence capable of controlling the expression of a coding sequence or functional RNA. In general, a coding sequence is located 3' (downstream) to a promoter sequence.
- Promoters may be derived in their entirety from a native gene, or be composed of different elements derived from different promoters found in nature, or even comprise synthetic nucleic acid segments. It is understood by those skilled in the art that different promoters may direct the expression of a gene in different cell types, or at different stages of development, or in response to different environmental or physiological conditions. Promoters can be constitutive promoters, inducible promoters, tunable promoters, hybrid promoters, synthetic promoters, tandem promoters, etc. Promoters which cause a gene to be expressed in most cell types at most times are commonly referred to as “constitutive promoters”.
- a functional promoter sequence controlling the expression of a gene of interest linked to the gene of interest refers to a promoter sequence which controls the transcription and translation of the coding sequence in a desired Gram-positive host cell.
- the present disclosure is directed to a polynucleotide comprising an upstream (5′) promoter (or 5′ promoter region, or tandem 5′ promoters and the like) functional in a Gram-positive cell, wherein the promoter region is operably linked to a nucleic acid sequence encoding a protein of interest.
- operably linked refers to the association of nucleic acid sequences on a single nucleic acid fragment so that the function of one is affected by the other.
- a promoter is operably linked with a coding sequence when it is capable of affecting the expression of that coding sequence (i.e., that the coding sequence is under the transcriptional control of the promoter).
- Coding sequences can be operably linked to regulatory sequences in sense or antisense orientation.
- a nucleic acid is “operably linked” when it is placed into a functional relationship with another nucleic acid sequence.
- DNA encoding a secretory leader is operably linked to DNA for a polypeptide if it is expressed as a pre-protein that participates in the secretion of the polypeptide; a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of the sequence; or a ribosome binding site is operably linked to a coding sequence if it is positioned so as to facilitate translation.
- “operably linked” means that the DNA sequences being linked are contiguous, and, in the case of a secretory leader, contiguous and in reading phase. However, enhancers do not have to be contiguous. Linking is accomplished by ligation at convenient restriction sites.
- suitable regulatory sequences refer to nucleotide sequences located upstream (5′ non-coding sequences), within, or downstream (3′ non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include promoters, transcription leader sequences, RNA processing site, effector binding site and stem-loop structures.
- ypuA signal sequence (abbreviated, “ypuAss”) comprises the amino acid sequence set forth in SEQ ID NO: 1.
- the native ypuAss (SEQ ID NO: 1) comprises twenty-five (25) amino acid residues shown as positions M1-A25 (FIG.1A).
- a modified (variant) B. licheniformis “ypuA signal sequence” (abbreviated, “mod- ypuAss”) comprises the amino acid sequence set forth in SEQ ID NO: 2.
- the mod-ypuAss (SEQ ID NO: 2) comprises twenty-five (25) amino acid residues shown as positions M 1 -A 25 (FIG.1A).
- a native B As used herein, a native B.
- AmyL signal sequence (abbreviated, “AmyLss”) comprises the amino acid sequence set forth in SEQ ID NO: 3.
- the native AmyLss (SEQ ID NO: 3) comprises twenty-nine (29) amino acid residues shown as positions M 1 - A 29 (FIG.1B).
- a modified (variant) B. licheniformis “AmyL signal sequence” (abbreviated, “mod- AmyLss”) comprises the amino acid sequence set forth in SEQ ID NO: 4.
- the AmyLss (SEQ ID NO: 4) comprises twenty-nine (29) amino acid residues shown as positions M 1 -A 29 (FIG. 1B).
- recombinant B. licheniformis strains expressing heterologous proteins of interest operably linked to the upstream (N-terminal) mod-AmyLss (SEQ ID NO: 4) have been described in PCT Publication No. WO 2023/023642.
- a B. licheniformis “serA upstream homology arm” (abbreviated, “serA.up”) comprises the nucleic acid sequence set forth in SEQ ID NO: 8, a B.
- licheniformis “serA downstream homology arm” (abbreviated, “serA.down”) comprises the nucleic acid sequence set forth in SEQ ID NO: 13, and the “open reading frame” (ORF) of the B. licheniformis serA gene comprises the nucleic acid sequence set forth in SEQ ID NO: 9.
- a B. licheniformis “lysA upstream homology arm” (abbreviated, “lysA.up”) comprises the nucleic acid sequence set forth in SEQ ID NO: 18, a B.
- licheniformis “lysA downstream homology arm” (abbreviated, “lysA.down”) comprises the nucleic acid sequence set forth in SEQ ID NO: 21, and the ORF of the B. licheniformis lysA gene comprises the nucleic acid sequence set forth in SEQ ID NO: 19.
- the “synthetic p3 promoter sequence” (abbreviated, “p3 pro”) comprises the nucleic acid sequence set forth in SEQ ID NO: 10
- the “synthetic p2 promoter sequence” (abbreviated, “p2 pro”) comprises the nucleic acid sequence set forth in SEQ ID NO: 20.
- subtilis “aprE 5 ⁇ -untranslated region” (abbreviated, “aprE 5 ⁇ - UTR”) comprises the nucleic acid sequence set forth in SEQ ID NO: 11.
- a wild-type B. licheniformis “AmyL terminator sequence” (abbreviated, “AmyL term”) comprises the nucleic acid sequence set forth in SEQ ID NO: 12.
- a “first copy of the amylase-1 integration (expression) cassette with a modified ypuA signal sequence” (abbreviated, “1 st copy amy-1 cassette with mod-ypuAss”) comprises the polynucleotide (DNA) sequence set forth in SEQ ID NO: 7.
- the 1 st copy amy-1 cassette with mod-ypuAss comprises an upstream (5 ⁇ ) homology arm (up) for the B. licheniformis serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream DNA comprising a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B.
- subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-ypuAss (SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13).
- a “second copy of the amylase-1 integration (expression) cassette with a modified ypuA signal sequence” comprises the polynucleotide sequence set forth in SEQ ID NO: 17. More particularly, as schematically presented in FIG. 2C, the 2 nd copy amy-1 cassette with mod-ypuAss comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream DNA comprising a B.
- licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-ypuAss (SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21).
- a “first copy of the amylase-1 integration (expression) cassette with a modified AmyL signal sequence” (abbreviated, “1 st copy amy-1 cassette with mod-AmyLss”) comprises the polynucleotide (DNA) sequence set forth in SEQ ID NO: 30. More particularly, as schematically presented in FIG.2E, the 1 st copy amy-1 cassette with mod-AmyLss comprises an upstream (5 ⁇ ) homology arm (up) for the B. licheniformis serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream DNA comprising a B.
- licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-AmyLss (SEQ ID NO: 4) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13).
- a “second copy of the amylase-1 integration (expression) cassette with a modified AmyL signal sequence” (abbreviated, “2 nd copy amy-1 cassette with mod-AmyLss”) comprises the polynucleotide sequence set forth in SEQ ID NO: 30. More particularly, as schematically presented in FIG.2G, the 2 nd copy amy-1 cassette with mod-AmyLss comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream DNA comprising a B.
- a “first copy of the amylase-2 cassette with a modified AmyL signal sequence” (abbreviated, “1 st copy amy-2 cassette with mod-AmyLss”) comprises the polynucleotide (DNA) sequence set forth in SEQ ID NO: 39. More particularly, as schematically presented in FIG.3A, the 1 st copy amy-2 cassette with mod-AmyLss comprises a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B.
- subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-AmyLss (SEQ ID NO: 4) operably linked to DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12).
- a “second copy of the amylase-2 integration (expression) cassette with a native ypuA signal sequence” (abbreviated, “2 nd copy amy-2 cassette with ypuAss”) comprises the polynucleotide sequence set forth in SEQ ID NO: 38. More particularly, as schematically presented in FIG.
- the 2 nd copy amy-2 cassette with ypuAss comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the ypuAss (SEQ ID NO: 1) operably linked to DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B.
- a “second copy of the amylase-2 integration (expression) cassette with a modified AmyL signal sequence” (abbreviated, “2 nd copy amy-2 cassette with mod-AmyLss”) comprises the polynucleotide sequence set forth in SEQ ID NO: 40.
- the 2 nd copy amy-2 cassette with mod-AmyLss comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising the B.
- subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-ypuAss (SEQ ID NO: 2) operably linked to DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21).
- Amylase is meant to include any amylase such as glucoamylases, ⁇ -amylases, ⁇ -amylases and wild-type ⁇ -amylases of bacteria such as Bacillus sp., such as B. licheniformis and B. subtilis.
- Amylase shall mean an enzyme that is, among other things, capable of catalyzing the degradation of starch.
- Amylases are hydrolases that cleave the ⁇ -D-(1 ⁇ 4) O-glycosidic linkages in starch.
- ⁇ -amylases (EC 3.2.1.1; ⁇ -D-(1 ⁇ 4)-glucan glucanohydrolase) are defined as endo-acting enzymes cleaving ⁇ -D-(1 ⁇ 4) O-glycosidic linkages within the starch molecule in a random fashion.
- the exo-acting amylolytic enzymes such as ⁇ -amylases (EC 3.2.1.2; ⁇ - D-(1 ⁇ 4)-glucan maltohydrolase) and some product-specific amylases like maltogenic ⁇ -amylase (EC 3.2.1.133) cleave the starch molecule from the non-reducing end of the substrate.
- ⁇ -Amylases ⁇ - glucosidases (EC 3.2.1.20; ⁇ -D-glucoside glucohydrolase), glucoamylases (EC 3.2.1.3; ⁇ -D-(1 ⁇ 4)-glucan glucohydrolase), and product-specific amylases can produce malto-oligosaccharides of a specific length from starch.
- Amylase-1 abbreviated “Amy-1”
- Amylase-2 abbreviated “Amy-2”
- an “Amylase-1 reporter protein” comprises the mature amino acid sequence set forth in SEQ ID NO: 6.
- the Amy-1 protein (SEQ ID NO: 6) and functional variants thereof have been described in PCT Publication No. WO2008/112459 (incorporated herein by reference in its entirety).
- an “Amylase-2 reporter protein” comprises the amino acid sequence set forth in SEQ ID NO: 37.
- the Amy-2 protein (SEQ ID NO: 37) and functional variants thereof have been described in PCT Publication No. WO2023/023642.
- the polynucleotide (cassette) of SEQ ID NO: 46 comprises an upstream promoter sequence (p3; SEQ ID NO: 10) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (mod-ypuAss; SEQ ID NO: 2) encoding a modified ypuA signal sequence operably linked to DNA (amy-1; SEQ ID NO: 6) encoding an amylase reporter protein.
- the polynucleotide (cassette) of SEQ ID NO: 47 comprises an upstream promoter sequence (p2; SEQ ID NO: 20) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (mod-ypuAss; SEQ ID NO: 2) encoding a modified ypuA signal sequence operably linked to DNA (amy-1; SEQ ID NO: 6) encoding an amylase reporter protein.
- the polynucleotide (cassette) of SEQ ID NO: 48 comprises an upstream promoter sequence (p3; SEQ ID NO: 10) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (ypuAss; SEQ ID NO: 1) encoding a native ypuA signal sequence operably linked to DNA (amy-1; SEQ ID NO: 6) encoding an amylase reporter protein.
- the polynucleotide (cassette) of SEQ ID NO: 49 comprises an upstream promoter sequence (p2; SEQ ID NO: 20) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (ypuAss; SEQ ID NO: 1) encoding a native ypuA signal sequence operably linked to DNA (amy-1; SEQ ID NO: 6) encoding an amylase reporter protein.
- the polynucleotide (cassette) of SEQ ID NO: 50 comprises an upstream promoter sequence (p3; SEQ ID NO: 10) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (mod-ypuAss; SEQ ID NO: 2) encoding a modified ypuA signal sequence operably linked to DNA (amy-2; SEQ ID NO: 37) encoding an amylase reporter protein.
- the polynucleotide (cassette) of SEQ ID NO: 51 comprises an upstream promoter sequence (p2; SEQ ID NO: 20) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (mod-ypuAss; SEQ ID NO: 2) encoding a modified ypuA signal sequence operably linked to DNA (amy-2; SEQ ID NO: 37) encoding an amylase reporter protein.
- the polynucleotide (cassette) of SEQ ID NO: 52 comprises an upstream promoter sequence (p3; SEQ ID NO: 10) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (ypuAss; SEQ ID NO: 1) encoding a native ypuA signal sequence operably linked to DNA (amy-2; SEQ ID NO: 37) encoding an amylase reporter protein.
- the polynucleotide (cassette) of SEQ ID NO: 53 comprises an upstream promoter sequence (p2; SEQ ID NO: 20) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (ypuAss; SEQ ID NO: 1) encoding a native ypuA signal sequence operably linked to DNA (amy-2; SEQ ID NO: 37) encoding an amylase reporter protein.
- one or more B. licheniformis strains of the disclosure comprise one or more genetic modifications introduced therein. For instance, in certain embodiments, recombinant B.
- licheniformis (host) cells of the disclosure may comprise one or more genetic modifications (e.g., gene deletions, disruptions, etc.) of one or more endogenous (native) genes encoding unwanted/undesired background proteins (e.g., proteases and the like).
- a B. licheniformis strain named “BF1965” comprises deletions ( ⁇ ) of its endogenous serA and lysA genes (abbreviated, “ ⁇ serA/ ⁇ lysA”).
- ⁇ serA/ ⁇ lysA the B. licheniformis BF1965 ( ⁇ serA/ ⁇ lysA) strain may be constructed as generally described in PCT Publication No. WO2023/023642.
- a B. licheniformis strain named “AL223” was constructed from the BF1965 strain ( ⁇ serA/ ⁇ lysA) and comprises a 1 st copy of the amy-1 cassette (SEQ ID NO: 7) with the mod-ypuAss integrated at the serA locus and a 2 nd copy of the amy-1 cassette (SEQ ID NO: 17) with the mod-ypuAss integrated at the lysA locus.
- a B. licheniformis strain named “AL223” was constructed from the BF1965 strain ( ⁇ serA/ ⁇ lysA) and comprises a 1 st copy of the amy-1 cassette (SEQ ID NO: 7) with the mod-ypuAss integrated at the serA locus and a 2 nd copy of the amy-1 cassette (SEQ ID NO: 17) with the mod-ypuAss integrated at the lysA locus.
- a B. licheniformis strain named “BF780” was constructed from the BF1965 strain ( ⁇ serA/ ⁇ lysA) and comprises a 1 st copy of the amy-2 cassette (SEQ ID NO: 39) with mod-AmyLss integrated at the serA locus.
- a B. licheniformis strain named “WS2756” was constructed from the BF780 strain and comprises a 1 st copy of the amy-2 cassette (SEQ ID NO: 39) with mod-AmyLss integrated at the serA locus and a 2 nd copy of the amy-2 cassette (SEQ ID NO: 38) with native ypuAss integrated at the lysA locus.
- a B. licheniformis strain named “WS2756” was constructed from the BF780 strain and comprises a 1 st copy of the amy-2 cassette (SEQ ID NO: 39) with mod-AmyLss integrated at the serA locus and a 2 nd copy of the amy-2 cassette (SEQ ID NO: 38) with native ypuAss integrated at the lysA locus.
- licheniformis strain named “BF822” was constructed from the BF780 strain and comprises a 1 st copy of the amy-2 cassette (SEQ ID NO: 39) with mod-AmyLss integrated at the serA locus and a 2 nd copy amy-2 cassette (SEQ ID NO: 40) with mod-AmyLss integrated at the lysA locus.
- the genus “Bacillus” includes all species within the genus “Bacillus”’ as known to those of skill in the art, including but not limited to B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B.
- a “host cell” refers to a cell that has the capacity to act as a host or expression vehicle for a newly introduced DNA sequence.
- the host cells are Bacillus sp. cells or E. coli cells.
- a “modified cell” refers to a recombinant cell that comprises at least one genetic modification which is not present in the parent or control cell from which the modified cell was derived or constructed from.
- POI protein of interest
- a recombinant (modified) cell when the expression and/or production of a protein of interest (POI) in a recombinant (modified) cell is being compared to the expression and/or production of the same POI in a control or parent cell, it will be understood that the modified and control (parent) cells are grown/cultivated/fermented under the same conditions (e.g., the same conditions such as media, temperature, pH and the like).
- “increasing” protein production or “increased” protein production is meant an increased amount of protein produced (e.g., a protein of interest).
- the protein may be produced inside the host cell, or secreted (or transported) into the culture medium.
- the protein of interest is produced (secreted) into the culture medium.
- Increased protein production may be detected for example, as higher maximal level of protein or enzymatic activity (e.g., such as amylase activity), or total extracellular protein produced as compared to the parental host cell.
- modification and “genetic modification” are used interchangeably and include, but are not limited to: (a) the introduction, substitution, or removal of one or more nucleotides in a gene (or an ORF/CDS thereof), or the introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of the gene or ORF thereof, (b) a gene disruption, (c) a gene conversion, (d) a gene deletion, (e) the down-regulation of a gene, (f) specific mutagenesis and/or (g) random mutagenesis of any one or more the genes disclosed herein.
- introducing includes methods known in the art for introducing polynucleotides into a cell, including, but not limited to protoplast fusion, natural or artificial transformation (e.g., calcium chloride, electroporation), transduction, transfection, conjugation and the like (e.g., see Ferrari et al., 1989).
- transformation e.g., calcium chloride, electroporation
- transduction e.g., transfection, conjugation and the like
- transfection e.g., see Ferrari et al., 1989.
- transformed or “transformation” mean a cell has been transformed by use of recombinant DNA techniques.
- Transformation typically occurs by insertion of one or more nucleotide sequences (e.g., a polynucleotide, an ORF or gene) into a cell.
- the inserted nucleotide sequence may be a heterologous nucleotide sequence (i.e., a sequence that is not naturally occurring in cell that is to be transformed). Transformation therefore generally refers to introducing an exogenous DNA into a host cell so that the DNA is maintained as a chromosomal integrant or a self-replicating extra-chromosomal vector.
- transforming DNA “transforming sequence”, and “DNA construct” refer to DNA that is used to introduce sequences into a host cell or organism.
- Transforming DNA is DNA used to introduce sequences into a host cell or organism.
- the DNA may be generated in vitro by PCR or any other suitable techniques.
- the transforming DNA comprises an incoming sequence, while in other embodiments it further comprises an incoming sequence flanked by homology boxes.
- the transforming DNA comprises other non-homologous sequences, added to the ends (i.e., stuffer sequences or flanks). The ends can be closed such that the transforming DNA forms a closed circle, such as, for example, insertion into a vector.
- a gene disruption includes, but is not limited to, frameshift mutations, premature stop codons (i.e., such that a functional protein is not made), substitutions eliminating or reducing activity of the protein internal deletions (such that a functional protein is not made), insertions disrupting the coding sequence, mutations removing the operable link between a native promoter required for transcription and the open reading frame, and the like.
- an incoming sequence refers to a DNA sequence that is introduced into the Bacillus sp. chromosome. In some embodiments, the incoming sequence is part of a DNA construct. In other embodiments, the incoming sequence encodes one or more proteins of interest. In some embodiments, the incoming sequence comprises a sequence that may or may not already be present in the genome of the cell to be transformed (i.e., it may be either a homologous or heterologous sequence). In some embodiments, the incoming sequence encodes one or more proteins of interest, a gene, and/or a mutated or modified gene.
- the incoming sequence encodes a functional wild- type gene or operon, a functional mutant gene or operon, or a nonfunctional gene or operon.
- the non-functional sequence may be inserted into a gene to disrupt function of the gene.
- the incoming sequence includes a selective marker.
- the incoming sequence includes two homology boxes. [0135] As used herein, “homology box” refers to a nucleic acid sequence, which is homologous to a sequence in the Bacillus chromosome.
- a homology box is an upstream or downstream region having between about 80 and 100% sequence identity, between about 90 and 100% sequence identity, or between about 95 and 100% sequence identity with the immediate flanking coding region of a gene or part of a gene to be deleted, disrupted, inactivated, down-regulated and the like, according to the invention. These sequences direct where in the Bacillus chromosome a DNA construct is integrated and directs what part of the Bacillus chromosome is replaced by the incoming sequence. While not meant to limit the present disclosure, a homology box may include about between 1 base pair (bp) to 200 kilobases (kb).
- a homology box includes about between 1 bp and 10.0 kb; between 1 bp and 5.0 kb; between 1 bp and 2.5 kb; between 1 bp and 1.0 kb, and between 0.25 kb and 2.5 kb.
- a homology box may also include about 10.0 kb, 5.0 kb, 2.5 kb, 2.0 kb, 1.5 kb, 1.0 kb, 0.5 kb, 0.25 kb and 0.1 kb.
- the 5' and 3' ends of a selective marker are flanked by a homology box wherein the homology box comprises nucleic acid sequences immediately flanking the coding region of the gene.
- a host cell “genome”, a bacterial (host) cell “genome”, or a Bacillus sp. (host) cell “genome” includes chromosomal and extrachromosomal genes.
- plasmid vector
- cassette refer to extrachromosomal elements, often carrying genes which are typically not part of the central metabolism of the cell, and usually in the form of circular double-stranded DNA molecules.
- Such elements may be autonomously replicating sequences, genome integrating sequences, phage or nucleotide sequences, linear or circular, of a single- stranded or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing a promoter fragment and DNA sequence for a selected gene product along with appropriate 3' untranslated sequence into a cell.
- plasmid refers to a circular double-stranded (ds) DNA construct used as a cloning vector, and which forms an extrachromosomal self-replicating genetic element in many bacteria and some eukaryotes.
- plasmids become incorporated into the genome of the host cell.
- plasmids exist in a parental cell and are lost in the daughter cell.
- a “transformation cassette” refers to a specific vector comprising a gene (or ORF thereof), and having elements in addition to the foreign gene that facilitate transformation of a particular host cell.
- the term “vector” refers to any nucleic acid that can be replicated (propagated) in cells and can carry new genes or DNA segments into cells. Thus, the term refers to a nucleic acid construct designed for transfer between different host cells.
- Vectors include viruses, bacteriophage, pro-viruses, plasmids, phagemids, transposons, and artificial chromosomes such as YACs (yeast artificial chromosomes), BACs (bacterial artificial chromosomes), PLACs (plant artificial chromosomes), and the like, that are “episomes” (i.e., replicate autonomously or can integrate into a chromosome of a host organism).
- An “expression vector” refers to a vector that has the ability to incorporate and express heterologous DNA in a cell. Many prokaryotic and eukaryotic expression vectors are commercially available and know to one skilled in the art.
- expression cassette and “expression vector” refer to a nucleic acid construct generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular nucleic acid in a target cell (i.e., these are vectors or vector elements, as described above).
- the recombinant expression cassette can be incorporated into a plasmid, chromosome, mitochondrial DNA, plastid DNA, virus, or nucleic acid fragment.
- the recombinant expression cassette portion of an expression vector includes, among other sequences, a nucleic acid sequence to be transcribed and a promoter.
- DNA constructs also include a series of specified nucleic acid elements that permit transcription of a particular nucleic acid in a target cell.
- a DNA construct of the disclosure comprises a selective marker and an inactivating chromosomal or gene or DNA segment as defined herein.
- a “targeting vector” is a vector that includes polynucleotide sequences that are homologous to a region in the chromosome of a host cell into which the targeting vector is transformed and that can drive homologous recombination at that region. For example, targeting vectors find use in introducing mutations into the chromosome of a host cell through homologous recombination.
- the targeting vector comprises other non-homologous sequences, e.g., added to the ends (i.e., stuffer sequences or flanking sequences).
- the ends can be closed such that the targeting vector forms a closed circle, such as, for example, insertion into a vector.
- a parental B. licheniformis (host) cell is modified (e.g., transformed) by introducing therein one or more “targeting vectors”.
- protein of interest or “POI” refers to a polypeptide of interest that is desired to be expressed in a modified B.
- a POI may be an enzyme, a substrate-binding protein, a surface-active protein, a structural protein, a receptor protein, and the like.
- a modified cell of the disclosure produces an increased amount of a heterologous protein of interest relative to a control (or parent) cell.
- an increased amount of a protein of interest produced by a modified cell of the disclosure is at least a 0.5% increase, at least a 1.0% increase, at least a 5.0% increase, or a greater than 5.0% increase, relative to the control (or parent) cell.
- a “gene of interest” or “GOI” refers a nucleic acid sequence (e.g., a polynucleotide, a gene or an ORF) which encodes a POI.
- a “gene of interest” encoding a “protein of interest” may be a naturally occurring gene, a mutated gene or a synthetic gene.
- the terms “polypeptide” and “protein” are used interchangeably, and refer to polymers of any length comprising amino acid residues linked by peptide bonds. The conventional one (1) letter or three (3) letter codes for amino acid residues are used herein.
- polypeptide may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids.
- polypeptide also encompasses an amino acid polymer that has been modified naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component.
- polypeptides containing one or more analogs of an amino acid including, for example, unnatural amino acids, etc.
- a gene of the instant disclosure encodes a commercially relevant industrial protein of interest, such as an enzyme (e.g., a acetyl esterases, aminopeptidases, amylases, arabinases, arabinofuranosidases, carbonic anhydrases, carboxypeptidases, catalases, cellulases, chitinases, chymosins, cutinases, deoxyribonucleases, epimerases, esterases, ⁇ -galactosidases, ⁇ -galactosidases, ⁇ -glucanases, glucan lysases, endo- ⁇ -glucanases, glucoamylases, glucose oxidases, ⁇ - glucosidases, ⁇ -glucosidases, glucuronidases, glycosyl hydrolases, hemicellulases, hexose oxidases, an enzyme (e.g.
- a “variant” polypeptide refers to a polypeptide that is derived from a parent (or reference) polypeptide by the substitution, addition, or deletion of one or more amino acids, typically by recombinant DNA techniques. Variant polypeptides may differ from a parent polypeptide by a small number of amino acid residues and may be defined by their level of primary amino acid sequence homology/identity with a parent (reference) polypeptide.
- variant polypeptides have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% amino acid sequence identity with a parent (reference) polypeptide sequence.
- a “variant” polynucleotide refers to a polynucleotide encoding a variant polypeptide, wherein the “variant polynucleotide” has a specified degree of sequence homology/identity with a parent polynucleotide, or hybridizes with a parent polynucleotide (or a complement thereof) under stringent hybridization conditions.
- a variant polynucleotide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% nucleotide sequence identity with a parent (reference) polynucleotide sequence.
- a “mutation” refers to any change or alteration in a nucleic acid sequence. Several types of mutations exist, including point mutations, deletion mutations, silent mutations, frame shift mutations, splicing mutations and the like.
- Mutations may be performed specifically (e.g., via site directed mutagenesis) or randomly (e.g., via chemical agents, passage through repair minus bacterial strains).
- substitution means the replacement (i.e., substitution) of one amino acid with another amino acid.
- homologous polynucleotides or polypeptides relates to homologous polynucleotides or polypeptides.
- homologous polynucleotides or polypeptides have a “degree of identity” of at least 60%, more preferably at least 70%, even more preferably at least 85%, still more preferably at least 90%, more preferably at least 95%, and most preferably at least 98%.
- the term “percent (%) identity” refers to the level of nucleic acid or amino acid sequence identity between the nucleic acid sequences that encode a polypeptide or the polypeptide's amino acid sequences, when aligned using a sequence alignment program.
- “specific productivity” is total amount of protein produced per cell per time over a given time period.
- the terms “purified”, “isolated” or “enriched” are meant that a biomolecule (e.g., a polypeptide or polynucleotide) is altered from its natural state by virtue of separating it from some, or all of, the naturally occurring constituents with which it is associated in nature.
- isolation or purification may be accomplished by art-recognized separation techniques such as ion exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, ammonium sulphate precipitation or other protein salt precipitation, centrifugation, size exclusion chromatography, filtration, microfiltration, gel electrophoresis or separation on a gradient to remove whole cells, cell debris, impurities, extraneous proteins, or enzymes undesired in the final composition. It is further possible to then add constituents to a purified or isolated biomolecule composition which provide additional benefits, for example, activating agents, anti-inhibition agents, desirable ions, compounds to control pH or other enzymes or chemicals. II.
- SIGNAL SEQUENCES FOR IMPROVED PROTEIN SECRETION [0156] Generally accepted methods of secreting heterologous proteins often rely on the use of native (protein) signal sequences of similar proteins (e.g., signal sequences from AmyL or AmyE for amylases, or AprE, NprE, or AprL for proteases), or signal sequences that are native to the heterologous protein sequence.
- protein translation, secretion and folding can be consecutive processes and/or concurrent processes, wherein the protein’s signal (secretion) sequence plays a dynamic role in all three processes.
- sub-optimal signal sequences can present particular problems, including among other things, poor or insufficient translation, secretion and/or folding of the heterologous protein, wherein non-optimal pairings of signal sequences to mature protein sequences can lead to bottlenecks in production and secretion of the protein, leading to misfolded/inactive product and/or the induction of cellular stress responses further decreasing the productivity of the host cell.
- Applicant screened B. licheniformis putative signal peptide sequences that were selected by using the SignalP peptide prediction algorithm (Armenteros et al., 2019), and have surprisingly identified a native B.
- licheniformis ypuA protein signal sequence (SEQ ID NO: 1) that is particularly useful for enhancing/increasing production (secretion) of heterologous proteins of interest. More particularly, as generally set forth below in the Examples, Applicant constructed recombinant B. licheniformis strains expressing exemplary reporter proteins (i.e., heterologous proteins of interest), wherein the mature (amino acid) sequences of the reporter proteins (e.g., Amy-1 reporter, Amy-2 reporter) were operably linked to upstream (N-terminal) signal peptide (secretion) sequences of the disclosure, such as the native B. licheniformis ypuA signal sequence (ypuAss) of SEQ ID NO: 1.
- exemplary reporter proteins i.e., heterologous proteins of interest
- the mature (amino acid) sequences of the reporter proteins e.g., Amy-1 reporter, Amy-2 reporter
- Applicant constructed a modified (variant) ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) which is particularly useful for enhancing/increasing production (secretion) of heterologous proteins of interest.
- the mod-ypuAss comprises a substitution of the aspartic acid (D24) to a serine (Ser; S24) residue at position 24, see FIG.1A.
- D24 aspartic acid
- Ser serine residues
- Applicant constructed a first (1 st ) copy of an amylase-1 integration cassette (1 st copy amy-1 cassette) and second (2 nd ) copy of an amylase-1 integration cassette (2 nd copy amy- 1 cassette) suitable for integration into a B. licheniformis host cell.
- the 1 st copy of the amy-1 cassette (FIG. 2A, SEQ ID NO: 7) with a modified ypuA signal sequence (mod- ypuAss; SEQ ID NO: 2) comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream DNA comprising a B.
- licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding a modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B.
- the mod-ypuAss (SEQ ID NO: 2) as compared to the native ypuA signal sequence (ypuAss; SEQ ID NO: 1) comprises a substitution of an aspartic acid (Asp; D) to a serine (Ser; S) residue at position 24 of SEQ ID NO: 2.
- a 2 nd copy of the amy-1 cassette (FIG.2C, SEQ ID NO: 17) with the modified ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to the lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (pro; SEQ ID NO: 20) operably linked to DNA encoding the B. subtilis aprE 5′UTR (SEQ ID NO: 11) operably linked to DNA encoding the modified B.
- licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21).
- myL term SEQ ID NO: 12
- licheniformis host strain named “AL223” was constructed by integrating the 1 st and 2 nd copies of the amy-1 expression cassettes into a parental B. licheniformis strain named “BF1965”, which BF1965 parent strain comprises deletions ( ⁇ ) of the wild-type serA and lysA genes ( ⁇ serA/ ⁇ lysA).
- a control strain (AL207) with the modified B was tested to test relative performance of the mod-ypuAss (SEQ ID NO: 2) on mature Amy-1 production.
- licheniformis AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) was constructed in the BF1965 strain by integration of a 1 st amy-1 cassette (SEQ ID NO: 30) at the serA locus and integration of a 2 nd amy-1 cassette (SEQ ID NO: 31) at the lysA locus. More particularly, the 1 st amy-1 cassette with mod-AmyLss (SEQ ID NO: 30) and the 2 nd amy-1 cassette with mod-AmyLss (SEQ ID NO: 31) are shown in FIG. 2E and FIG. 2G, respectively.
- the native AmyL signal sequence (AmyLss; SEQ ID NO: 3) is generally known to be an optimal signal sequence for producing/secreting many amylase proteins.
- the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) has been shown to further enhance the production/secretion of amylase proteins as compared to the native AmyLss (SEQ ID NO: 3), as generally described in PCT Publication No.2023/023642.
- the mod-AmyLss (SEQ ID NO: 4) comprises a serine (S28) at the -2 position (Tjalsma et al., 2000) relative to the native AmyLss (SEQ ID NO: 3), see FIG.
- the B. licheniformis AL223 strain with the mod-ypuAss was assayed for production of the Amy-1 reporter and compared to the B. licheniformis control AL207 strain comprising two (2) integrated copies of the amylase-1 expression cassettes with the modified (control) AmyL signal sequence (mod-AmyLss control; SEQ ID NO: 4) using standard small-scale conditions.
- the Amy-1 protein production was quantified using the method of Bradford assay, wherein the relative improvement in production of Amy-1 from the AL223 strain is compared to the production of the Amy-1 from the AL207 (control) strain.
- the mod-ypuAss (SEQ ID NO: 2) demonstrates an approximately 50% increase in Amy-1 protein production (strain AL223) relative to Amy-1 production (strain AL207) using the mod-AmyLss (SEQ ID NO: 4).
- Applicant constructed certain amylase-2 integration (expression) cassettes suitable for integration into a B. licheniformis host cell. More particularly, as described in Example 5 and shown in FIG. 3, the second (2 nd ) copy of an amylase-2 expression cassette (FIG.
- SEQ ID NO: 38 with a native ypuA signal sequence (ypuAss; SEQ ID NO: 1) comprises an upstream (5′) homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to a downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the native B.
- licheniformis YpuA signal sequence (ypuAss; SEQ ID NO: 1) operably linked to DNA encoding the Amy-2 protein (SEQ ID NO: 37) operably linked to a B. licheniformis amyL transcriptional terminator (AmyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21).
- MyL term SEQ ID NO: 12
- Example 6 of the disclosure describes the construction of B. licheniformis host cells (strains) expressing two (2) copies of the mature Amy-2 (SEQ ID NO: 37) reporter protein.
- the 2 nd copy amylase-2 expression cassette (SEQ ID NO: 38) constructed in Example 5 was integrated into the lysA locus of the B. licheniformis strain named “BF780”, resulting in the modified B. licheniformis host strain named “WS2756” (Example 6).
- the BF780 strain was constructed from the parental B. licheniformis BF1965 strain ( ⁇ serA/ ⁇ lysA), wherein the BF1965 strain comprises an introduced first (1 st ) copy of an amylase-2 cassette with the mod-AmyLss (FIG. 3A; SEQ ID NO: 39) integrated into the serA locus.
- licheniformis strain WS2756 comprises 2-copies of the amy-2 integration (expression) cassettes, wherein the 1 st copy amy-2 cassette integrated at the serA locus (FIG. 3A; SEQ ID NO: 39) and having the mod-AmyLss (SEQ ID NO: 4) and the 2 nd copy amy-2 cassette integrated into the lysA locus (FIG.3C; SEQ ID NO: 38) and having the native ypuAss (SEQ ID NO: 1).
- the BF822 control strain described in Example 6 comprises 2-copies of the amy-2 integration (expression) cassettes, wherein the 1 st copy amy-2 cassette (FIG.3A; SEQ ID NO: 39) is integrated at the serA locus and having the mod- AmyLss (SEQ ID NO: 4) and the 2 nd copy amy-2 cassette (FIG.3E; SEQ ID NO: 40) is integrated into the lysA locus and having the mod-AmyLss (SEQ ID NO: 4).
- the WS2756 strain and BF822 (control) strain constructed in Example 6 were assayed for production of the Amy-2 (SEQ ID NO: 37) reporter using standard small- scale conditions.
- the WS2756 strain (with native ypuAss; SEQ ID NO: 1) demonstrates an approximately 18% increase in Amy-2 reporter production as compared to the control BF822 strain (with mod-AmyLss; SEQ ID NO: 4).
- certain embodiments of the disclosure provide, inter alia, nucleic acids, polynucleotides, vectors, expression cassettes, regulatory elements, and the like, suitable for use in constructing recombinant (modified) Bacillus host cells.
- polynucleotides e.g., expression cassettes
- an upstream (5′) promoter (pro) sequence operably linked to a downstream nucleic acid sequence (ss) encoding a protein signal (secretion) sequence operably linked to a downstream nucleic acid sequence (poi) encoding a protein of interest.
- Scheme 1 a generic polynucleotide sequence encoding an amino (N) terminal signal sequence in operable combination with a mature protein of interest (POI) is shown below in Scheme 1: Scheme 1: 5′-[ss]-[poi]-3′ wherein the nucleic acid (ss) sequence encoding the N-terminal signal sequence (SS) is upstream and operably linked to a downstream nucleic acid (poi) sequence encoding a mature protein of interest (POI).
- polynucleotide expression cassettes may be described generically as shown in Scheme 2: Scheme 2: 5′-[pro]-[ss]-[poi]-3′ wherein the promoter (pro) sequence is upstream (5′) and operably linked to a nucleic acid (ss) sequence encoding the N-terminal signal sequence (SS), which is upstream and operably linked to a nucleic acid (poi) sequence encoding a protein of interest (POI).
- the polynucleotide may further comprise a terminator (term) sequence downstream and operably linked to the nucleic acid (poi) sequence encoding the mature POI.
- the disclosure relates to polynucleotide constructs (e.g., Scheme 2) encoding a heterologous POI (e.g., a mature amylase protein sequence), wherein the nucleic acid (DNA) encoding the POI is operably linked to an upstream nucleic acid encoding an N-terminal ypuA signal sequence (ypuAss; SEQ ID NO: 1) and/or operably linked to an upstream nucleic acid encoding an N- terminal modified (variant) ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2).
- ypuAss N-terminal ypuA signal sequence
- mod-ypuAss N-terminal ypuA signal sequence
- certain other embodiments of the disclosure provide recombinant (modified) B.
- licheniformis strains/cells comprising one or more introduced polynucleotide constructs (e.g., expression cassettes) encoding one or more mature amylases comprising an N-terminal ypuA signal sequence (SED ID NO: 1) and/or encoding one or more mature amylases comprising an N-terminal modified (variant) ypuA signal sequence (SED ID NO: 2).
- SED ID NO: 1 the native B. licheniformis ypuA signal sequence
- FIG. 1 the native B. licheniformis ypuA signal sequence
- ypuAss comprises twenty-five (25) amino acid residues with an aspartic acid (Asp; D24) residue at position 24, whereas the modified (variant) ypuA signal sequence (FIG. 1B, mod-YpuAss; SED IDNO: 2) comprises a substitution of the aspartic acid (D24) to a serine (Ser; S24) residue at position 24.
- the amino acid positions of a particular protein signal sequence may be described and numbered from the amino-terminus (NH 2 ), as indicated in FIG. 1.
- the amino acid positions may be described and numbered according to the cleavage site of a particular signal sequence.
- the most C-terminal amino acid position of the ypuAss (SEQ ID NO: 1, residue A 25 ) can be designated with a negative 1 (-1) amino acid (position), the amino acid position to its left a negative 2 (-2), etc..
- other native and/or modified signal sequences of the disclosure may be designated with similar specificity.
- certain other embodiments of the disclosure provide recombinant (modified) B. licheniformis cells (strains) capable of secreting enhanced amounts of amylase proteins. III.
- certain embodiments of the disclosure are related to recombinant (modified) Bacillus cells capable of producing increased amounts of heterologous proteins of interest. Certain embodiments are therefore related to methods for constructing such recombinant Bacillus cells having increased protein production capabilities.
- one or more expression cassettes encoding a protein of intertest are introduced into Bacillus cells of the disclosure.
- the cassettes are integrated into the genome of the cell.
- expression cassettes encoding a protein of interest were integrated into the lysA locus and the serA locus of a parental B.
- Bacillus cells of the disclosure are rendered deficient in the production of one or more native (endogenous) genes.
- Bacillus cells of the disclosure are rendered deficient in the production of one or more native (endogenous) proteases.
- a host cell of the disclosure is a Bacillus licheniformis cell deficient in the production of one or more native proteases selected from the group consisting of wprA, nprE, mpr, aprL, bprE, htrA, vpr and ispA.
- recombinant cells of the disclosure may be constructed by one of skill using standard and routine recombinant DNA and molecular cloning techniques well known in the art.
- Methods for genetically modifying cells include, but are not limited to, (a) the introduction, substitution, or removal of one or more nucleotides in a gene, or the introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of the gene, (b) a gene disruption, (c) a gene conversion, (d) a gene deletion, (e) a gene down-regulation, (f) site specific mutagenesis and/or (g) random mutagenesis.
- modified cells of the disclosure may be constructed by reducing or eliminating the expression of a gene, using methods well known in the art, for example, insertions, disruptions, replacements, or deletions.
- the portion of the gene to be modified or inactivated may be, for example, the coding region or a regulatory element required for expression of the coding region.
- An example of such a regulatory or control sequence may be a promoter sequence or a functional part thereof, (i.e., a part which is sufficient for affecting expression of the nucleic acid sequence).
- Other control sequences for modification include, but are not limited to, a leader sequence, a pro-peptide sequence, a signal sequence, a transcription terminator, a transcriptional activator and the like.
- a modified cell is constructed by gene deletion to eliminate or reduce the expression of the gene.
- Gene deletion techniques enable the partial or complete removal of the gene(s), thereby eliminating their expression, or expressing a non-functional (or reduced activity) protein product.
- the deletion of the gene(s) may be accomplished by homologous recombination using a plasmid that has been constructed to contiguously contain the 5' and 3' regions flanking the gene.
- the contiguous 5' and 3' regions may be introduced into a Bacillus cell, for example, on a temperature-sensitive plasmid, such as pE194, in association with a second selectable marker at a permissive temperature to allow the plasmid to become established in the cell.
- the cell is then shifted to a non-permissive temperature to select for cells that have the plasmid integrated into the chromosome at one of the homologous flanking regions. Selection for integration of the plasmid is affected by selection for the second selectable marker. After integration, a recombination event at the second homologous flanking region is stimulated by shifting the cells to the permissive temperature for several generations without selection. The cells are plated to obtain single colonies and the colonies are examined for loss of both selectable markers.
- a person of skill in the art may readily identify nucleotide regions in the gene’s coding sequence and/or the gene’s non- coding sequence suitable for complete or partial deletion.
- a modified cell is constructed by introducing, substituting, or removing one or more nucleotides in the gene or a regulatory element required for the transcription or translation thereof.
- nucleotides may be inserted or removed so as to result in the introduction of a stop codon, the removal of the start codon, or a frame-shift of the open reading frame.
- Such a modification may be accomplished by site-directed mutagenesis or PCR generated mutagenesis in accordance with methods known in the art.
- a gene of the disclosure is inactivated by complete or partial deletion.
- a modified cell is constructed by the process of gene conversion.
- a nucleic acid sequence corresponding to the gene(s) is mutagenized in vitro to produce a defective nucleic acid sequence, which is then transformed into the parental Bacillus cell to produce a defective gene.
- the defective nucleic acid sequence replaces the endogenous gene.
- the defective gene or gene fragment also encodes a marker which may be used for selection of transformants containing the defective gene.
- the defective gene may be introduced on a non-replicating or temperature-sensitive plasmid in association with a selectable marker. Selection for integration of the plasmid is affected by selection for the marker under conditions not permitting plasmid replication.
- a modified cell is constructed by established anti-sense techniques using a nucleotide sequence complementary to the nucleic acid sequence of the gene. More specifically, expression of the gene by a Bacillus cell may be reduced (down-regulated) or eliminated by introducing a nucleotide sequence complementary to the nucleic acid sequence of the gene, which may be transcribed in the cell and is capable of hybridizing to the mRNA produced in the cell.
- RNA interference RNA interference
- siRNA small interfering RNA
- miRNA microRNA
- antisense oligonucleotides and the like, all of which are well known to the skilled artisan.
- a modified cell is produced/constructed via CRISPR-Cas9 editing.
- a gene encoding a protein of interest can be edited or disrupted (or deleted or down-regulated) by means of nucleic acid guided endonucleases, that find their target DNA by binding either a guide RNA (e.g., Cas9) and Cpf1 or a guide DNA (e.g., NgAgo), which recruits the endonuclease to the target sequence on the DNA, wherein the endonuclease can generate a single or double stranded break in the DNA.
- a guide RNA e.g., Cas9 and Cpf1 or a guide DNA (e.g., NgAgo)
- This targeted DNA break becomes a substrate for DNA repair, and can recombine with a provided editing template to disrupt or delete the gene.
- the gene encoding the nucleic acid guided endonuclease (for this purpose Cas9 from S. pyogenes), or a codon optimized gene encoding the Cas9 nuclease is operably linked to a promoter active in the Bacillus cell and a terminator active in Bacillus cell, thereby creating a Bacillus Cas9 expression cassette.
- a promoter active in the Bacillus cell and a terminator active in Bacillus cell, thereby creating a Bacillus Cas9 expression cassette.
- target sites unique to the gene of interest are readily identified by a person skilled in the art.
- variable targeting domain will comprise nucleotides of the target site which are 5′ of the (PAM) proto-spacer adjacent motif (TGG), which nucleotides are fused to DNA encoding the Cas9 endonuclease recognition domain for S. pyogenes Cas9 (CER).
- PAM proto-spacer adjacent motif
- CER Cas9 endonuclease recognition domain for S. pyogenes Cas9
- a Bacillus expression cassette for the gRNA is created by operably linking the DNA encoding the gRNA to a promoter active in Bacillus cells and a terminator active in Bacillus cells.
- the DNA break induced by the endonuclease is repaired/replaced with an incoming sequence.
- a nucleotide editing template is provided, such that the DNA repair machinery of the cell can utilize the editing template.
- about 500bp 5′ of targeted gene can be fused to about 500bp 3′ of the targeted gene to generate an editing template, which template is used by the Bacillus host’s machinery to repair the DNA break generated by the RGEN.
- the Cas9 expression cassette, the gRNA expression cassette and the editing template can be co- delivered to filamentous fungal cells using many different methods (e.g., protoplast fusion, electroporation, natural competence, or induced competence).
- the transformed cells are screened by PCR amplifying the target gene locus, by amplifying the locus with a forward and reverse primer. These primers can amplify the wild-type locus or the modified locus that has been edited by the RGEN.
- a modified cell is constructed by random or specific mutagenesis using methods well known in the art, including, but not limited to, chemical mutagenesis and transposition. Modification of the gene may be performed by subjecting the parental cell to mutagenesis and screening for mutant cells in which expression of the gene has been reduced or eliminated.
- the mutagenesis which may be specific or random, may be performed, for example, by use of a suitable physical or chemical mutagenizing agent, use of a suitable oligonucleotide, or subjecting the DNA sequence to PCR generated mutagenesis.
- the mutagenesis may be performed by use of any combination of these mutagenizing methods.
- Examples of a physical or chemical mutagenizing agent suitable for the present purpose include ultraviolet (UV) irradiation, hydroxylamine, N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), N-methyl- N'-nitrosoguanidine (NTG), O-methyl hydroxylamine, nitrous acid, ethyl methane sulphonate (EMS), sodium bisulphite, formic acid, and nucleotide analogues.
- UV ultraviolet
- MNNG N-methyl-N'-nitro-N-nitrosoguanidine
- NTG N-methyl- N'-nitrosoguanidine
- EMS ethyl methane sulphonate
- sodium bisulphite formic acid
- nucleotide analogues examples of mutagenesis is typically performed by incubating the parental cell to be mutagenized in the presence of the mutagenizing agent of choice under suitable conditions, and selecting for mutant cells exhibiting reduced or no expression of the gene.
- WO2003/083125 discloses methods for modifying Bacillus cells, such as the creation of Bacillus deletion strains and DNA constructs using PCR fusion to bypass E. coli.
- PCT Publication No. WO2002/14490 discloses methods for modifying Bacillus cells including (1) the construction and transformation of an integrative plasmid (pComK), (2) random mutagenesis of coding sequences, signal sequences and pro-peptide sequences, (3) homologous recombination, (4) increasing transformation efficiency by adding non-homologous flanks to the transformation DNA, (5) optimizing double cross-over integrations, (6) site directed mutagenesis and (7) marker-less deletion.
- pComK integrative plasmid
- bacterial cells e.g., E. coli and Bacillus
- transformation including protoplast transformation and congression, transduction, and protoplast fusion are known and suited for use in the present disclosure.
- Methods of transformation are particularly preferred to introduce a DNA construct of the present disclosure into a host cell.
- host cells are directly transformed (i.e., an intermediate cell is not used to amplify, or otherwise process, the DNA construct prior to introduction into the host cell).
- DNA constructs are co-transformed with a plasmid without being inserted into the plasmid.
- a selective marker is deleted or substantially excised from the modified Bacillus strain by methods known in the art.
- resolution of the vector from a host chromosome leaves the flanking regions in the chromosome, while removing the indigenous chromosomal region.
- Promoters and promoter sequence regions for use in the expression of genes, open reading frames (ORFs) thereof and/or variant sequences thereof in Bacillus cells are generally known on one of skill in the art.
- Promoter sequences of the disclosure are generally chosen so that they are functional in the Bacillus cells (e.g., B. licheniformis cells, B. subtilis cells and the like).
- promoters useful for driving gene expression in Bacillus cells include, but are not limited to, the B. subtilis alkaline protease (aprE) promoter, the ⁇ -amylase promoter (amyE) of B. subtilis, the ⁇ -amylase promoter (amyL) of B.
- certain embodiments are related to methods of producing proteins of interest in Bacillus cells by fermenting the cells in a suitable medium. Fermentation methods well known in the art can be applied to ferment Bacillus cells of the disclosure.
- the cells are cultured under batch or continuous fermentation conditions.
- a classical batch fermentation is a closed system, where the composition of the medium is set at the beginning of the fermentation and is not altered during the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism(s). In this method, fermentation is permitted to occur without the addition of any components to the system.
- a batch fermentation qualifies as a “batch” with respect to the addition of the carbon source, and attempts are often made to control factors such as pH and oxygen concentration.
- the metabolite and biomass compositions of the batch system change constantly up to the time the fermentation is stopped.
- cells can progress through a static lag phase to a high growth log phase, and finally to a stationary phase, where growth rate is diminished or halted. If untreated, cells in the stationary phase eventually die.
- cells in log phase are responsible for the bulk of production of product.
- a suitable variation on the standard batch system is the “fed-batch” fermentation system.
- the substrate is added in increments as the fermentation progresses.
- Fed-batch systems are useful when catabolite repression likely inhibits the metabolism of the cells and where it is desirable to have limited amounts of substrate in the medium.
- Continuous fermentation is an open system where a defined fermentation medium is added continuously to a bioreactor, and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the cultures at a constant high density, where cells are primarily in log phase growth. Continuous fermentation allows for the modulation of one or more factors that affect cell growth and/or product concentration. For example, in one embodiment, a limiting nutrient, such as the carbon source or nitrogen source, is maintained at a fixed rate and all other parameters are allowed to moderate.
- a limiting nutrient such as the carbon source or nitrogen source
- a number of factors affecting growth can be altered continuously while the cell concentration, measured by media turbidity, is kept constant. Continuous systems strive to maintain steady state growth conditions. Thus, cell loss due to medium being drawn off should be balanced against the cell growth rate in the fermentation. Methods of modulating nutrients and growth factors for continuous fermentation processes, as well as techniques for maximizing the rate of product formation, are well known in the art of industrial microbiology.
- a protein of interest expressed/produced by a Bacillus cell of the disclosure may be recovered from the culture medium by conventional procedures including separating the host cells from the medium by centrifugation or filtration, or if necessary, disrupting the cells and removing the supernatant from the cellular fraction and debris.
- the proteinaceous components of the supernatant or filtrate are precipitated by means of a salt, e.g., ammonium sulfate.
- the precipitated proteins are then solubilized and may be purified by a variety of chromatographic procedures, e.g., ion exchange chromatography, gel filtration.
- the cells are cultured under batch or continuous fermentation conditions.
- a classical batch fermentation is a closed system, where the composition of the medium is set at the beginning of the fermentation and is not altered during the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism(s). In this method, fermentation is permitted to occur without the addition of any components to the system.
- a batch fermentation qualifies as a “batch” with respect to the addition of the carbon source, and attempts are often made to control factors such as pH and oxygen concentration.
- the metabolite and biomass compositions of the batch system change constantly up to the time the fermentation is stopped.
- cells can progress through a static lag phase to a high growth log phase, and finally to a stationary phase, where growth rate is diminished or halted. If untreated, cells in the stationary phase eventually die. In general, cells in log phase are responsible for the bulk of production of product.
- a suitable variation on the standard batch system is the “fed-batch” fermentation system. In this variation of a typical batch system, the substrate is added in increments as the fermentation progresses.
- Fed-batch systems are useful when catabolite repression likely inhibits the metabolism of the cells and where it is desirable to have limited amounts of substrate in the medium. Measurement of the actual substrate concentration in fed-batch systems is difficult and is therefore estimated on the basis of the changes of measurable factors, such as pH, dissolved oxygen and the partial pressure of waste gases, such as CO2. Batch and fed-batch fermentations are common and known in the art. [0195] Continuous fermentation is an open system where a defined fermentation medium is added continuously to a bioreactor, and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the cultures at a constant high density, where cells are primarily in log phase growth. Continuous fermentation allows for the modulation of one or more factors that affect cell growth and/or product concentration.
- a limiting nutrient such as the carbon source or nitrogen source
- a number of factors affecting growth can be altered continuously while the cell concentration, measured by media turbidity, is kept constant. Continuous systems strive to maintain steady state growth conditions. Thus, cell loss due to medium being drawn off should be balanced against the cell growth rate in the fermentation.
- a protein of interest expressed/produced by a Bacillus cell of the disclosure may be recovered from the culture medium by conventional procedures including separating the host cells from the medium by centrifugation or filtration, or if necessary, disrupting the cells and removing the supernatant from the cellular fraction and debris.
- the proteinaceous components of the supernatant or filtrate are precipitated by means of a salt, e.g., ammonium sulfate.
- the precipitated proteins are then solubilized and may be purified by a variety of chromatographic procedures, e.g., ion exchange chromatography, gel filtration. VI.
- a protein of interest (POI) of the instant disclosure can be any endogenous or heterologous protein, and it may be a variant of such a POI.
- the protein can contain one or more disulfide bridges or is a protein whose functional form is a monomer or a multimer, i.e., the protein has a quaternary structure and is composed of a plurality of identical (homologous) or non-identical (heterologous) subunits, wherein the POI or a variant POI thereof is preferably one with properties of interest.
- a recombinant (modified) Bacillus cell of the disclosure produces at least about 0.1% more, at least about 0.5% more, at least about 1% more, at least about 5% more, at least about 6% more, at least about 7% more, at least about 8% more, at least about 9% more, or at least about 10% or more of a POI, relative to a control (or parent) Bacillus cell.
- a modified Bacillus cell of the disclosure exhibits an increased specific productivity (Qp) of a POI relative the control (or parent) Bacillus cell.
- the detection of specific productivity (Qp) is a suitable method for evaluating protein production.
- a modified Bacillus cell of the disclosure comprises a specific productivity (Qp) increase of at least about 0.1%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more, relative to the control (or parent) Bacillus cell.
- a POI or a variant POI thereof is selected from the group consisting of acetyl esterases, aminopeptidases, amylases, arabinases, arabinofuranosidases, carbonic anhydrases, carboxypeptidases, catalases, cellulases, chitinases, chymosins, cutinases, deoxyribonucleases, epimerases, esterases, ⁇ -galactosidases, ⁇ -galactosidases, ⁇ -glucanases, glucan lysases, endo- ⁇ -glucanases, glucoamylases, glucose oxidases, ⁇ -glucosidases, ⁇ -glucosidases, glucuronidases, glycosyl hydrolases, hemicellulases, hexose oxidases, hydrolases, invertases, isomerase
- a POI is an amylase.
- compositions and methods disclosed herein are as follows: [0203] 1. An isolated nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) comprising SEQ ID NO: 2. [0204] 2. A polynucleotide comprising an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) comprising SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI). [0205] 3.
- the polynucleotide of embodiment 2 comprising an upstream 5′-UTR sequence operably linked to the nucleic acid encoding the ypuAss.
- POI heterologous protein of interest
- the polynucleotide of embodiment 4 comprising an upstream 5′-UTR sequence operably linked to the nucleic acid encoding the mod-ypuAss.
- An expression cassette comprising an upstream promoter operably linked to a downstream polynucleotide of any one of embodiments 2-10, wherein the promoter is functional in a Bacillus sp. cell.
- a modified Bacillus sp. cell comprising an introduced polynucleotide comprising an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI), or comprising an introduced polynucleotide comprising an upstream nucleic acid encoding a modified ypuA signal sequence (mod- ypuAss) of SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous POI.
- ypuAss native ypuA signal sequence
- POI heterologous protein of interest
- a modified Bacillus sp. cell comprising at least two introduced polynucleotides, wherein the first polynucleotide comprises an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI), or comprising an upstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous POI
- the second polynucleotide comprises an upstream nucleic acid encoding a ypuAss of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous POI, or comprising an upstream nucleic acid encoding a mod-ypuAss of SEQ ID NO: 2 operably linked to a downstream nucleic acid
- the modified cell of embodiment 16 or embodiment 17, wherein the heterologous POI is an amylase. [0227] 24.
- a modified Bacillus sp. cell comprising an introduced expression cassette of any one of embodiments 11-13.
- 26. A modified Bacillus sp. cell comprising at least two introduced expression cassettes of any one of embodiments 11-13.
- 27. A method for producing a heterologous protein of interest (POI) in a Bacillus sp. cell comprising (a) introducing into a Bacillus sp.
- POI heterologous protein of interest
- an expression cassette comprising an upstream (5′) promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding a heterologous POI, and (b) fermenting the modified cell under suitable conditions for the production of the POI.
- ypuAss native ypuA signal sequence
- mod-ypuAss modified ypuA signal sequence
- the modified cell produces an increased amount of POI relative to a control cell producing the same POI when fermented under the same conditions
- the control cell comprises an introduced expression cassette comprising the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a modified AmyL signal sequence (mod-AmyLss) of SEQ ID NO: 4 operably linked to a downstream nucleic acid encoding the heterologous POI.
- a method for producing a heterologous protein of interest (POI) in a Bacillus sp. cell comprising (a) introducing into a Bacillus sp.
- the first cassette comprises an upstream (5′) promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI
- the second cassette comprises an upstream promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a ypuAss of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a mod-ypuAss of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI, and (b) fermenting the modified cell under suitable
- the modified cell produces an increased amount of POI relative to a control cell producing the same POI when fermented under the same conditions
- the control cell comprises at least two introduced expression cassettes
- the first cassette comprises the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI
- the second cassette comprises the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic
- the polynucleotide further comprises a downstream terminator sequence operably linked to the nucleic acid encoding the POI.
- the terminator sequence comprises at least 95% identity to the B. licheniformis amyL terminator sequence of SEQ ID NO: 12.
- the promoter comprises at least 95% identity to SEQ ID NO: 10 or SEQ ID NO: 20. [0243] 40.
- EXAMPLE 1 CONSTRUCTION OF A TEMPLATE PLASMID FOR A FIRST COPY OF AN AMYLASE-1 INTEGRATION CASSETTE [0247]
- the instant example describes construction of a template DNA (SEQ ID NO: 5) by overlapping extension PCR for the integration of a first (1 st ) copy of an Amylase-1 (Amy-1; SEQ ID NO: 6) expression cassette.
- the 1 st copy of the amy-1 cassette with mod-ypuAss (1 st copy amy-1 cassette mod- ypuAss; SEQ ID NO: 7) comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream DNA comprising a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding a modified B.
- licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13).
- the modified ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) as compared to the native ypuA signal sequence (ypuAss; SEQ ID NO: 1) comprises a substitution of an aspartic acid (Asp; D) to a serine (Ser; S) residue at position 24 of SEQ ID NO: 2 (i.e., the (-2) position relative to the signal peptidase cleavage site of SEQ ID NO: 2).
- the DNA fragments were amplified using Q5 DNA polymerase per the manufacturer’s instructions.
- the PCR products were purified using Zymo clean and concentrate five (5) columns per manufacturer’s instructions.
- the DNA fragments were assembled by overlapping extension PCR, generating the template (SEQ ID NO: 5) for the 1 st amylase-1 integration cassette w/ mod-ypuAss (SEQ ID NO: 7; FIG.2A/FIG.2B).
- SEQ ID NO: 5 the template for the 1 st amylase-1 integration cassette w/ mod-ypuAss
- SEQ ID NO: 7 the 1 st amylase-1 integration cassette w/ mod-ypuAss
- the 2 nd copy of the amy-1 cassette with mod-ypuAss (2 nd copy amy-1 cassette mod-ypuAss;; SEQ ID NO: 17) comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to the lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (pro; SEQ ID NO: 20) operably linked to DNA encoding the B. subtilis aprE 5′UTR (SEQ ID NO: 11) operably linked to DNA encoding the modified B.
- licheniformis ypuA signal sequence (mod- ypuAss; SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21).
- the 1 st copy amy-1 cassette (SEQ ID NO: 7) and the 2 nd copy amy-1 cassette (SEQ ID NO: 17) comprise the same modified ypuA signal sequence (i.e., mod-ypuAss; SEQ ID NO: 2) shown in FIG.1A.
- the DNA fragments were amplified using Q5 DNA polymerase per the manufacturer’s instructions.
- the PCR products were purified using Zymo clean and concentrate 5 columns per manufacturer’s instructions.
- the DNA fragments were assembled by overlapping extension PCR, generating the template (SEQ ID NO: 16) for 2 nd Amylase 1 cassette (SEQ ID NO: 16; FIG.2C/FIG.2D).
- the recombinant B. licheniformis AL223 strain was constructed by integrating the 1 st amy-1 cassette (SEQ ID NO: 7) into the serA locus, and the 2 nd amy-1 cassette (SEQ ID NO: 17) into the lysA locus of the parental BF1965 strain ( ⁇ serA/ ⁇ lysA).
- the 1 st amy-1 cassette (SEQ ID NO: 7) integration fragment was generated by PCR amplification from a template for 1 st amy-1 cassette (SEQ ID NO: 5) with the “oAL106” (SEQ ID NO: 14) and “oAL108” (SEQ ID NO: 15) primer pair.
- the 2 nd amy-1 cassette (SEQ ID NO: 17) integration fragment was generated by PCR amplification from a template for 2 nd amy-1 cassette (SEQ ID NO: 16) with the “oAL079” (SEQ ID NO: 22) and “oAL080” (SEQ ID NO: 23) primer pair.
- the 1 st and 2 nd amy-1 expression cassettes (SEQ ID NO: 7 and SEQ ID NO: 17) were transformed into parental BF1965 strain ( ⁇ serA/ ⁇ lysA) strain using the methods described in PCT Publication No. WO2019/040412.
- the BF1965 competent cells were generated by growing the strain overnight in L broth containing one hundred (100) ppm spectinomycin at 37°C with 250 RPM shaking. The culture was diluted the next day to OD600 of 0.7 of fresh L broth containing one hundred (100) ppm spectinomycin. This new culture was grown for one (1) hour at 37°C, 250 RPM shaking. D-xylose was added to 0.1% w ⁇ v -1 . The culture was grown for an additional four (4) hours at 37°C and 250 RPM shaking. The cells were harvested at 1700 ⁇ g for seven (7) minutes, and used as competent cells for transformation.
- This PCR product a 3,709 bp fragment (SEQ ID NO: 26), was sequenced using the method of Sanger and the “oAL081” (SEQ ID NO: 32), “seq1” (SEQ ID NO: 33) and “seq2” (SEQ ID NO: 34) primers.
- a colony with the correct integration of the cassette (SEQ ID NO: 7) was stored as strain AL217.
- the AL217 competent cells were generated as described above. One hundred (100) ⁇ l of AL217 competent cells were mixed with twenty (20) ⁇ l of the 2 nd Amy-1 cassette (SEQ ID NO: 17) integration fragment. The cell/DNA mixture was incubated at 1200 RPM, 37°C for one and a half (1.5) hours.
- licheniformis AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) was constructed in the BF1965 strain by integration of a 1 st amy-1 cassette (SEQ ID NO: 30) at the serA locus and integration of a 2 nd amy-1 cassette (SEQ ID NO: 31) at the lysA locus.
- the modified B. licheniformis AmyL signal sequence (mod-AmyLss) has been described PCT Publication No.2023/023642, and is an improved signal (secretion) sequence compared to the native B. licheniformis AmyL signal sequence.
- the modified (control) AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) as compared to the native AmyL signal sequence (AmyLss; SEQ ID NO: 3) comprises a substitution of an alanine (Ala, A) to a serine (Ser; S) residue at position 28 of SEQ ID NO: 4 (i.e., the (-2) position relative to the signal peptidase cleavage site of SEQ ID NO: 4).
- the 1 st amy-1 cassette with mod-AmyLss (SEQ ID NO: 30; FIG. 2E/FIG.
- 2F comprises an upstream homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to DNA comprising a serA (ORF; SEQ ID NO: 9) operably linked to the synthetic p3 promoter (SEQ ID NO: 10) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) operably linked to DNA encoding Amy- 1 (SEQ ID NO: 6) operably linked to a B.
- the 2 nd amy-1 cassette with mod-AmyLss comprises an upstream homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to DNA comprising a lysA ORF (SEQ ID NO: 19) operably linked to the synthetic p2 promoter (SEQ ID NO: 20) operably linked to DNA comprising the B.
- subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-AmyLss (SEQ ID NO: 4) operably linked to DNA encoding Amy-1 (SEQ ID NO: 6) operably linked to the B. licheniformis [0258] amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21).
- amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21).
- EXAMPLE 4 EFFECT OF THE MODIFIED YPUA SIGNAL SEQUENCE ON AMYLASE-1 PRODUCTION [0259] In the present example, the B.
- licheniformis AL223 strain comprising two integrated copies of the amylase-1 expression cassettes with the modified B. licheniformis ypuA signal sequence (SEQ ID NO: 2) were assayed for production of the Amy-1 reporter and compared to the control AL207 strain comprising two integrated copies of the amylase-1 expression cassettes with the modified (control) AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) using standard small-scale conditions, as described in PCT Publication No. WO2019/055261 (incorporated herein by reference).
- the Amy-1 protein production was quantified using the method of Bradford assay, wherein the relative improvement in production of Amy-1 from the AL223 strain is compared to the production of the Amy-1 from the AL207 (control) strain, as presented below in TABLE 1 Strain Name Signal Sequence SEQ ID NO Relative Amy1 Production ⁇ CV AL223 mod-ypuAss 2 1.50 ⁇ 0.07 R ELATIVE PERFORMANCE OF MODIFIED YPUA SIGNAL SEQUENCE VERSUS MODIFIED AMYL SIGNAL SEQUENCES ON AMY-1 PRODUCTION [0260]
- the modified ypuA signal sequence (SEQ ID NO: 2) demonstrates an approximately 50% improvement in Amy-1 protein production (strain AL223) relative to the Amy-1 protein production (control strain AL207) with the modified AmyL (control) signal sequence (SEQ ID NO: 4).
- EXAMPLE 5 CONSTRUCTION OF A TEMPLATE PLASMID FOR A SECOND COPY OF AN AMYLASE-2 INTEGRATION CASSETTE
- Applicant identified a native B. licheniformis ypuA protein signal sequence (SEQ ID NO: 1) that is particularly useful for enhancing/increasing production (secretion) of Amylase-2 (Amy-2; SEQ ID NO: 37).
- the instant example describes construction of a template plasmid named “pWS704” (SEQ ID NO: 36) for the integration of second copy of Amylase-2 (Amy-2; SEQ ID NO: 37) expression cassette.
- the 2 nd copy amy-2 cassette with a native ypuA signal sequence comprises an upstream (5′) homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to a downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the native B.
- licheniformis YpuA signal sequence (ypuAss; SEQ ID NO: 1) operably linked to DNA encoding the Amy- 2 protein (SEQ ID NO: 37) operably linked to a B. licheniformis amyL transcriptional terminator (AmyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21).
- the amino acid positions of the native B. licheniformis ypuA signal sequence (ypuAss) comprise amino acid (residue) positions M 1 - A 25 , as presented in FIG.1A (SEQ ID NO: 1).
- the DNA fragments were amplified using Q5 DNA polymerase per the manufacturer’s instructions.
- the PCR products were purified using Zymo clean and concentrate five (5) columns per manufacturer’s instructions.
- the DNA fragments were assembled into a plasmid named “pRS426” purchased from ATCC (ATCC Catalogue No. 7710) by using yeast gap-repair cloning method (Joska et al., 2014), generating plasmid pWS704 (SEQ ID NO: 36).
- EXAMPLE 6 CONSTRUCTION OF A BACILLUS LICHENIFORMIS STRAIN EXPRESSING TWO COPIES OF AMYLASE-2
- the 2 nd copy amy-2 expression cassette (SEQ ID NO: 38) constructed in Example 5 was integrated into the lysA locus of the B. licheniformis strain BF780, resulting in the modified B. licheniformis host strain named “WS2756”. More particularly, the WS2756 host was constructed by integrating the 2 nd copy of the amy-2 cassette (SEQ ID NO: 38) into the lysA locus of the BF780 strain.
- amy-2 expression cassette (with native ypuAss) was generated by PCR amplification from the plasmid template pWS704 (SEQ ID NO: 36) with the ws683 (SEQ ID NO: 41) and ws688 (SEQ ID NO: 42) primer pair.
- the amy-2 cassette integration fragment was transformed into the BF780 strain ( ⁇ serA::[p3-mod-AmyLss-amy-2]serA- ⁇ lysA) using the methods described in PCT Publication No. WO2023/023642.
- the BF780 competent cells were generated by growing the strain overnight in L broth containing one hundred (100) ppm spectinomycin at 37°C with 250 RPM shaking. The culture was diluted the next day to OD600 of 0.7 of fresh L broth containing one hundred (100) ppm spectinomycin. This new culture was grown for one (1) hour at 37°C, 250 RPM shaking. D-xylose was added to 0.1% w ⁇ v -1 . The culture was grown for an additional four (4) hours at 37°C and 250 RPM shaking. The cells were harvested at 1700 ⁇ g for seven (7) minutes, and used as competent cells for transformation.
- This PCR product a 1,892 bp fragment (SEQ ID NO: 45), was sequenced using the method of Sanger and the ws775 and ws776 primers.
- a colony with the correct integration of the 2 nd copy amy-2 cassette (FIG. 3C/FIG. 3D; SEQ ID NO: 38) was stored as strain WS2756 (serA::[p3-mod-AmyLss-amylase 2] serA lysA::[p2-ypuAss-amy-2] lysA).
- the 2 nd copy amy-2 cassette with mod- AmyLss comprises an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to a DNA sequence encoding lysA (ORF; SEQ ID NO: 19) operably linked to the synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the modified B.
- licheniformis AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) operably linked to DNA encoding the Amy-2 reporter protein (SEQ ID NO: 37) operably linked to the B. licheniformis amyL transcriptional terminator (SEQ ID NO: 12) operably linked to the downstream homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21).
- a colony with the correct integration of the amy-2 cassette (SEQ ID NO: 38) was stored as a control strain BF822 (serA::[p3-mod-AmyLss-amylase 2] serA lysA::[p2-AmyLss-amy-2] lysA).
- the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) relative to the native AmyL signal sequence (AmyLss; SEQ ID NO: 3) comprises a substitution of alanine (A) to serine (S) at the minus (-) 2 position, relative to the signal peptidase cleavage site.
- EXAMPLE 7 EFFECT OF THE YPUA SIGNAL SEQUENCE ON AMYLASE-2 PRODUCTION [0267]
- the WS2756 strain and the BF822 (control) strain constructed and described in Example 6 were assayed for production of the Amy-2 (SEQ ID NO: 37) reporter using standard small-scale conditions, as described in PCT Publication No.
- licheniformis ypuA signal sequence demonstrates an improvement in Amy-2 reporter protein production in the WS2756 strain relative to the B. licheniformis (control) strain BF822, comprising the modified B. licheniformis AmyL signal sequence (modAmyLss).
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Organic Chemistry (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Wood Science & Technology (AREA)
- Genetics & Genomics (AREA)
- Zoology (AREA)
- Biochemistry (AREA)
- Molecular Biology (AREA)
- General Health & Medical Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Medicinal Chemistry (AREA)
- General Engineering & Computer Science (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Biotechnology (AREA)
- Microbiology (AREA)
- Chemical Kinetics & Catalysis (AREA)
- General Chemical & Material Sciences (AREA)
- Biomedical Technology (AREA)
- Gastroenterology & Hepatology (AREA)
- Biophysics (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
Abstract
The present disclosure is generally related to recombinant microbial cells expressing heterologous proteins of interest. Certain aspects of the disclosure are therefore related to, inter alia, recombinant Bacillus cells having enhanced protein production capabilities, novel protein signal sequences, recombinant polynucleotides encoding heterologous proteins of interest, and related compositions and/or methods thereof. Thus, as exemplified herein, the recombinant Bacillus cells of the instant disclosure are particularly suitable for use in the expression, production and secretion of heterologous proteins.
Description
METHODS AND COMPOSITIONS FOR ENHANCED PROTEIN PRODUCTION IN BACILLUS CELLS FIELD [0001] The present disclosure is generally related to the fields of microbial cells, molecular biology, fermentation, protein production, and the like. Certain aspects of the disclosure are related to, inter alia, recombinant Bacillus cells having enhanced protein production capabilities. CROSS REFERENCE TO RELATED APPLICATIONS [0002] This application claims benefit to U.S. Provisional Patent Application No. 63/597,527, filed November 9, 2023, which is incorporated herein by referenced in its entirety. REFERENCE TO A SEQUENCE LISTING [0003] The contents of the electronic submission of the text file Sequence Listing, named “NB42165-US- PSP_SequenceListing.txt” was created on November 07, 2023 and is 157 KB in size, which is hereby incorporated by reference in its entirety. BACKGROUND [0004] Gram-positive bacteria such as Bacillus subtilis, Bacillus licheniformis, Bacillus amyloliquefaciens and the like are frequently used as microbial factories for the production of industrial relevant proteins, due to their excellent fermentation properties and high yields (e.g., up to 25 grams per liter culture; Van Dijl and Hecker, 2013). For example, Bacillus sp. cells are well known for their production of amylases (Jensen et al., 2000; Raul et al., 2014) and proteases (Brode et al., 1996) necessary for food, textile, laundry, medical instrument cleaning, pharmaceutical industries and the like (Westers et al., 2004). Because these non- pathogenic Gram-positive bacteria produce proteins that completely lack toxic by-products (e.g., lipopolysaccharides; LPS, also known as endotoxins) they have obtained the “Qualified Presumption of Safety” (QPS) status of the European Food Safety Authority, and many of their products gained a “Generally Recognized As Safe” (GRAS) status from the US Food and Drug Administration (Olempska- Beer et al., 2006; Earl et al., 2008; Caspers et al., 2010). Thus, the production of proteins (e.g., enzymes, antibodies, receptors, etc.) in Gram-positive bacterial cells is an area of high interest in the biotechnological arts, wherein small improvements in protein yield are quite significant when the protein is produced in large industrial quantities. [0005] As described hereinafter, the instant disclosure is related to the highly desirable and unmet needs for obtaining, constructing, producing and the like, Gram-positive host cells having increased protein production capabilities. More specifically, certain aspects of the instant disclosure are related to, among
other things, compositions and methods for constructing recombinant Bacillus cells having enhanced protein production capabilities, which recombinant cells are particularly useful in the production of heterologous proteins. SUMMARY [0006] As briefly set forth above, certain embodiments of the disclosure are related to recombinant (modified) Bacillus cells (strains) capable of expressing/producing increased amounts of proteins of interest. Certain embodiments of the disclosure therefore provide, inter alia, nucleic acids, polynucleotides, vectors, expression cassettes, signal (peptide) sequences, proteins of interest, microbial cells, methods for constructing Bacillus cells producing proteins of interest, methods for cultivating Bacillus cells for the expression/production of proteins of interest and the like. [0007] In particular embodiments, the disclosure provides an isolated nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) set for in SEQ ID NO: 2. In other embodiments, the disclosure provides a polynucleotide comprising an upstream (5ʹ) nucleic acid encoding a native ypuA signal sequence (ypuAss) set forth in SEQ ID NO: 1 operably linked to a downstream (3ʹ) nucleic acid encoding a heterologous protein of interest (POI). In other embodiments, the disclosure provides a polynucleotide comprising an upstream (5ʹ) nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to a downstream (3ʹ) nucleic acid encoding a heterologous POI. [0008] In certain other embodiments, the disclosure is related to expression cassettes comprising an upstream (heterologous) promoter operably linked to a downstream 5′-UTR operably linked to a downstream polynucleotide comprising an upstream nucleic acid encoding the ypuAss (SEQ ID NO: 1) operably linked to the downstream nucleic acid encoding the POI, wherein the promoter and 5′-UTR sequences are functional in Bacillus sp. cells. In other embodiments, the disclosure is related to expression cassettes comprising an upstream (heterologous) promoter operably linked to a downstream 5′-UTR operably linked to a downstream polynucleotide comprising an upstream nucleic acid encoding the mod- ypuAss (SEQ ID NO: 2) operably linked to the downstream nucleic acid encoding the POI, wherein the promoter and 5′-UTR sequences are functional in Bacillus sp. cells. Certain other embodiments provide modified Bacillus sp. cells comprising an introduced polynucleotide or expression cassette of the disclosure. [0009] In other one or more other embodiments, the disclosure is related to methods for constructing modified Bacillus sp. cells producing a heterologous POI. In certain embodiments, the disclosure is related to methods for producing a heterologous POI in a Bacillus sp. cell comprising (a) introducing into a Bacillus sp. cell an expression cassette comprising an upstream (5′) promoter operably linked to a downstream 5′- UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA
signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding a heterologous POI, and (b) fermenting the modified cell under suitable conditions for the production of the POI. In a related embodiment of the methods, the modified cell produces an increased amount of the POI relative to a control cell producing the same POI, wherein the control cell comprises an introduced expression cassette comprising the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a modified AmyL signal sequence (mod-AmyLss) of SEQ ID NO: 4 operably linked to a downstream nucleic acid encoding the heterologous POI. BRIEF DESCRIPTION OF DRAWINGS [0010] Figure 1 presents amino acid sequence alignments of certain native and modified protein signal sequences of the disclosure. More particularly, FIG.1A shows an alignment of the native B. licheniformis ypuA signal sequence (ypuAss; SEQ ID NO: 1) relative to the modified ypuA signal sequence (mod- ypuAss; SEQ ID NO: 2). As presented in FIG. 1A, the native ypuA signal sequence (SEQ ID NO: 1) comprises twenty-five (25) amino acid residues with an aspartic acid (Asp; D24) residue at position 24 of SEQ ID NO: 1, whereas the modified ypuA signal sequence (SEQ ID NO: 2) comprises a substitution of the aspartic acid (D24) to a serine (Ser; S24) residue at position 24 of SEQ ID NO: 2. Similarly, FIG. 1B shows an alignment of the native B. licheniformis AmyL signal sequence (AmyLss; SEQ ID NO: 3) relative to the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4). As presented in FIG.1B, the native B. licheniformis AmyL signal sequence (AmyLss; SEQ ID NO: 3) comprises twenty-nine (29) amino acid residues with an alanine (Ala; A28) residue at position 28 of SEQ ID NO: 3, whereas the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) comprises a substitution of the alanine (A28) to a serine (Ser; S28) residue at position 28 of SEQ ID NO: 4. [0011] Figure 2 presents schematics showing the general design of the amylase-1 expression cassettes suitable for integration into Gram-positive bacterial cells of the disclosure. In particular, as shown in FIG.2A, a first copy of the amylase-1 integration cassette (1st copy amy-1 cassette with modified ypuA signal sequence; SEQ ID NO: 7) comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream (DNA) encoding a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to a downstream DNA encoding the Amylase-1 (Amy-1; SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13). FIG. 2B presents the sequence identification numbers (SEQ) of the 1st copy amylase-1 cassette (SEQ ID NO: 7) shown in FIG. 2A. Similarly, FIG. 2C shows the design of the second copy of the amylase-1 integration
cassette (2nd copy amy-1 cassette with modified ypuA signal sequence; SEQ ID NO: 17) comprising an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to a downstream DNA encoding the Amylase-1 (Amy-1; SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21). FIG.2D presents the sequence identification numbers (SEQ) of the 2nd copy amylase-1 integration cassette (SEQ ID NO: 17) shown in FIG.2C. Likewise, as presented in FIG. 2E-FIG. 2H, a first copy amylase-1 integration cassette (1st copy amy-1 cassette with modified AmyL signal sequence; SEQ ID NO: 30) and a second copy amylase-1 integration cassette (2nd copy amy-1 cassette with modified AmyL signal sequence; SEQ ID NO: 31) were constructed as controls. More particularly, as shown in FIG.2E, the 1st copy amy-1 cassette (SEQ ID NO: 30) comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream (DNA) encoding a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B. licheniformis AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) operably linked to a downstream DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL- term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13). FIG.2F presents the sequence identification numbers (SEQ) of the 1st copy amylase-1 cassette (SEQ ID NO: 30) shown in FIG.2E. As shown in FIG.2G, the 2nd copy amy-1 cassette (SEQ ID NO: 31) comprises an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B. licheniformis Amyl signal sequence (mod-AmyLss; SEQ ID NO: 4) operably linked to a downstream DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21). FIG.2H presents the sequence identification numbers (SEQ) of the 2nd copy amylase-1 integration cassette (SEQ ID NO: 31) shown in FIG.2G. [0012] Figure 3 presents schematics showing the general design of certain amylase-2 expression cassettes suitable for integration into Gram-positive bacterial cells of the disclosure. In particular, as shown in
FIG.3A, a first copy of the amylase-2 cassette (1st copy amy-2 cassette with mod-AmyLss; SEQ ID NO: 39) comprises a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B. licheniformis mod-AmyLss (SEQ ID NO: 4) operably linked to a downstream DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO: 12). FIG. 3B presents the sequence identification numbers (SEQ) of the 1st copy amylase-2 cassette (SEQ ID NO: 39) shown in FIG. 3A. As shown in FIG. 3C, the 2nd copy amy-2 cassette with ypuAss (SEQ ID NO: 38) comprises an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the native B. licheniformis ypuA signal sequence (ypuAss; SEQ ID NO: 1) operably linked to a downstream DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL-term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21). FIG.3D presents the sequence identification numbers (SEQ) of the 2nd copy amylase-2 integration cassette (SEQ ID NO: 38) shown in FIG. 3C. As presented in FIG. 3E, a 2nd copy amy-2 cassette with mod-AmyLss (SEQ ID NO: 40) comprises an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream (DNA) encoding a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to a downstream DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to a downstream DNA encoding the modified B. licheniformis AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) operably linked to a downstream DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL- term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21). FIG.3F presents the sequence identification numbers (SEQ) of the 2nd copy amylase-2 integration cassette (SEQ ID NO: 40) shown in FIG.3E. BRIEF DESCRIPTION OF THE BIOLOGICAL SEQUENCES [0013] SEQ ID NO: 1 is the amino acid sequence of the native B. licheniformis ypuA signal sequence (ypuAss). [0014] SEQ ID NO: 2 is the amino acid sequence of a modified ypuA signal sequence (mod-ypuAss). [0015] SEQ ID NO: 3 is the amino acid sequence the native B. licheniformis AmyL signal sequence (AmyLss). [0016] SEQ ID NO: 4 is the amino acid sequence of a modified AmyL signal sequence (mod-AmyLss).
[0017] SEQ ID NO: 5 is a DNA template for the first (1st) copy amylase-1 expression cassette. [0018] SEQ ID NO: 6 is the mature amino acid sequence of an engineered (variant) B. licheniformis Amylase-1 (Amy-1) reporter protein. [0019] SEQ ID NO: 7 is an integration (expression) cassette encoding the 1st copy Amyl-1 reporter protein with mod-ypuAss. [0020] SEQ ID NO: 8 is a B. licheniformis serA upstream (serA.up) homology arm (DNA) sequence. [0021] SEQ ID NO: 9 is the open reading frame (ORF) of the B. licheniformis serA gene. [0022] SEQ ID NO: 10 is a synthetic p3 promoter sequence. [0023] SEQ ID NO: 11 is a wild-type B. subtilis aprE 5ʹ-UTR sequence. [0024] SEQ ID NO: 12 is a wild-type B. licheniformis AmyL terminator sequence. [0025] SEQ ID NO: 13 is a B. licheniformis serA downstream (serA.down) homology arm sequence. [0026] SEQ ID NO: 14 is a synthetic primer sequence named “oAL106”. [0027] SEQ ID NO: 15 is a synthetic primer sequence named “oAL108”. [0028] SEQ ID NO: 16 is a DNA template for the second (2nd) copy amylase-1 expression cassette. [0029] SEQ ID NO: 17 is an integration (expression) cassette encoding the 2nd copy Amy-1 reporter protein with mod-ypuAss. [0030] SEQ ID NO: 18 is a B. licheniformis lysA upstream (lysA.up) homology arm sequence. [0031] SEQ ID NO: 19 is the open reading frame (ORF) of the B. licheniformis lysA gene. [0032] SEQ ID NO: 20 is a synthetic p2 promoter sequence. [0033] SEQ ID NO: 21 is a B. licheniformis lysA downstream (lysA.down) homology arm sequence. [0034] SEQ ID NO: 22 is a synthetic primer sequence named “oAL079”. [0035] SEQ ID NO: 23 is a synthetic primer sequence named “oAL080”. [0036] SEQ ID NO: 24 is a synthetic primer sequence named “oAL069”. [0037] SEQ ID NO: 25 is a synthetic primer sequence named “oAL076”. [0038] SEQ ID NO: 26 is a 3,709 bp DNA fragment for screening integration of the 1st copy amy-1 cassette (SEQ ID NO: 7) with the mod-ypuAss. [0039] SEQ ID NO: 27 is a synthetic primer sequence named “oAL082”. [0040] SEQ ID NO: 28 is a synthetic primer sequence named “oAL083”. [0041] SEQ ID NO: 29 is a 2,133 bp DNA fragment for screening integration of the 2nd copy amy-1 cassette (SEQ ID NO: 17) with the mod-ypuAss. [0042] SEQ ID NO: 30 is a control integration (expression) cassette encoding the 1st copy Amyl-1 reporter protein with mod-AmyLss. [0043] SEQ ID NO: 31 is a control integration (expression) cassette encoding the 2nd copy Amyl-1 reporter protein with mod-AmyLss.
[0044] SEQ ID NO: 32 is a synthetic primer sequence named “oAL081”. [0045] SEQ ID NO: 33 is a synthetic primer sequence named “seq1”. [0046] SEQ ID NO: 34 is a synthetic primer sequence named “seq2”. [0047] SEQ ID NO: 35 is a synthetic primer sequence named “oAL086”. [0048] SEQ ID NO: 36 is a plasmid (DNA) template (pWS704) for a second (2nd) copy amylase-2 integration (expression) cassette encoding a 2nd copy of the Amy-2 protein (SEQ ID NO: 37) with the native ypuAss (SEQ ID NO: 1) [0049] SEQ ID NO: 37 is the mature amino acid sequence of an engineered (variant) Cytophaga sp. Amylase-2 (Amy-2) reporter protein. [0050] SEQ ID NO: 38 is an integration (expression) cassette encoding the 2nd copy Amy-2 reporter protein with ypuAss. [0051] SEQ ID NO: 39 is an integration (expression) cassette encoding the 1st copy Amy-2 reporter protein with mod-AmyLss. [0052] SEQ ID NO: 40 is an integration (expression) cassette encoding the 2nd copy Amy-2 reporter protein with mod-AmyLss. [0053] SEQ ID NO: 41 is a synthetic primer sequence named “ws683”. [0054] SEQ ID NO: 42 is a synthetic primer sequence named “ws688”. [0055] SEQ ID NO: 43 is a synthetic primer sequence named “ws775”. [0056] SEQ ID NO: 44 is a synthetic primer sequence named “ws776”. [0057] SEQ ID NO: 45 is a 1,892 bp DNA fragment for screening integration of the 2nd copy amy-2 cassette (SEQ ID NO: 38). [0058] SEQ ID NO: 46 is an expression cassette named “(p3 pro)-(aprE 5′-UTR)-(mod-ypuAss)-(amy-1)”. [0059] SEQ ID NO: 47 is an expression cassette named “(p2 pro)-(aprE 5′-UTR)-(mod-ypuAss)-(amy-1)”. [0060] SEQ ID NO: 48 is an expression cassette named “(p3 pro)-(aprE 5′-UTR)-(ypuAss)-(amy-1)”. [0061] SEQ ID NO: 49 is an expression cassette named “(p2 pro)-(aprE 5′-UTR)-(ypuAss)-(amy-1)”. [0062] SEQ ID NO: 50 is an expression cassette named “(p3 pro)-(aprE 5′-UTR)-(mod-ypuAss)-(amy-2)”. [0063] SEQ ID NO: 51 is an expression cassette named “(p2 pro)-(aprE 5′-UTR)-(mod-ypuAss)-(amy-2)”. [0064] SEQ ID NO: 52 is an expression cassette named “(p3 pro)-(aprE 5′-UTR)-(ypuAss)-(amy-2)”. [0065] SEQ ID NO: 53 is an expression cassette named “(p2 pro)-(aprE 5′-UTR)-(ypuAss)-(amy-2)”. DETAILED DESCRIPTION [0066] As described herein, the instant disclosure addresses numerous ongoing and unmet needs in the art, particularly as related to the industrial scale production recombinant proteins. Certain embodiments of the instant disclosure provide, inter alia, recombinant Bacillus cells capable expressing increased amounts of
proteins of interest. Certain aspects of the disclosure therefore provide, among other things, novel (recombinant) Bacillus cells comprising introduced nucleic acids (e.g., vectors, expression cassettes) encoding proteins of interest, polynucleotide constructs encoding modified (protein) signal sequences operably linked to a downstream nucleic acid a encoding protein of interest, recombinant Bacillus cells comprising one or more introduced polynucleotide constructs, and related methods for cultivating and expressing heterologous proteins of interest in a recombinant Bacillus cell of the disclosure and the like. I. DEFINITIONS [0067] In view of the recombinant cells, nucleic acids, polynucleotides, proteins of interest and the like, the following terms and phrases are defined. Terms not defined herein should be accorded their ordinary meaning as used in the art. [0068] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present compositions and methods apply. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present compositions and methods, representative illustrative methods and materials are now described. All publications and patents cited herein are incorporated by reference in their entirety. [0069] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely”, “only”, “excluding”, “not including” and the like, in connection with the recitation of claim elements, or use of a “negative” limitation or proviso thereof. [0070] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present compositions and methods described herein. Any recited method can be carried out in the order of events recited or in any other order which is logically possible. [0071] As used herein, the terms “recombinant” or “non-natural” refer to an organism, microorganism, cell, nucleic acid molecule, or vector that has at least one engineered genetic alteration, or has been modified by the introduction of a heterologous nucleic acid molecule, or refer to a cell (e.g., a Gram-positive cell) that has been altered such that the expression of a heterologous nucleic acid molecule or an endogenous nucleic acid molecule or a gene can be controlled. Recombinant also refers to a cell that is derived from a non-natural cell, or is progeny of a non-natural cell having one or more such modifications. Genetic alterations include, for example, modifications introducing expressible nucleic acid molecules encoding proteins, or other nucleic acid molecule additions, deletions, substitutions or other functional alteration of a cell’s genetic material. For example, recombinant cells may express genes or other nucleic acid molecules
(e.g., polynucleotide expression constructs) that are not found in identical or homologous form within a native (wild-type) cell, or may provide an altered expression pattern of endogenous genes, such as being over-expressed, under-expressed, minimally expressed, or not expressed at all. “Recombination”, “recombining” or generating a “recombined” nucleic acid is generally the assembly of two or more nucleic acid fragments wherein the assembly gives rise to a chimeric DNA sequence that would not otherwise be found in the genome. [0072] The term “derived” encompasses the terms “originated”, “obtained”, “obtainable”, and “created” and generally indicates that one specified material or composition finds its origin in another specified material or composition, or has features that can be described with reference to the other specified material or composition. [0073] As used herein, “nucleic acid” refers to a nucleotide or polynucleotide sequence, and fragments or portions thereof, as well as to DNA, cDNA, and RNA of genomic or synthetic origin, which may be double- stranded or single-stranded, whether representing the sense or antisense strand. It will be understood that as a result of the degeneracy of the genetic code, a multitude of nucleotide sequences may encode a given protein. [0074] It is understood that the polynucleotides (or nucleic acid molecules) described herein include “genes”, “vectors” and “plasmids”. [0075] Accordingly, the term “gene”, refers to a polynucleotide that codes for a particular sequence of amino acids, which comprise all, or part of a protein coding sequence, and may include regulatory (non- transcribed) DNA sequences, such as promoter sequences, which determine for example the conditions under which the gene is expressed. The transcribed region of the gene may include untranslated regions (UTRs), including introns, 5′-untranslated regions (UTRs), and 3′-UTRs, as well as the coding sequence. [0076] As used herein, an “endogenous gene” refers to a gene in its natural location in the genome of an organism. [0077] As used herein, a “heterologous” gene, a “non-endogenous” gene, or a “foreign” gene refer to a gene not normally found in the host organism, but that is introduced into the host organism by gene transfer. The term “foreign” gene(s) comprises native genes inserted into a non-native organism and/or chimeric genes inserted into a native or non-native organism. [0078] As used herein, a “heterologous control sequence”, refers to a gene expression control sequence (e.g., promoters, enhancers, terminators, etc.) which does not function in nature to regulate (control) the expression of the gene of interest. Generally, heterologous nucleic acids are not endogenous (native) to the cell, or a part of the genome in which they are present, and have been added to the cell, by infection, transfection, transduction, transformation, microinjection, electroporation, and the like. A “heterologous” nucleic acid construct may contain a control sequence/DNA coding (ORF) sequence combination that is
the same as, or different, from a control sequence/DNA coding sequence combination found in the native host cell. [0079] As used herein, the terms “signal sequence” and “signal peptide” refer to a sequence of amino acid residues that may participate in the secretion or direct transport of a mature protein or precursor form of a protein. The signal sequence is typically located N-terminal to the precursor or mature protein sequence. The signal sequence may be endogenous or exogenous. A signal sequence is normally absent from the mature protein. A signal sequence is typically cleaved from the protein by a signal peptidase during translocation. [0080] As used herein, the term “expression” refers to the transcription and stable accumulation of sense (mRNA) or anti-sense RNA, derived from a nucleic acid molecule of the disclosure. Expression may also refer to translation of mRNA into a polypeptide. Thus, the term “expression” includes any steps involved in the production of the polypeptide including, but not limited to, transcription, post-transcriptional modification, translation, post-translational modification, secretion and the like. [0081] As used herein, the term “coding sequence” refers to a nucleotide sequence, which directly specifies the amino acid sequence of its (encoded) protein product. The boundaries of the coding sequence are generally determined by an open reading frame (hereinafter, “ORF”), which usually begins with an ATG start codon. The coding sequence typically includes DNA, cDNA, and recombinant nucleotide sequences. [0082] The term “promoter” as used herein refers to a nucleic acid sequence capable of controlling the expression of a coding sequence or functional RNA. In general, a coding sequence is located 3' (downstream) to a promoter sequence. Promoters may be derived in their entirety from a native gene, or be composed of different elements derived from different promoters found in nature, or even comprise synthetic nucleic acid segments. It is understood by those skilled in the art that different promoters may direct the expression of a gene in different cell types, or at different stages of development, or in response to different environmental or physiological conditions. Promoters can be constitutive promoters, inducible promoters, tunable promoters, hybrid promoters, synthetic promoters, tandem promoters, etc. Promoters which cause a gene to be expressed in most cell types at most times are commonly referred to as “constitutive promoters”. It is further recognized that since in most cases the exact boundaries of regulatory sequences have not been completely defined, DNA fragments of different lengths may have identical promoter activity. [0083] As used herein, “a functional promoter sequence controlling the expression of a gene of interest linked to the gene of interest’s protein coding sequence” refers to a promoter sequence which controls the transcription and translation of the coding sequence in a desired Gram-positive host cell. For example, in certain embodiments, the present disclosure is directed to a polynucleotide comprising an upstream (5′)
promoter (or 5′ promoter region, or tandem 5′ promoters and the like) functional in a Gram-positive cell, wherein the promoter region is operably linked to a nucleic acid sequence encoding a protein of interest. [0084] The term “operably linked” as used herein refers to the association of nucleic acid sequences on a single nucleic acid fragment so that the function of one is affected by the other. For example, a promoter is operably linked with a coding sequence when it is capable of affecting the expression of that coding sequence (i.e., that the coding sequence is under the transcriptional control of the promoter). Coding sequences can be operably linked to regulatory sequences in sense or antisense orientation. [0085] A nucleic acid is “operably linked” when it is placed into a functional relationship with another nucleic acid sequence. For example, DNA encoding a secretory leader (i.e., a signal sequence), is operably linked to DNA for a polypeptide if it is expressed as a pre-protein that participates in the secretion of the polypeptide; a promoter or enhancer is operably linked to a coding sequence if it affects the transcription of the sequence; or a ribosome binding site is operably linked to a coding sequence if it is positioned so as to facilitate translation. Generally, “operably linked” means that the DNA sequences being linked are contiguous, and, in the case of a secretory leader, contiguous and in reading phase. However, enhancers do not have to be contiguous. Linking is accomplished by ligation at convenient restriction sites. If such sites do not exist, the synthetic oligonucleotide adaptors or linkers are used in accordance with conventional practice. [0086] As used herein, “suitable regulatory sequences” refer to nucleotide sequences located upstream (5′ non-coding sequences), within, or downstream (3′ non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences may include promoters, transcription leader sequences, RNA processing site, effector binding site and stem-loop structures. [0087] As used herein, a native B. licheniformis “ypuA signal sequence” (abbreviated, “ypuAss”) comprises the amino acid sequence set forth in SEQ ID NO: 1. In particular, as shown in FIG. 1A, the native ypuAss (SEQ ID NO: 1) comprises twenty-five (25) amino acid residues shown as positions M1-A25 (FIG.1A). [0088] As used herein, a modified (variant) B. licheniformis “ypuA signal sequence” (abbreviated, “mod- ypuAss”) comprises the amino acid sequence set forth in SEQ ID NO: 2. In particular, as shown in FIG. 1A, the mod-ypuAss (SEQ ID NO: 2) comprises twenty-five (25) amino acid residues shown as positions M1-A25 (FIG.1A). [0089] As used herein, a native B. licheniformis “AmyL signal sequence” (abbreviated, “AmyLss”) comprises the amino acid sequence set forth in SEQ ID NO: 3. In particular, as shown in FIG. 1B, the native AmyLss (SEQ ID NO: 3) comprises twenty-nine (29) amino acid residues shown as positions M1- A29 (FIG.1B).
[0090] As used herein, a modified (variant) B. licheniformis “AmyL signal sequence” (abbreviated, “mod- AmyLss”) comprises the amino acid sequence set forth in SEQ ID NO: 4. As shown in FIG.1B, the AmyLss (SEQ ID NO: 4) comprises twenty-nine (29) amino acid residues shown as positions M1-A29 (FIG. 1B). For instance, recombinant B. licheniformis strains expressing heterologous proteins of interest operably linked to the upstream (N-terminal) mod-AmyLss (SEQ ID NO: 4) have been described in PCT Publication No. WO 2023/023642. [0091] As used herein, a B. licheniformis “serA upstream homology arm” (abbreviated, “serA.up”) comprises the nucleic acid sequence set forth in SEQ ID NO: 8, a B. licheniformis “serA downstream homology arm” (abbreviated, “serA.down”) comprises the nucleic acid sequence set forth in SEQ ID NO: 13, and the “open reading frame” (ORF) of the B. licheniformis serA gene comprises the nucleic acid sequence set forth in SEQ ID NO: 9. [0092] As used herein, a B. licheniformis “lysA upstream homology arm” (abbreviated, “lysA.up”) comprises the nucleic acid sequence set forth in SEQ ID NO: 18, a B. licheniformis “lysA downstream homology arm” (abbreviated, “lysA.down”) comprises the nucleic acid sequence set forth in SEQ ID NO: 21, and the ORF of the B. licheniformis lysA gene comprises the nucleic acid sequence set forth in SEQ ID NO: 19. [0093] As ushed herein, the “synthetic p3 promoter sequence” (abbreviated, “p3 pro”) comprises the nucleic acid sequence set forth in SEQ ID NO: 10, and the “synthetic p2 promoter sequence” (abbreviated, “p2 pro”) comprises the nucleic acid sequence set forth in SEQ ID NO: 20. [0094] As used herein, a wild-type B. subtilis “aprE 5ʹ-untranslated region” (abbreviated, “aprE 5ʹ- UTR”) comprises the nucleic acid sequence set forth in SEQ ID NO: 11. [0095] As used herein, a wild-type B. licheniformis “AmyL terminator sequence” (abbreviated, “AmyL term”) comprises the nucleic acid sequence set forth in SEQ ID NO: 12. [0096] As used herein, a “first copy of the amylase-1 integration (expression) cassette with a modified ypuA signal sequence” (abbreviated, “1st copy amy-1 cassette with mod-ypuAss”) comprises the polynucleotide (DNA) sequence set forth in SEQ ID NO: 7. More particularly, as schematically presented in FIG. 2A, the 1st copy amy-1 cassette with mod-ypuAss comprises an upstream (5ʹ) homology arm (up) for the B. licheniformis serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream DNA comprising a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-ypuAss (SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13).
[0097] As used herein, a “second copy of the amylase-1 integration (expression) cassette with a modified ypuA signal sequence” (abbreviated, “2nd copy amy-1 cassette with mod-ypuAss) comprises the polynucleotide sequence set forth in SEQ ID NO: 17. More particularly, as schematically presented in FIG. 2C, the 2nd copy amy-1 cassette with mod-ypuAss comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-ypuAss (SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21). [0098] As used herein, a “first copy of the amylase-1 integration (expression) cassette with a modified AmyL signal sequence” (abbreviated, “1st copy amy-1 cassette with mod-AmyLss”) comprises the polynucleotide (DNA) sequence set forth in SEQ ID NO: 30. More particularly, as schematically presented in FIG.2E, the 1st copy amy-1 cassette with mod-AmyLss comprises an upstream (5ʹ) homology arm (up) for the B. licheniformis serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream DNA comprising a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-AmyLss (SEQ ID NO: 4) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13). [0099] As used herein, a “second copy of the amylase-1 integration (expression) cassette with a modified AmyL signal sequence” (abbreviated, “2nd copy amy-1 cassette with mod-AmyLss”) comprises the polynucleotide sequence set forth in SEQ ID NO: 30. More particularly, as schematically presented in FIG.2G, the 2nd copy amy-1 cassette with mod-AmyLss comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-ypuAss (SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21).
[0100] As used herein, a “first copy of the amylase-2 cassette with a modified AmyL signal sequence” (abbreviated, “1st copy amy-2 cassette with mod-AmyLss”) comprises the polynucleotide (DNA) sequence set forth in SEQ ID NO: 39. More particularly, as schematically presented in FIG.3A, the 1st copy amy-2 cassette with mod-AmyLss comprises a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-AmyLss (SEQ ID NO: 4) operably linked to DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12). [0101] As used herein, a “second copy of the amylase-2 integration (expression) cassette with a native ypuA signal sequence” (abbreviated, “2nd copy amy-2 cassette with ypuAss”) comprises the polynucleotide sequence set forth in SEQ ID NO: 38. More particularly, as schematically presented in FIG. 3C, the 2nd copy amy-2 cassette with ypuAss comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the ypuAss (SEQ ID NO: 1) operably linked to DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21). [0102] As used herein, a “second copy of the amylase-2 integration (expression) cassette with a modified AmyL signal sequence” (abbreviated, “2nd copy amy-2 cassette with mod-AmyLss”) comprises the polynucleotide sequence set forth in SEQ ID NO: 40. More particularly, as schematically presented in FIG.3E, the 2nd copy amy-2 cassette with mod-AmyLss comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-ypuAss (SEQ ID NO: 2) operably linked to DNA encoding the Amy-2 (SEQ ID NO: 37) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21). [0103] As used herein, the term “Amylase” is meant to include any amylase such as glucoamylases, α-amylases, β-amylases and wild-type α-amylases of bacteria such as Bacillus sp., such as B. licheniformis and B. subtilis. In certain embodiments, “Amylase” shall mean an enzyme that is, among other things, capable of catalyzing the degradation of starch. Amylases are hydrolases that cleave the α-D-(1→4) O-glycosidic linkages in starch. Generally, α-amylases (EC 3.2.1.1; α-D-(1→4)-glucan glucanohydrolase)
are defined as endo-acting enzymes cleaving α-D-(1→4) O-glycosidic linkages within the starch molecule in a random fashion. In contrast, the exo-acting amylolytic enzymes, such as β-amylases (EC 3.2.1.2; α- D-(1→4)-glucan maltohydrolase) and some product-specific amylases like maltogenic α-amylase (EC 3.2.1.133) cleave the starch molecule from the non-reducing end of the substrate. β-Amylases, α- glucosidases (EC 3.2.1.20; α-D-glucoside glucohydrolase), glucoamylases (EC 3.2.1.3; α-D-(1→4)-glucan glucohydrolase), and product-specific amylases can produce malto-oligosaccharides of a specific length from starch. [0104] As used herein, the terms “Amylase-1” (abbreviated “Amy-1”) and/or “Amylase-2” (abbreviated “Amy-2”) are not meant to be limiting, but rather refer to exemplary amylase reporter proteins of the disclosure. For example, in certain aspects, the disclosure demonstrates enhanced expression of such exemplary amylase reporters. [0105] As used herein, an “Amylase-1 reporter protein” comprises the mature amino acid sequence set forth in SEQ ID NO: 6. In particular, the Amy-1 protein (SEQ ID NO: 6) and functional variants thereof have been described in PCT Publication No. WO2008/112459 (incorporated herein by reference in its entirety). [0106] As used herein, an “Amylase-2 reporter protein” comprises the amino acid sequence set forth in SEQ ID NO: 37. In particular, the Amy-2 protein (SEQ ID NO: 37) and functional variants thereof have been described in PCT Publication No. WO2023/023642. [0107] As used herein, the polynucleotide (cassette) of SEQ ID NO: 46 comprises an upstream promoter sequence (p3; SEQ ID NO: 10) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (mod-ypuAss; SEQ ID NO: 2) encoding a modified ypuA signal sequence operably linked to DNA (amy-1; SEQ ID NO: 6) encoding an amylase reporter protein. [0108] As used herein, the polynucleotide (cassette) of SEQ ID NO: 47 comprises an upstream promoter sequence (p2; SEQ ID NO: 20) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (mod-ypuAss; SEQ ID NO: 2) encoding a modified ypuA signal sequence operably linked to DNA (amy-1; SEQ ID NO: 6) encoding an amylase reporter protein. [0109] As used herein, the polynucleotide (cassette) of SEQ ID NO: 48 comprises an upstream promoter sequence (p3; SEQ ID NO: 10) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (ypuAss; SEQ ID NO: 1) encoding a native ypuA signal sequence operably linked to DNA (amy-1; SEQ ID NO: 6) encoding an amylase reporter protein. [0110] As used herein, the polynucleotide (cassette) of SEQ ID NO: 49 comprises an upstream promoter sequence (p2; SEQ ID NO: 20) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (ypuAss; SEQ ID NO: 1) encoding a native ypuA signal sequence operably linked to DNA (amy-1; SEQ ID NO: 6) encoding an amylase reporter protein.
[0111] As used herein, the polynucleotide (cassette) of SEQ ID NO: 50 comprises an upstream promoter sequence (p3; SEQ ID NO: 10) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (mod-ypuAss; SEQ ID NO: 2) encoding a modified ypuA signal sequence operably linked to DNA (amy-2; SEQ ID NO: 37) encoding an amylase reporter protein. [0112] As used herein, the polynucleotide (cassette) of SEQ ID NO: 51 comprises an upstream promoter sequence (p2; SEQ ID NO: 20) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (mod-ypuAss; SEQ ID NO: 2) encoding a modified ypuA signal sequence operably linked to DNA (amy-2; SEQ ID NO: 37) encoding an amylase reporter protein. [0113] As used herein, the polynucleotide (cassette) of SEQ ID NO: 52 comprises an upstream promoter sequence (p3; SEQ ID NO: 10) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (ypuAss; SEQ ID NO: 1) encoding a native ypuA signal sequence operably linked to DNA (amy-2; SEQ ID NO: 37) encoding an amylase reporter protein. [0114] As used herein, the polynucleotide (cassette) of SEQ ID NO: 53 comprises an upstream promoter sequence (p2; SEQ ID NO: 20) operably linked to a 5′-UTR sequence (aprE; SEQ ID NO: 11) operably linked to DNA (ypuAss; SEQ ID NO: 1) encoding a native ypuA signal sequence operably linked to DNA (amy-2; SEQ ID NO: 37) encoding an amylase reporter protein. [0115] In certain embodiments, one or more B. licheniformis strains of the disclosure comprise one or more genetic modifications introduced therein. For instance, in certain embodiments, recombinant B. licheniformis (host) cells of the disclosure may comprise one or more genetic modifications (e.g., gene deletions, disruptions, etc.) of one or more endogenous (native) genes encoding unwanted/undesired background proteins (e.g., proteases and the like). [0116] As used herein, a B. licheniformis strain named “BF1965” comprises deletions (Δ) of its endogenous serA and lysA genes (abbreviated, “∆serA/∆lysA”). For example, the B. licheniformis BF1965 (∆serA/∆lysA) strain may be constructed as generally described in PCT Publication No. WO2023/023642. [0117] As used herein, a B. licheniformis strain named “AL223” was constructed from the BF1965 strain (∆serA/∆lysA) and comprises a 1st copy of the amy-1 cassette (SEQ ID NO: 7) with the mod-ypuAss integrated at the serA locus and a 2nd copy of the amy-1 cassette (SEQ ID NO: 17) with the mod-ypuAss integrated at the lysA locus. [0118] As used herein, a B. licheniformis strain named “AL207” was constructed from the BF1965 strain (∆serA/∆lysA) and comprises a 1st copy of the amy-1 cassette (SEQ ID NO: 30) with the mod-AmyLss integrated at the serA locus and a 2nd copy amy-1 cassette (SEQ ID NO: 31) with the mod-AmyLss integrated at the lysA locus.
[0119] As used herein, a B. licheniformis strain named “BF780” was constructed from the BF1965 strain (∆serA/∆lysA) and comprises a 1st copy of the amy-2 cassette (SEQ ID NO: 39) with mod-AmyLss integrated at the serA locus. [0120] As used herein, a B. licheniformis strain named “WS2756” was constructed from the BF780 strain and comprises a 1st copy of the amy-2 cassette (SEQ ID NO: 39) with mod-AmyLss integrated at the serA locus and a 2nd copy of the amy-2 cassette (SEQ ID NO: 38) with native ypuAss integrated at the lysA locus. [0121] As used herein, a B. licheniformis strain named “BF822” was constructed from the BF780 strain and comprises a 1st copy of the amy-2 cassette (SEQ ID NO: 39) with mod-AmyLss integrated at the serA locus and a 2nd copy amy-2 cassette (SEQ ID NO: 40) with mod-AmyLss integrated at the lysA locus. [0122] [0123] As used herein, the genus “Bacillus” includes all species within the genus “Bacillus”’ as known to those of skill in the art, including but not limited to B. subtilis, B. licheniformis, B. lentus, B. brevis, B. stearothermophilus, B. alkalophilus, B. amyloliquefaciens, B. clausii, B. halodurans, B. megaterium, B. coagulans, B. circulans, B. lautus, and B. thuringiensis. It is recognized that the genus Bacillus continues to undergo taxonomical reorganization. Thus, it is intended that the genus include species that have been reclassified, including but not limited to such organisms as B. stearothermophilus, which is now named “Geobacillus stearothermophilus”. [0124] As used herein, a “host cell” refers to a cell that has the capacity to act as a host or expression vehicle for a newly introduced DNA sequence. For example, in certain embodiments of the disclosure, the host cells are Bacillus sp. cells or E. coli cells. [0125] As used herein, a “modified cell” refers to a recombinant cell that comprises at least one genetic modification which is not present in the parent or control cell from which the modified cell was derived or constructed from. [0126] As used herein, when the expression and/or production of a protein of interest (POI) in a recombinant (modified) cell is being compared to the expression and/or production of the same POI in a control or parent cell, it will be understood that the modified and control (parent) cells are grown/cultivated/fermented under the same conditions (e.g., the same conditions such as media, temperature, pH and the like). [0127] As used herein, an “increased amount”, when used in phrases such as “a recombinant cell ‘expresses/produces an increased amount’ of a protein of interest relative to a control (parental) cell”, particularly refers to an “increased amount” of the protein of interest (POI) expressed/produced by the recombinant cell, which “increased amount” is always relative to the control (parent) cell expressing/producing the same POI, wherein the modified and unmodified cells are grown/cultured/fermented under the same conditions.
[0128] As used herein, “increasing” protein production or “increased” protein production is meant an increased amount of protein produced (e.g., a protein of interest). The protein may be produced inside the host cell, or secreted (or transported) into the culture medium. In certain embodiments, the protein of interest is produced (secreted) into the culture medium. Increased protein production may be detected for example, as higher maximal level of protein or enzymatic activity (e.g., such as amylase activity), or total extracellular protein produced as compared to the parental host cell. [0129] As used herein, the terms “modification” and “genetic modification” are used interchangeably and include, but are not limited to: (a) the introduction, substitution, or removal of one or more nucleotides in a gene (or an ORF/CDS thereof), or the introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of the gene or ORF thereof, (b) a gene disruption, (c) a gene conversion, (d) a gene deletion, (e) the down-regulation of a gene, (f) specific mutagenesis and/or (g) random mutagenesis of any one or more the genes disclosed herein. [0130] As used herein, the term “introducing”, as used in phrases such as “introducing into a bacterial cell” or “introducing into a B. licheniformis cell at least one polynucleotide open reading frame (ORF), or a gene thereof, or a vector thereof, includes methods known in the art for introducing polynucleotides into a cell, including, but not limited to protoplast fusion, natural or artificial transformation (e.g., calcium chloride, electroporation), transduction, transfection, conjugation and the like (e.g., see Ferrari et al., 1989). [0131] As used herein, “transformed” or “transformation” mean a cell has been transformed by use of recombinant DNA techniques. Transformation typically occurs by insertion of one or more nucleotide sequences (e.g., a polynucleotide, an ORF or gene) into a cell. The inserted nucleotide sequence may be a heterologous nucleotide sequence (i.e., a sequence that is not naturally occurring in cell that is to be transformed). Transformation therefore generally refers to introducing an exogenous DNA into a host cell so that the DNA is maintained as a chromosomal integrant or a self-replicating extra-chromosomal vector. [0132] As used herein, “transforming DNA”, “transforming sequence”, and “DNA construct” refer to DNA that is used to introduce sequences into a host cell or organism. Transforming DNA is DNA used to introduce sequences into a host cell or organism. The DNA may be generated in vitro by PCR or any other suitable techniques. In some embodiments, the transforming DNA comprises an incoming sequence, while in other embodiments it further comprises an incoming sequence flanked by homology boxes. In yet a further embodiment, the transforming DNA comprises other non-homologous sequences, added to the ends (i.e., stuffer sequences or flanks). The ends can be closed such that the transforming DNA forms a closed circle, such as, for example, insertion into a vector. [0133] As used herein, “disruption of a gene” or a “gene disruption”, are used interchangeably and refer broadly to any genetic modification that substantially prevents a host cell from producing a functional gene product (e.g., a protein). Thus, as used herein, a gene disruption includes, but is not limited to, frameshift
mutations, premature stop codons (i.e., such that a functional protein is not made), substitutions eliminating or reducing activity of the protein internal deletions (such that a functional protein is not made), insertions disrupting the coding sequence, mutations removing the operable link between a native promoter required for transcription and the open reading frame, and the like. [0134] As used herein “an incoming sequence” refers to a DNA sequence that is introduced into the Bacillus sp. chromosome. In some embodiments, the incoming sequence is part of a DNA construct. In other embodiments, the incoming sequence encodes one or more proteins of interest. In some embodiments, the incoming sequence comprises a sequence that may or may not already be present in the genome of the cell to be transformed (i.e., it may be either a homologous or heterologous sequence). In some embodiments, the incoming sequence encodes one or more proteins of interest, a gene, and/or a mutated or modified gene. In alternative embodiments, the incoming sequence encodes a functional wild- type gene or operon, a functional mutant gene or operon, or a nonfunctional gene or operon. In some embodiments, the non-functional sequence may be inserted into a gene to disrupt function of the gene. In another embodiment, the incoming sequence includes a selective marker. In a further embodiment the incoming sequence includes two homology boxes. [0135] As used herein, “homology box” refers to a nucleic acid sequence, which is homologous to a sequence in the Bacillus chromosome. More specifically, a homology box is an upstream or downstream region having between about 80 and 100% sequence identity, between about 90 and 100% sequence identity, or between about 95 and 100% sequence identity with the immediate flanking coding region of a gene or part of a gene to be deleted, disrupted, inactivated, down-regulated and the like, according to the invention. These sequences direct where in the Bacillus chromosome a DNA construct is integrated and directs what part of the Bacillus chromosome is replaced by the incoming sequence. While not meant to limit the present disclosure, a homology box may include about between 1 base pair (bp) to 200 kilobases (kb). Preferably, a homology box includes about between 1 bp and 10.0 kb; between 1 bp and 5.0 kb; between 1 bp and 2.5 kb; between 1 bp and 1.0 kb, and between 0.25 kb and 2.5 kb. A homology box may also include about 10.0 kb, 5.0 kb, 2.5 kb, 2.0 kb, 1.5 kb, 1.0 kb, 0.5 kb, 0.25 kb and 0.1 kb. In some embodiments, the 5' and 3' ends of a selective marker are flanked by a homology box wherein the homology box comprises nucleic acid sequences immediately flanking the coding region of the gene. [0136] As used herein, a host cell “genome”, a bacterial (host) cell “genome”, or a Bacillus sp. (host) cell “genome” includes chromosomal and extrachromosomal genes. [0137] As used herein, the terms “plasmid”, “vector” and “cassette” refer to extrachromosomal elements, often carrying genes which are typically not part of the central metabolism of the cell, and usually in the form of circular double-stranded DNA molecules. Such elements may be autonomously replicating sequences, genome integrating sequences, phage or nucleotide sequences, linear or circular, of a single-
stranded or double-stranded DNA or RNA, derived from any source, in which a number of nucleotide sequences have been joined or recombined into a unique construction which is capable of introducing a promoter fragment and DNA sequence for a selected gene product along with appropriate 3' untranslated sequence into a cell. [0138] As used herein, the term “plasmid” refers to a circular double-stranded (ds) DNA construct used as a cloning vector, and which forms an extrachromosomal self-replicating genetic element in many bacteria and some eukaryotes. In some embodiments, plasmids become incorporated into the genome of the host cell. in some embodiments plasmids exist in a parental cell and are lost in the daughter cell. [0139] A used herein, a “transformation cassette” refers to a specific vector comprising a gene (or ORF thereof), and having elements in addition to the foreign gene that facilitate transformation of a particular host cell. [0140] As used herein, the term “vector” refers to any nucleic acid that can be replicated (propagated) in cells and can carry new genes or DNA segments into cells. Thus, the term refers to a nucleic acid construct designed for transfer between different host cells. Vectors include viruses, bacteriophage, pro-viruses, plasmids, phagemids, transposons, and artificial chromosomes such as YACs (yeast artificial chromosomes), BACs (bacterial artificial chromosomes), PLACs (plant artificial chromosomes), and the like, that are “episomes” (i.e., replicate autonomously or can integrate into a chromosome of a host organism). [0141] An “expression vector” refers to a vector that has the ability to incorporate and express heterologous DNA in a cell. Many prokaryotic and eukaryotic expression vectors are commercially available and know to one skilled in the art. Selection of appropriate expression vectors is within the knowledge of one skilled in the art. [0142] As used herein, the terms “expression cassette” and “expression vector” refer to a nucleic acid construct generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular nucleic acid in a target cell (i.e., these are vectors or vector elements, as described above). The recombinant expression cassette can be incorporated into a plasmid, chromosome, mitochondrial DNA, plastid DNA, virus, or nucleic acid fragment. Typically, the recombinant expression cassette portion of an expression vector includes, among other sequences, a nucleic acid sequence to be transcribed and a promoter. In some embodiments, DNA constructs also include a series of specified nucleic acid elements that permit transcription of a particular nucleic acid in a target cell. In certain embodiments, a DNA construct of the disclosure comprises a selective marker and an inactivating chromosomal or gene or DNA segment as defined herein. [0143] As used herein, a “targeting vector” is a vector that includes polynucleotide sequences that are homologous to a region in the chromosome of a host cell into which the targeting vector is transformed and
that can drive homologous recombination at that region. For example, targeting vectors find use in introducing mutations into the chromosome of a host cell through homologous recombination. In some embodiments, the targeting vector comprises other non-homologous sequences, e.g., added to the ends (i.e., stuffer sequences or flanking sequences). The ends can be closed such that the targeting vector forms a closed circle, such as, for example, insertion into a vector. For example, in certain embodiments, a parental B. licheniformis (host) cell is modified (e.g., transformed) by introducing therein one or more “targeting vectors”. [0144] As used herein, the term “protein of interest” or “POI” refers to a polypeptide of interest that is desired to be expressed in a modified B. licheniformis (daughter) host cell, wherein the POI is preferably expressed at increased levels (i.e., relative to a control cell). Thus, as used herein, a POI may be an enzyme, a substrate-binding protein, a surface-active protein, a structural protein, a receptor protein, and the like. In certain embodiments, a modified cell of the disclosure produces an increased amount of a heterologous protein of interest relative to a control (or parent) cell. In particular embodiments, an increased amount of a protein of interest produced by a modified cell of the disclosure is at least a 0.5% increase, at least a 1.0% increase, at least a 5.0% increase, or a greater than 5.0% increase, relative to the control (or parent) cell. [0145] As used herein, a “gene of interest” or “GOI” refers a nucleic acid sequence (e.g., a polynucleotide, a gene or an ORF) which encodes a POI. A “gene of interest” encoding a “protein of interest” may be a naturally occurring gene, a mutated gene or a synthetic gene. [0146] As used herein, the terms “polypeptide” and “protein” are used interchangeably, and refer to polymers of any length comprising amino acid residues linked by peptide bonds. The conventional one (1) letter or three (3) letter codes for amino acid residues are used herein. The polypeptide may be linear or branched, it may comprise modified amino acids, and it may be interrupted by non-amino acids. The term polypeptide also encompasses an amino acid polymer that has been modified naturally or by intervention; for example, disulfide bond formation, glycosylation, lipidation, acetylation, phosphorylation, or any other manipulation or modification, such as conjugation with a labeling component. Also included within the definition are, for example, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids, etc.), as well as other modifications known in the art. [0147] In certain embodiments, a gene of the instant disclosure encodes a commercially relevant industrial protein of interest, such as an enzyme (e.g., a acetyl esterases, aminopeptidases, amylases, arabinases, arabinofuranosidases, carbonic anhydrases, carboxypeptidases, catalases, cellulases, chitinases, chymosins, cutinases, deoxyribonucleases, epimerases, esterases, α-galactosidases, β-galactosidases, α-glucanases, glucan lysases, endo-β-glucanases, glucoamylases, glucose oxidases, α- glucosidases, β-glucosidases, glucuronidases, glycosyl hydrolases, hemicellulases, hexose oxidases, hydrolases, invertases, isomerases, laccases, lipases, lyases, mannosidases, oxidases, oxidoreductases, pectate lyases, pectin acetyl esterases,
pectin depolymerases, pectin methyl esterases, pectinolytic enzymes, perhydrolases, polyol oxidases, peroxidases, phenoloxidases, phytases, polygalacturonases, proteases, peptidases, rhamno-galacturonases, ribonucleases, transferases, transport proteins, transglutaminases, xylanases, hexose oxidases, and combinations thereof). [0148] As used herein, a “variant” polypeptide refers to a polypeptide that is derived from a parent (or reference) polypeptide by the substitution, addition, or deletion of one or more amino acids, typically by recombinant DNA techniques. Variant polypeptides may differ from a parent polypeptide by a small number of amino acid residues and may be defined by their level of primary amino acid sequence homology/identity with a parent (reference) polypeptide. [0149] Preferably, variant polypeptides have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% amino acid sequence identity with a parent (reference) polypeptide sequence. As used herein, a “variant” polynucleotide refers to a polynucleotide encoding a variant polypeptide, wherein the “variant polynucleotide” has a specified degree of sequence homology/identity with a parent polynucleotide, or hybridizes with a parent polynucleotide (or a complement thereof) under stringent hybridization conditions. Preferably, a variant polynucleotide has at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or even at least 99% nucleotide sequence identity with a parent (reference) polynucleotide sequence. [0150] As used herein, a “mutation” refers to any change or alteration in a nucleic acid sequence. Several types of mutations exist, including point mutations, deletion mutations, silent mutations, frame shift mutations, splicing mutations and the like. Mutations may be performed specifically (e.g., via site directed mutagenesis) or randomly (e.g., via chemical agents, passage through repair minus bacterial strains). [0151] As used herein, in the context of a polypeptide or a sequence thereof, the term “substitution” means the replacement (i.e., substitution) of one amino acid with another amino acid. [0152] As used herein, the term “homology” relates to homologous polynucleotides or polypeptides. If two or more polynucleotides or two or more polypeptides are homologous, this means that the homologous polynucleotides or polypeptides have a “degree of identity” of at least 60%, more preferably at least 70%, even more preferably at least 85%, still more preferably at least 90%, more preferably at least 95%, and most preferably at least 98%. Whether two polynucleotide or polypeptide sequences have a sufficiently high degree of identity to be homologous as defined herein, can suitably be investigated by aligning the two sequences using a computer program known in the art, such as “GAP” provided in the GCG program package (Program Manual for the Wisconsin Package, Version 8, August 1994, Genetics Computer Group, 575 Science Drive, Madison, Wisconsin, USA 53711) (Needleman and Wunsch, (1970). Using GAP with
the following settings for DNA sequence comparison: GAP creation penalty of 5.0 and GAP extension penalty of 0.3. [0153] As used herein, the term “percent (%) identity” refers to the level of nucleic acid or amino acid sequence identity between the nucleic acid sequences that encode a polypeptide or the polypeptide's amino acid sequences, when aligned using a sequence alignment program. [0154] As used herein, “specific productivity” is total amount of protein produced per cell per time over a given time period. [0155] As used herein, the terms “purified”, “isolated” or “enriched” are meant that a biomolecule (e.g., a polypeptide or polynucleotide) is altered from its natural state by virtue of separating it from some, or all of, the naturally occurring constituents with which it is associated in nature. Such isolation or purification may be accomplished by art-recognized separation techniques such as ion exchange chromatography, affinity chromatography, hydrophobic separation, dialysis, protease treatment, ammonium sulphate precipitation or other protein salt precipitation, centrifugation, size exclusion chromatography, filtration, microfiltration, gel electrophoresis or separation on a gradient to remove whole cells, cell debris, impurities, extraneous proteins, or enzymes undesired in the final composition. It is further possible to then add constituents to a purified or isolated biomolecule composition which provide additional benefits, for example, activating agents, anti-inhibition agents, desirable ions, compounds to control pH or other enzymes or chemicals. II. SIGNAL SEQUENCES FOR IMPROVED PROTEIN SECRETION [0156] Generally accepted methods of secreting heterologous proteins often rely on the use of native (protein) signal sequences of similar proteins (e.g., signal sequences from AmyL or AmyE for amylases, or AprE, NprE, or AprL for proteases), or signal sequences that are native to the heterologous protein sequence. For example, protein translation, secretion and folding can be consecutive processes and/or concurrent processes, wherein the protein’s signal (secretion) sequence plays a dynamic role in all three processes. More particularly, as appreciated by one of skill in the art, the use of sub-optimal signal sequences can present particular problems, including among other things, poor or insufficient translation, secretion and/or folding of the heterologous protein, wherein non-optimal pairings of signal sequences to mature protein sequences can lead to bottlenecks in production and secretion of the protein, leading to misfolded/inactive product and/or the induction of cellular stress responses further decreasing the productivity of the host cell. [0157] As set forth herein, Applicant screened B. licheniformis putative signal peptide sequences that were selected by using the SignalP peptide prediction algorithm (Armenteros et al., 2019), and have surprisingly identified a native B. licheniformis ypuA protein signal sequence (SEQ ID NO: 1) that is particularly useful for enhancing/increasing production (secretion) of heterologous proteins of interest. More particularly, as
generally set forth below in the Examples, Applicant constructed recombinant B. licheniformis strains expressing exemplary reporter proteins (i.e., heterologous proteins of interest), wherein the mature (amino acid) sequences of the reporter proteins (e.g., Amy-1 reporter, Amy-2 reporter) were operably linked to upstream (N-terminal) signal peptide (secretion) sequences of the disclosure, such as the native B. licheniformis ypuA signal sequence (ypuAss) of SEQ ID NO: 1. In other embodiments, Applicant constructed a modified (variant) ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) which is particularly useful for enhancing/increasing production (secretion) of heterologous proteins of interest. In particular, the mod-ypuAss (SEQ ID NO: 2) comprises a substitution of the aspartic acid (D24) to a serine (Ser; S24) residue at position 24, see FIG.1A. For example, as described in Tjalsma et al. (2000), serine residues are the most abundant at the -2 position relative to the signal peptidase (SPase I) cleavage site. [0158] Thus, in certain embodiments, Applicant constructed a first (1st) copy of an amylase-1 integration cassette (1st copy amy-1 cassette) and second (2nd) copy of an amylase-1 integration cassette (2nd copy amy- 1 cassette) suitable for integration into a B. licheniformis host cell. As generally described in Example 1, the 1st copy of the amy-1 cassette (FIG. 2A, SEQ ID NO: 7) with a modified ypuA signal sequence (mod- ypuAss; SEQ ID NO: 2) comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream DNA comprising a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding a modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13). As shown in FIG.1A, the mod-ypuAss (SEQ ID NO: 2) as compared to the native ypuA signal sequence (ypuAss; SEQ ID NO: 1) comprises a substitution of an aspartic acid (Asp; D) to a serine (Ser; S) residue at position 24 of SEQ ID NO: 2. [0159] Likewise, as generally described in Example 2, a 2nd copy of the amy-1 cassette (FIG.2C, SEQ ID NO: 17) with the modified ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to the lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (pro; SEQ ID NO: 20) operably linked to DNA encoding the B. subtilis aprE 5′UTR (SEQ ID NO: 11) operably linked to DNA encoding the modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21).
[0160] Thus, the integration (expression) cassettes constructed/described in Example 1 and Example 2 were integrated into exemplary B. licheniformis host cells (strains) set forth and described in Example 3. For instance, a B. licheniformis host strain named “AL223” was constructed by integrating the 1st and 2nd copies of the amy-1 expression cassettes into a parental B. licheniformis strain named “BF1965”, which BF1965 parent strain comprises deletions (∆) of the wild-type serA and lysA genes (∆serA/∆lysA). As further described in Example 3, to test relative performance of the mod-ypuAss (SEQ ID NO: 2) on mature Amy-1 production, a control strain (AL207) with the modified B. licheniformis AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) was constructed in the BF1965 strain by integration of a 1st amy-1 cassette (SEQ ID NO: 30) at the serA locus and integration of a 2nd amy-1 cassette (SEQ ID NO: 31) at the lysA locus. More particularly, the 1st amy-1 cassette with mod-AmyLss (SEQ ID NO: 30) and the 2nd amy-1 cassette with mod-AmyLss (SEQ ID NO: 31) are shown in FIG. 2E and FIG. 2G, respectively. For instance, the native AmyL signal sequence (AmyLss; SEQ ID NO: 3) is generally known to be an optimal signal sequence for producing/secreting many amylase proteins. Likewise, the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) has been shown to further enhance the production/secretion of amylase proteins as compared to the native AmyLss (SEQ ID NO: 3), as generally described in PCT Publication No.2023/023642. In particular, the mod-AmyLss (SEQ ID NO: 4) comprises a serine (S28) at the -2 position (Tjalsma et al., 2000) relative to the native AmyLss (SEQ ID NO: 3), see FIG. 1B.As presented in Example 4, the B. licheniformis AL223 strain with the mod-ypuAss (SEQ ID NO: 2) was assayed for production of the Amy-1 reporter and compared to the B. licheniformis control AL207 strain comprising two (2) integrated copies of the amylase-1 expression cassettes with the modified (control) AmyL signal sequence (mod-AmyLss control; SEQ ID NO: 4) using standard small-scale conditions. For example, the Amy-1 protein production was quantified using the method of Bradford assay, wherein the relative improvement in production of Amy-1 from the AL223 strain is compared to the production of the Amy-1 from the AL207 (control) strain. More particularly, as presented in TABLE 1 (Example 4), the mod-ypuAss (SEQ ID NO: 2) demonstrates an approximately 50% increase in Amy-1 protein production (strain AL223) relative to Amy-1 production (strain AL207) using the mod-AmyLss (SEQ ID NO: 4). [0161] In other embodiments, Applicant constructed certain amylase-2 integration (expression) cassettes suitable for integration into a B. licheniformis host cell. More particularly, as described in Example 5 and shown in FIG. 3, the second (2nd) copy of an amylase-2 expression cassette (FIG. 3C; SEQ ID NO: 38) with a native ypuA signal sequence (ypuAss; SEQ ID NO: 1) comprises an upstream (5′) homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to a downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the native B. licheniformis YpuA signal sequence (ypuAss; SEQ ID NO: 1) operably
linked to DNA encoding the Amy-2 protein (SEQ ID NO: 37) operably linked to a B. licheniformis amyL transcriptional terminator (AmyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21). [0162] Likewise, Example 6 of the disclosure describes the construction of B. licheniformis host cells (strains) expressing two (2) copies of the mature Amy-2 (SEQ ID NO: 37) reporter protein. More particularly, the 2nd copy amylase-2 expression cassette (SEQ ID NO: 38) constructed in Example 5 was integrated into the lysA locus of the B. licheniformis strain named “BF780”, resulting in the modified B. licheniformis host strain named “WS2756” (Example 6). In particular, the BF780 strain was constructed from the parental B. licheniformis BF1965 strain (∆serA/∆lysA), wherein the BF1965 strain comprises an introduced first (1st) copy of an amylase-2 cassette with the mod-AmyLss (FIG. 3A; SEQ ID NO: 39) integrated into the serA locus. As further described in Example 6, the modified B. licheniformis strain WS2756 comprises 2-copies of the amy-2 integration (expression) cassettes, wherein the 1st copy amy-2 cassette integrated at the serA locus (FIG. 3A; SEQ ID NO: 39) and having the mod-AmyLss (SEQ ID NO: 4) and the 2nd copy amy-2 cassette integrated into the lysA locus (FIG.3C; SEQ ID NO: 38) and having the native ypuAss (SEQ ID NO: 1). To further test and evaluate the relative performance of the native ypuA signal sequence (ypuAss; SEQ ID NO: 1) on Amy-2 (SEQ ID NO: 37) reporter production, a control strain named “BF822” with the mod-AmyLss (SEQ ID NO: 4) was constructed in the BF780 strain by the integration of the amy-2 cassette (SEQ ID NO: 40) at the lysA locus. (FIG.3E). Thus, the BF822 control strain described in Example 6 comprises 2-copies of the amy-2 integration (expression) cassettes, wherein the 1st copy amy-2 cassette (FIG.3A; SEQ ID NO: 39) is integrated at the serA locus and having the mod- AmyLss (SEQ ID NO: 4) and the 2nd copy amy-2 cassette (FIG.3E; SEQ ID NO: 40) is integrated into the lysA locus and having the mod-AmyLss (SEQ ID NO: 4). [0163] As set forth and described in Example 7, the WS2756 strain and BF822 (control) strain constructed in Example 6 were assayed for production of the Amy-2 (SEQ ID NO: 37) reporter using standard small- scale conditions. More particularly, as shown in TABLE 2, the WS2756 strain (with native ypuAss; SEQ ID NO: 1) demonstrates an approximately 18% increase in Amy-2 reporter production as compared to the control BF822 strain (with mod-AmyLss; SEQ ID NO: 4). [0164] Thus, certain embodiments of the disclosure provide, inter alia, nucleic acids, polynucleotides, vectors, expression cassettes, regulatory elements, and the like, suitable for use in constructing recombinant (modified) Bacillus host cells. Certain aspects are therefore related to polynucleotides (e.g., expression cassettes) comprising an upstream (5′) promoter (pro) sequence operably linked to a downstream nucleic acid sequence (ss) encoding a protein signal (secretion) sequence operably linked to a downstream nucleic acid sequence (poi) encoding a protein of interest.
[0165] For example, a generic polynucleotide sequence encoding an amino (N) terminal signal sequence in operable combination with a mature protein of interest (POI) is shown below in Scheme 1: Scheme 1: 5′-[ss]-[poi]-3′ wherein the nucleic acid (ss) sequence encoding the N-terminal signal sequence (SS) is upstream and operably linked to a downstream nucleic acid (poi) sequence encoding a mature protein of interest (POI). [0166] In certain embodiments, polynucleotide expression cassettes may be described generically as shown in Scheme 2: Scheme 2: 5′-[pro]-[ss]-[poi]-3′ wherein the promoter (pro) sequence is upstream (5′) and operably linked to a nucleic acid (ss) sequence encoding the N-terminal signal sequence (SS), which is upstream and operably linked to a nucleic acid (poi) sequence encoding a protein of interest (POI). In certain other embodiments, the polynucleotide may further comprise a terminator (term) sequence downstream and operably linked to the nucleic acid (poi) sequence encoding the mature POI. [0167] In certain embodiments, the disclosure relates to polynucleotide constructs (e.g., Scheme 2) encoding a heterologous POI (e.g., a mature amylase protein sequence), wherein the nucleic acid (DNA) encoding the POI is operably linked to an upstream nucleic acid encoding an N-terminal ypuA signal sequence (ypuAss; SEQ ID NO: 1) and/or operably linked to an upstream nucleic acid encoding an N- terminal modified (variant) ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2). [0168] Thus, certain other embodiments of the disclosure provide recombinant (modified) B. licheniformis strains/cells comprising one or more introduced polynucleotide constructs (e.g., expression cassettes) encoding one or more mature amylases comprising an N-terminal ypuA signal sequence (SED ID NO: 1) and/or encoding one or more mature amylases comprising an N-terminal modified (variant) ypuA signal sequence (SED ID NO: 2). In particular, as presented in FIG. 1, the native B. licheniformis ypuA signal sequence (FIG. 1A, ypuAss; SEQ ID NO: 1) comprises twenty-five (25) amino acid residues with an aspartic acid (Asp; D24) residue at position 24, whereas the modified (variant) ypuA signal sequence (FIG. 1B, mod-YpuAss; SED IDNO: 2) comprises a substitution of the aspartic acid (D24) to a serine (Ser; S24) residue at position 24. Thus, in certain embodiments, the amino acid positions of a particular protein signal sequence may be described and numbered from the amino-terminus (NH2), as indicated in FIG. 1. Alternatively, the amino acid positions may be described and numbered according to the cleavage site of a particular signal sequence. For example, the most C-terminal amino acid position of the ypuAss (SEQ ID NO: 1, residue A25) can be designated with a negative 1 (-1) amino acid (position), the amino acid position to its left a negative 2 (-2), etc.. Likewise, other native and/or modified signal sequences of the disclosure may be designated with similar specificity.
[0169] Thus, certain other embodiments of the disclosure provide recombinant (modified) B. licheniformis cells (strains) capable of secreting enhanced amounts of amylase proteins. III. RECOMBINANT POLYNUCLEOTIDES AND MOLECULAR BIOLOGY [0170] As generally set forth above and further described below in the Examples, certain embodiments of the disclosure are related to recombinant (modified) Bacillus cells capable of producing increased amounts of heterologous proteins of interest. Certain embodiments are therefore related to methods for constructing such recombinant Bacillus cells having increased protein production capabilities. In certain embodiments, one or more expression cassettes encoding a protein of intertest are introduced into Bacillus cells of the disclosure. In exemplary embodiments, the cassettes are integrated into the genome of the cell. For example, in certain exemplary embodiments, expression cassettes encoding a protein of interest were integrated into the lysA locus and the serA locus of a parental B. licheniformis cell comprising deletions of its native (endogenous) lysA (∆lysA) and serA (∆serA) genes. Thus, certain embodiments are related to, among other things, nucleic acids, polynucleotides (e.g., vectors, expression cassettes), regulatory elements, and the like, suitable for use in constructing recombinant (modified) Bacillus host cells. [0171] In certain other embodiments, Bacillus cells of the disclosure are rendered deficient in the production of one or more native (endogenous) genes. In certain embodiments, Bacillus cells of the disclosure are rendered deficient in the production of one or more native (endogenous) proteases. For example, in certain embodiments, a host cell of the disclosure is a Bacillus licheniformis cell deficient in the production of one or more native proteases selected from the group consisting of wprA, nprE, mpr, aprL, bprE, htrA, vpr and ispA. [0172] Accordingly, as presented in the Examples and generally described herein, recombinant cells of the disclosure may be constructed by one of skill using standard and routine recombinant DNA and molecular cloning techniques well known in the art. Methods for genetically modifying cells include, but are not limited to, (a) the introduction, substitution, or removal of one or more nucleotides in a gene, or the introduction, substitution, or removal of one or more nucleotides in a regulatory element required for the transcription or translation of the gene, (b) a gene disruption, (c) a gene conversion, (d) a gene deletion, (e) a gene down-regulation, (f) site specific mutagenesis and/or (g) random mutagenesis. [0173] In certain embodiments, modified cells of the disclosure may be constructed by reducing or eliminating the expression of a gene, using methods well known in the art, for example, insertions, disruptions, replacements, or deletions. The portion of the gene to be modified or inactivated may be, for example, the coding region or a regulatory element required for expression of the coding region.
[0174] An example of such a regulatory or control sequence may be a promoter sequence or a functional part thereof, (i.e., a part which is sufficient for affecting expression of the nucleic acid sequence). Other control sequences for modification include, but are not limited to, a leader sequence, a pro-peptide sequence, a signal sequence, a transcription terminator, a transcriptional activator and the like. [0175] In certain other embodiments a modified cell is constructed by gene deletion to eliminate or reduce the expression of the gene. Gene deletion techniques enable the partial or complete removal of the gene(s), thereby eliminating their expression, or expressing a non-functional (or reduced activity) protein product. In such methods, the deletion of the gene(s) may be accomplished by homologous recombination using a plasmid that has been constructed to contiguously contain the 5' and 3' regions flanking the gene. The contiguous 5' and 3' regions may be introduced into a Bacillus cell, for example, on a temperature-sensitive plasmid, such as pE194, in association with a second selectable marker at a permissive temperature to allow the plasmid to become established in the cell. The cell is then shifted to a non-permissive temperature to select for cells that have the plasmid integrated into the chromosome at one of the homologous flanking regions. Selection for integration of the plasmid is affected by selection for the second selectable marker. After integration, a recombination event at the second homologous flanking region is stimulated by shifting the cells to the permissive temperature for several generations without selection. The cells are plated to obtain single colonies and the colonies are examined for loss of both selectable markers. Thus, a person of skill in the art may readily identify nucleotide regions in the gene’s coding sequence and/or the gene’s non- coding sequence suitable for complete or partial deletion. [0176] In other embodiments, a modified cell is constructed by introducing, substituting, or removing one or more nucleotides in the gene or a regulatory element required for the transcription or translation thereof. For example, nucleotides may be inserted or removed so as to result in the introduction of a stop codon, the removal of the start codon, or a frame-shift of the open reading frame. Such a modification may be accomplished by site-directed mutagenesis or PCR generated mutagenesis in accordance with methods known in the art. Thus, in certain embodiments, a gene of the disclosure is inactivated by complete or partial deletion. [0177] In another embodiment, a modified cell is constructed by the process of gene conversion. For example, in the gene conversion method, a nucleic acid sequence corresponding to the gene(s) is mutagenized in vitro to produce a defective nucleic acid sequence, which is then transformed into the parental Bacillus cell to produce a defective gene. By homologous recombination, the defective nucleic acid sequence replaces the endogenous gene. It may be desirable that the defective gene or gene fragment also encodes a marker which may be used for selection of transformants containing the defective gene. For example, the defective gene may be introduced on a non-replicating or temperature-sensitive plasmid in association with a selectable marker. Selection for integration of the plasmid is affected by selection for
the marker under conditions not permitting plasmid replication. Selection for a second recombination event leading to gene replacement is affected by examination of colonies for loss of the selectable marker and acquisition of the mutated gene. Alternatively, the defective nucleic acid sequence may contain an insertion, substitution, or deletion of one or more nucleotides of the gene, as described below. [0178] In other embodiments, a modified cell is constructed by established anti-sense techniques using a nucleotide sequence complementary to the nucleic acid sequence of the gene. More specifically, expression of the gene by a Bacillus cell may be reduced (down-regulated) or eliminated by introducing a nucleotide sequence complementary to the nucleic acid sequence of the gene, which may be transcribed in the cell and is capable of hybridizing to the mRNA produced in the cell. Under conditions allowing the complementary anti-sense nucleotide sequence to hybridize to the mRNA, the amount of protein translated is thus reduced or eliminated. Such anti-sense methods include, but are not limited to RNA interference (RNAi), small interfering RNA (siRNA), microRNA (miRNA), antisense oligonucleotides, and the like, all of which are well known to the skilled artisan. [0179] In other embodiments, a modified cell is produced/constructed via CRISPR-Cas9 editing. For example, a gene encoding a protein of interest can be edited or disrupted (or deleted or down-regulated) by means of nucleic acid guided endonucleases, that find their target DNA by binding either a guide RNA (e.g., Cas9) and Cpf1 or a guide DNA (e.g., NgAgo), which recruits the endonuclease to the target sequence on the DNA, wherein the endonuclease can generate a single or double stranded break in the DNA. This targeted DNA break becomes a substrate for DNA repair, and can recombine with a provided editing template to disrupt or delete the gene. For example, the gene encoding the nucleic acid guided endonuclease (for this purpose Cas9 from S. pyogenes), or a codon optimized gene encoding the Cas9 nuclease is operably linked to a promoter active in the Bacillus cell and a terminator active in Bacillus cell, thereby creating a Bacillus Cas9 expression cassette. Likewise, one or more target sites unique to the gene of interest are readily identified by a person skilled in the art. For example, to build a DNA construct encoding a gRNA -directed to a target site within the gene of interest, the variable targeting domain (VT) will comprise nucleotides of the target site which are 5′ of the (PAM) proto-spacer adjacent motif (TGG), which nucleotides are fused to DNA encoding the Cas9 endonuclease recognition domain for S. pyogenes Cas9 (CER). The combination of the DNA encoding a VT domain and the DNA encoding the CER domain thereby generate a DNA encoding a gRNA. Thus, a Bacillus expression cassette for the gRNA is created by operably linking the DNA encoding the gRNA to a promoter active in Bacillus cells and a terminator active in Bacillus cells. [0180] In certain embodiments, the DNA break induced by the endonuclease is repaired/replaced with an incoming sequence. For example, to precisely repair the DNA break generated by the Cas9 expression cassette and the gRNA expression cassette described above, a nucleotide editing template is provided, such
that the DNA repair machinery of the cell can utilize the editing template. For example, about 500bp 5′ of targeted gene can be fused to about 500bp 3′ of the targeted gene to generate an editing template, which template is used by the Bacillus host’s machinery to repair the DNA break generated by the RGEN. [0181] The Cas9 expression cassette, the gRNA expression cassette and the editing template can be co- delivered to filamentous fungal cells using many different methods (e.g., protoplast fusion, electroporation, natural competence, or induced competence). The transformed cells are screened by PCR amplifying the target gene locus, by amplifying the locus with a forward and reverse primer. These primers can amplify the wild-type locus or the modified locus that has been edited by the RGEN. These fragments are then sequenced using a sequencing primer to identify edited colonies. [0182] In yet other embodiments, a modified cell is constructed by random or specific mutagenesis using methods well known in the art, including, but not limited to, chemical mutagenesis and transposition. Modification of the gene may be performed by subjecting the parental cell to mutagenesis and screening for mutant cells in which expression of the gene has been reduced or eliminated. The mutagenesis, which may be specific or random, may be performed, for example, by use of a suitable physical or chemical mutagenizing agent, use of a suitable oligonucleotide, or subjecting the DNA sequence to PCR generated mutagenesis. Furthermore, the mutagenesis may be performed by use of any combination of these mutagenizing methods. [0183] Examples of a physical or chemical mutagenizing agent suitable for the present purpose include ultraviolet (UV) irradiation, hydroxylamine, N-methyl-N'-nitro-N-nitrosoguanidine (MNNG), N-methyl- N'-nitrosoguanidine (NTG), O-methyl hydroxylamine, nitrous acid, ethyl methane sulphonate (EMS), sodium bisulphite, formic acid, and nucleotide analogues. When such agents are used, the mutagenesis is typically performed by incubating the parental cell to be mutagenized in the presence of the mutagenizing agent of choice under suitable conditions, and selecting for mutant cells exhibiting reduced or no expression of the gene. [0184] PCT Publication No. WO2003/083125 discloses methods for modifying Bacillus cells, such as the creation of Bacillus deletion strains and DNA constructs using PCR fusion to bypass E. coli. PCT Publication No. WO2002/14490 discloses methods for modifying Bacillus cells including (1) the construction and transformation of an integrative plasmid (pComK), (2) random mutagenesis of coding sequences, signal sequences and pro-peptide sequences, (3) homologous recombination, (4) increasing transformation efficiency by adding non-homologous flanks to the transformation DNA, (5) optimizing double cross-over integrations, (6) site directed mutagenesis and (7) marker-less deletion. [0185] Those of skill in the art are well aware of suitable methods for introducing polynucleotide sequences into bacterial cells (e.g., E. coli and Bacillus). Indeed, such methods as transformation including protoplast transformation and congression, transduction, and protoplast fusion are known and suited for use
in the present disclosure. Methods of transformation are particularly preferred to introduce a DNA construct of the present disclosure into a host cell. [0186] In addition to commonly used methods, in some embodiments, host cells are directly transformed (i.e., an intermediate cell is not used to amplify, or otherwise process, the DNA construct prior to introduction into the host cell). Introduction of the DNA construct into the host cell includes those physical and chemical methods known in the art to introduce DNA into a host cell, without insertion into a plasmid or vector. Such methods include, but are not limited to, calcium chloride precipitation, electroporation, naked DNA, liposomes and the like. In additional embodiments, DNA constructs are co-transformed with a plasmid without being inserted into the plasmid. In further embodiments, a selective marker is deleted or substantially excised from the modified Bacillus strain by methods known in the art. In some embodiments, resolution of the vector from a host chromosome leaves the flanking regions in the chromosome, while removing the indigenous chromosomal region. [0187] Promoters and promoter sequence regions for use in the expression of genes, open reading frames (ORFs) thereof and/or variant sequences thereof in Bacillus cells are generally known on one of skill in the art. Promoter sequences of the disclosure are generally chosen so that they are functional in the Bacillus cells (e.g., B. licheniformis cells, B. subtilis cells and the like). For example, promoters useful for driving gene expression in Bacillus cells include, but are not limited to, the B. subtilis alkaline protease (aprE) promoter, the α-amylase promoter (amyE) of B. subtilis, the α-amylase promoter (amyL) of B. licheniformis, the α-amylase promoter of B. amyloliquefaciens, the neutral protease (nprE) promoter from B. subtilis, a mutant aprE promoter, or any other promoter from B licheniformis or other related Bacilli. Methods for screening and creating promoter libraries with a range of activities (promoter strength) in Bacillus cells is describe in Publication No. WO2002/14490. IV. FERMENTING BACILLUS CELLS FOR PRODUCTION OF A PROTEIN OF INTEREST [0188] As generally described above, certain embodiments are related to compositions and methods for constructing and obtaining Bacillus cells having increased protein production phenotypes. Thus, certain embodiments are related to methods of producing proteins of interest in Bacillus cells by fermenting the cells in a suitable medium. Fermentation methods well known in the art can be applied to ferment Bacillus cells of the disclosure. [0189] In some embodiments, the cells are cultured under batch or continuous fermentation conditions. A classical batch fermentation is a closed system, where the composition of the medium is set at the beginning of the fermentation and is not altered during the fermentation. At the beginning of the fermentation, the medium is inoculated with the desired organism(s). In this method, fermentation is permitted to occur without the addition of any components to the system. Typically, a batch fermentation qualifies as a “batch” with respect to the addition of the carbon source, and attempts are often made to control factors such as pH
and oxygen concentration. The metabolite and biomass compositions of the batch system change constantly up to the time the fermentation is stopped. Within typical batch cultures, cells can progress through a static lag phase to a high growth log phase, and finally to a stationary phase, where growth rate is diminished or halted. If untreated, cells in the stationary phase eventually die. In general, cells in log phase are responsible for the bulk of production of product. [0190] A suitable variation on the standard batch system is the “fed-batch” fermentation system. In this variation of a typical batch system, the substrate is added in increments as the fermentation progresses. Fed-batch systems are useful when catabolite repression likely inhibits the metabolism of the cells and where it is desirable to have limited amounts of substrate in the medium. Measurement of the actual substrate concentration in fed-batch systems is difficult and is therefore estimated on the basis of the changes of measurable factors, such as pH, dissolved oxygen and the partial pressure of waste gases, such as CO2. Batch and fed-batch fermentations are common and known in the art. [0191] Continuous fermentation is an open system where a defined fermentation medium is added continuously to a bioreactor, and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the cultures at a constant high density, where cells are primarily in log phase growth. Continuous fermentation allows for the modulation of one or more factors that affect cell growth and/or product concentration. For example, in one embodiment, a limiting nutrient, such as the carbon source or nitrogen source, is maintained at a fixed rate and all other parameters are allowed to moderate. In other systems, a number of factors affecting growth can be altered continuously while the cell concentration, measured by media turbidity, is kept constant. Continuous systems strive to maintain steady state growth conditions. Thus, cell loss due to medium being drawn off should be balanced against the cell growth rate in the fermentation. Methods of modulating nutrients and growth factors for continuous fermentation processes, as well as techniques for maximizing the rate of product formation, are well known in the art of industrial microbiology. [0192] In certain embodiments, a protein of interest expressed/produced by a Bacillus cell of the disclosure may be recovered from the culture medium by conventional procedures including separating the host cells from the medium by centrifugation or filtration, or if necessary, disrupting the cells and removing the supernatant from the cellular fraction and debris. Typically, after clarification, the proteinaceous components of the supernatant or filtrate are precipitated by means of a salt, e.g., ammonium sulfate. The precipitated proteins are then solubilized and may be purified by a variety of chromatographic procedures, e.g., ion exchange chromatography, gel filtration. [0193] In some embodiments, the cells are cultured under batch or continuous fermentation conditions. A classical batch fermentation is a closed system, where the composition of the medium is set at the beginning of the fermentation and is not altered during the fermentation. At the beginning of the fermentation, the
medium is inoculated with the desired organism(s). In this method, fermentation is permitted to occur without the addition of any components to the system. Typically, a batch fermentation qualifies as a “batch” with respect to the addition of the carbon source, and attempts are often made to control factors such as pH and oxygen concentration. The metabolite and biomass compositions of the batch system change constantly up to the time the fermentation is stopped. Within typical batch cultures, cells can progress through a static lag phase to a high growth log phase, and finally to a stationary phase, where growth rate is diminished or halted. If untreated, cells in the stationary phase eventually die. In general, cells in log phase are responsible for the bulk of production of product. [0194] A suitable variation on the standard batch system is the “fed-batch” fermentation system. In this variation of a typical batch system, the substrate is added in increments as the fermentation progresses. Fed-batch systems are useful when catabolite repression likely inhibits the metabolism of the cells and where it is desirable to have limited amounts of substrate in the medium. Measurement of the actual substrate concentration in fed-batch systems is difficult and is therefore estimated on the basis of the changes of measurable factors, such as pH, dissolved oxygen and the partial pressure of waste gases, such as CO2. Batch and fed-batch fermentations are common and known in the art. [0195] Continuous fermentation is an open system where a defined fermentation medium is added continuously to a bioreactor, and an equal amount of conditioned medium is removed simultaneously for processing. Continuous fermentation generally maintains the cultures at a constant high density, where cells are primarily in log phase growth. Continuous fermentation allows for the modulation of one or more factors that affect cell growth and/or product concentration. For example, in one embodiment, a limiting nutrient, such as the carbon source or nitrogen source, is maintained at a fixed rate and all other parameters are allowed to moderate. In other systems, a number of factors affecting growth can be altered continuously while the cell concentration, measured by media turbidity, is kept constant. Continuous systems strive to maintain steady state growth conditions. Thus, cell loss due to medium being drawn off should be balanced against the cell growth rate in the fermentation. Methods of modulating nutrients and growth factors for continuous fermentation processes, as well as techniques for maximizing the rate of product formation, are well known in the art of industrial microbiology. [0196] In certain embodiments, a protein of interest expressed/produced by a Bacillus cell of the disclosure may be recovered from the culture medium by conventional procedures including separating the host cells from the medium by centrifugation or filtration, or if necessary, disrupting the cells and removing the supernatant from the cellular fraction and debris. Typically, after clarification, the proteinaceous components of the supernatant or filtrate are precipitated by means of a salt, e.g., ammonium sulfate. The precipitated proteins are then solubilized and may be purified by a variety of chromatographic procedures, e.g., ion exchange chromatography, gel filtration.
VI. PROTEINS OF INTEREST [0197] A protein of interest (POI) of the instant disclosure can be any endogenous or heterologous protein, and it may be a variant of such a POI. The protein can contain one or more disulfide bridges or is a protein whose functional form is a monomer or a multimer, i.e., the protein has a quaternary structure and is composed of a plurality of identical (homologous) or non-identical (heterologous) subunits, wherein the POI or a variant POI thereof is preferably one with properties of interest. [0198] For example, in certain embodiments, a recombinant (modified) Bacillus cell of the disclosure produces at least about 0.1% more, at least about 0.5% more, at least about 1% more, at least about 5% more, at least about 6% more, at least about 7% more, at least about 8% more, at least about 9% more, or at least about 10% or more of a POI, relative to a control (or parent) Bacillus cell. [0199] In certain embodiments, a modified Bacillus cell of the disclosure exhibits an increased specific productivity (Qp) of a POI relative the control (or parent) Bacillus cell. For example, the detection of specific productivity (Qp) is a suitable method for evaluating protein production. The specific productivity (Qp) can be determined using the following equation: “Qp = gP/gDCW•hr” wherein, “gP” is grams of protein produced in the tank; “gDCW” is grams of dry cell weight (DCW) in the tank and “hr” is fermentation time in hours from the time of inoculation, which includes the time of production as well as growth time. [0200] Thus, in certain other embodiments, a modified Bacillus cell of the disclosure comprises a specific productivity (Qp) increase of at least about 0.1%, at least about 1%, at least about 5%, at least about 6%, at least about 7%, at least about 8%, at least about 9%, or at least about 10% or more, relative to the control (or parent) Bacillus cell. [0201] In certain embodiments, a POI or a variant POI thereof is selected from the group consisting of acetyl esterases, aminopeptidases, amylases, arabinases, arabinofuranosidases, carbonic anhydrases, carboxypeptidases, catalases, cellulases, chitinases, chymosins, cutinases, deoxyribonucleases, epimerases, esterases, α-galactosidases, β-galactosidases, α-glucanases, glucan lysases, endo-β-glucanases, glucoamylases, glucose oxidases, α-glucosidases, β-glucosidases, glucuronidases, glycosyl hydrolases, hemicellulases, hexose oxidases, hydrolases, invertases, isomerases, laccases, ligases, lipases, lyases, mannosidases, oxidases, oxidoreductases, pectate lyases, pectin acetyl esterases, pectin depolymerases, pectin methyl esterases, pectinolytic enzymes, perhydrolases, polyol oxidases, peroxidases, phenoloxidases, phytases, polygalacturonases, proteases, peptidases, rhamno-galacturonases, ribonucleases, transferases, transport proteins, transglutaminases, xylanases, hexose oxidases, and combinations thereof. In certain preferred embodiments, a POI is an amylase.
VI. EXEMPLARY EMBODIMENTS [0202] Non-limiting embodiments of compositions and methods disclosed herein are as follows: [0203] 1. An isolated nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) comprising SEQ ID NO: 2. [0204] 2. A polynucleotide comprising an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) comprising SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI). [0205] 3. The polynucleotide of embodiment 2, comprising an upstream 5′-UTR sequence operably linked to the nucleic acid encoding the ypuAss. [0206] 4. A polynucleotide comprising an upstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) comprising SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI). [0207] 5. The polynucleotide of embodiment 4, comprising an upstream 5′-UTR sequence operably linked to the nucleic acid encoding the mod-ypuAss. [0208] 6. The polynucleotide of embodiment 5, wherein the 5′-UTR sequence comprises at least 95% identity to the B. subtilis aprE 5′-UTR sequence of SEQ ID NO: 11. [0209] 7. The polynucleotide of embodiment 2 or embodiment 4, comprising a downstream terminator sequence operably linked to the nucleic acid encoding the POI. [0210] 8. The polynucleotide of embodiment 7, wherein the terminator sequence comprises at least 95% identity to the B. licheniformis amyL terminator sequence of SEQ ID NO: 12. [0211] 9. The polynucleotide of embodiment 2 or embodiment 4, wherein the heterologous POI is an amylase. [0212] 10. The polynucleotide of embodiment 9, wherein the amylase comprises at least 90% identity to SEQ ID NO: 6 or SEQ ID NO: 37. [0213] 11. An expression cassette comprising an upstream promoter operably linked to a downstream polynucleotide of any one of embodiments 2-10, wherein the promoter is functional in a Bacillus sp. cell. [0214] 12. The cassette of embodiment 11, wherein the promoter comprises at least 95% identity to SEQ ID NO: 10 or SEQ ID NO: 20. [0215] 13. An expression cassette comprising at least 90% identity to any one of SEQ ID NO: 46-53. [0216] 14. The expression cassette of any one of embodiments 11-13, introduced into a Bacillus sp. cell. [0217] 15. At least two expression cassettes of any one of embodiments 11-13, introduced into a Bacillus sp. cell.
[0218] 16. A modified Bacillus sp. cell comprising an introduced polynucleotide comprising an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI), or comprising an introduced polynucleotide comprising an upstream nucleic acid encoding a modified ypuA signal sequence (mod- ypuAss) of SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous POI. [0219] 17. A modified Bacillus sp. cell comprising at least two introduced polynucleotides, wherein the first polynucleotide comprises an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI), or comprising an upstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous POI, and the second polynucleotide comprises an upstream nucleic acid encoding a ypuAss of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous POI, or comprising an upstream nucleic acid encoding a mod-ypuAss of SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous POI. [0220] 18. The modified cell of embodiment 16, wherein the polynucleotide comprises an upstream 5′- UTR sequence operably linked to the nucleic acid encoding the ypuAss or mod-ypuAss. [0221] 19. The modified cell of embodiment 17, wherein the first and second polynucleotides comprise an upstream 5′-UTR sequence operably linked to the nucleic acid encoding the ypuAss and/or mod-ypuAss. [0222] 20. The modified cell of embodiment 18 or embodiment 19, wherein the 5′-UTR sequence comprises at least 95% identity to the B. subtilis aprE 5′-UTR sequence of SEQ ID NO: 11. [0223] 21. The modified cell of embodiment 16, wherein the polynucleotide comprises a downstream terminator sequence operably linked to the nucleic acid encoding the heterologous POI. [0224] 22. The modified cell of embodiment 17, wherein the first and second polynucleotides comprise a downstream terminator sequence operably linked to the nucleic acid encoding the heterologous POI. [0225] 22. The modified cell of embodiment 21 or embodiment 22, wherein the terminator sequence comprises at least 95% identity to the B. licheniformis amyL terminator sequence of SEQ ID NO: 12 [0226] 23. The modified cell of embodiment 16 or embodiment 17, wherein the heterologous POI is an amylase. [0227] 24. The modified cell of embodiment 23, wherein the amylase comprises at least 90% identity to SEQ ID NO: 6 or SEQ ID NO: 37. [0228] 25. A modified Bacillus sp. cell comprising an introduced expression cassette of any one of embodiments 11-13. [0229] 26. A modified Bacillus sp. cell comprising at least two introduced expression cassettes of any one of embodiments 11-13.
[0230] 27. A method for producing a heterologous protein of interest (POI) in a Bacillus sp. cell comprising (a) introducing into a Bacillus sp. cell an expression cassette comprising an upstream (5′) promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding a heterologous POI, and (b) fermenting the modified cell under suitable conditions for the production of the POI. [0231] 28. The method of embodiment 27, wherein the modified cell produces an increased amount of POI relative to a control cell producing the same POI when fermented under the same conditions, wherein the control cell comprises an introduced expression cassette comprising the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a modified AmyL signal sequence (mod-AmyLss) of SEQ ID NO: 4 operably linked to a downstream nucleic acid encoding the heterologous POI. [0232] 29. A method for producing a heterologous protein of interest (POI) in a Bacillus sp. cell comprising (a) introducing into a Bacillus sp. cell at least two expression cassettes, wherein (i) the first cassette comprises an upstream (5′) promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI, and (ii) the second cassette comprises an upstream promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a ypuAss of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a mod-ypuAss of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI, and (b) fermenting the modified cell under suitable conditions for the production of the POI. [0233] 30. The method of embodiment 29, wherein the modified cell produces an increased amount of POI relative to a control cell producing the same POI when fermented under the same conditions, wherein the control cell comprises at least two introduced expression cassettes, wherein (i) the first cassette comprises the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI, and (ii) the second cassette comprises the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1,
or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI. [0234] 31. The method of embodiment 27 or embodiment 29, wherein the POI is secreted into the fermentation broth when fermented under suitable conditions [0235] 32. The method of embodiment 27 or embodiment 29, wherein the POI is an amylase. [0236] 33. The method of embodiment 32, wherein the amylase comprises at least 90% identity to SEQ ID NO: 6 or SEQ ID NO: 37. [0237] 34. The method of embodiment 27, wherein the increased amount of POI produced by the modified cell is at least about 10% increased as compared the same POI produced by the control cell. [0238] 35. The method of embodiment 29, wherein the increased amount of POI produced by the modified cell is at least about 40% increased as compared the same POI produced by the control cell. [0239] 36. The method of embodiment 27 or embodiment 29, wherein the 5′-UTR sequence comprises at least 95% identity to the B. subtilis aprE 5′-UTR sequence of SEQ ID NO: 11. [0240] 37. The method of embodiment 27 or embodiment 29, wherein the polynucleotide further comprises a downstream terminator sequence operably linked to the nucleic acid encoding the POI. [0241] 38. The method of embodiment 37, wherein the terminator sequence comprises at least 95% identity to the B. licheniformis amyL terminator sequence of SEQ ID NO: 12. [0242] 39. The method of embodiment 27 or embodiment 29, wherein the promoter comprises at least 95% identity to SEQ ID NO: 10 or SEQ ID NO: 20. [0243] 40. The method of embodiment 27, wherein cassette comprises at least 90% identity to any one of SEQ ID NO: 46-53. [0244] 41. The method of embodiment 29, wherein the first and second cassettes comprise at least 90% identity to any one of SEQ ID NO: 46-53. [0245] 42. The method of embodiment 27 or embodiment 29, wherein the Bacillus sp. cell is selected from the group consisting of a B. subtilis cell, a B. licheniformis cell, a B. lentus cell, a B. brevis cell, a B. stearothermophilus cell, a B. alkalophilus cell, a B. amyloliquefaciens cell, a B. clausii cell, a B. halodurans cell, a B. megaterium cell, a B. coagulans cell, a B. circulans cell, a B. lautus cell and a B. thuringiensis cell.
EXAMPLES [0246] Certain aspects of the present disclosure may be further understood in light of the following examples, which should not be construed as limiting. Modifications to materials and methods will be apparent to those skilled in the art. Standard recombinant DNA and molecular cloning techniques used herein are well known in the art (Ausubel et al., 1987; Sambrook et al., 1989). EXAMPLE 1 CONSTRUCTION OF A TEMPLATE PLASMID FOR A FIRST COPY OF AN AMYLASE-1 INTEGRATION CASSETTE [0247] The instant example describes construction of a template DNA (SEQ ID NO: 5) by overlapping extension PCR for the integration of a first (1st) copy of an Amylase-1 (Amy-1; SEQ ID NO: 6) expression cassette. For example, the 1st copy of the amy-1 cassette with mod-ypuAss (1st copy amy-1 cassette mod- ypuAss; SEQ ID NO: 7) comprises an upstream (5′) homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to downstream DNA comprising a B. licheniformis serA ORF (SEQ ID NO: 9) operably linked to a synthetic p3 promoter (p3 pro; SEQ ID NO: 10) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding a modified B. licheniformis ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13). More particularly, as shown in FIG.1A, the modified ypuA signal sequence (mod-ypuAss; SEQ ID NO: 2) as compared to the native ypuA signal sequence (ypuAss; SEQ ID NO: 1) comprises a substitution of an aspartic acid (Asp; D) to a serine (Ser; S) residue at position 24 of SEQ ID NO: 2 (i.e., the (-2) position relative to the signal peptidase cleavage site of SEQ ID NO: 2). [0248] The DNA fragments were amplified using Q5 DNA polymerase per the manufacturer’s instructions. The PCR products were purified using Zymo clean and concentrate five (5) columns per manufacturer’s instructions. The DNA fragments were assembled by overlapping extension PCR, generating the template (SEQ ID NO: 5) for the 1st amylase-1 integration cassette w/ mod-ypuAss (SEQ ID NO: 7; FIG.2A/FIG.2B). EXAMPLE 2 CONSTRUCTION OF A TEMPLATE PLASMID FOR A SECOND COPY OF AN AMYLASE-1 INTEGRATION CASSETTE
[0249] The present example describes construction of a template DNA by overlapping extension PCR (SEQ ID NO: 16) for the integration of a second (2nd) copy of the Amy-1 (SEQ ID NO: 6) expression cassette. More specifically, the 2nd copy of the amy-1 cassette with mod-ypuAss (2nd copy amy-1 cassette mod-ypuAss;; SEQ ID NO: 17) comprises an upstream (5′) homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to the lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (pro; SEQ ID NO: 20) operably linked to DNA encoding the B. subtilis aprE 5′UTR (SEQ ID NO: 11) operably linked to DNA encoding the modified B. licheniformis ypuA signal sequence (mod- ypuAss; SEQ ID NO: 2) operably linked to DNA encoding the Amy-1 (SEQ ID NO: 6) reporter protein operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream (3′) homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21). Thus, the 1st copy amy-1 cassette (SEQ ID NO: 7) and the 2nd copy amy-1 cassette (SEQ ID NO: 17) comprise the same modified ypuA signal sequence (i.e., mod-ypuAss; SEQ ID NO: 2) shown in FIG.1A. [0250] The DNA fragments were amplified using Q5 DNA polymerase per the manufacturer’s instructions. The PCR products were purified using Zymo clean and concentrate 5 columns per manufacturer’s instructions. The DNA fragments were assembled by overlapping extension PCR, generating the template (SEQ ID NO: 16) for 2nd Amylase 1 cassette (SEQ ID NO: 16; FIG.2C/FIG.2D). EXAMPLE 3 CONSTRUCTION OF A BACILLUS LICHENIFORMIS STRAIN EXPRESSING TWO COPIES OF AMYLASE-1 WITH A MODIFIED YPUA SIGNAL SEQUENCE [0251] In the instant example, the expression cassettes constructed and described in the preceding examples were integrated into an exemplary B. licheniformis host strain. More particularly, the B. licheniformis host strain named “AL223” was constructed by integrating the 1st and 2nd copies of the amy-1 cassettes into a B. licheniformis parental strain named “BF1965”, which BF1965 strain comprises deletions of the wild-type serA and lysA genes (abbreviated “∆serA/∆lysA”). For example, the recombinant B. licheniformis AL223 strain was constructed by integrating the 1st amy-1 cassette (SEQ ID NO: 7) into the serA locus, and the 2nd amy-1 cassette (SEQ ID NO: 17) into the lysA locus of the parental BF1965 strain (∆serA/∆lysA). [0252] Thus, the 1st amy-1 cassette (SEQ ID NO: 7) integration fragment was generated by PCR amplification from a template for 1st amy-1 cassette (SEQ ID NO: 5) with the “oAL106” (SEQ ID NO: 14) and “oAL108” (SEQ ID NO: 15) primer pair. The 2nd amy-1 cassette (SEQ ID NO: 17) integration fragment was generated by PCR amplification from a template for 2nd amy-1 cassette (SEQ ID NO: 16) with the “oAL079” (SEQ ID NO: 22) and “oAL080” (SEQ ID NO: 23) primer pair.
[0253] Subsequently, the 1st and 2nd amy-1 expression cassettes (SEQ ID NO: 7 and SEQ ID NO: 17) were transformed into parental BF1965 strain (∆serA/∆lysA) strain using the methods described in PCT Publication No. WO2019/040412. Briefly, the BF1965 competent cells were generated by growing the strain overnight in L broth containing one hundred (100) ppm spectinomycin at 37°C with 250 RPM shaking. The culture was diluted the next day to OD600 of 0.7 of fresh L broth containing one hundred (100) ppm spectinomycin. This new culture was grown for one (1) hour at 37°C, 250 RPM shaking. D-xylose was added to 0.1% w·v-1. The culture was grown for an additional four (4) hours at 37°C and 250 RPM shaking. The cells were harvested at 1700·g for seven (7) minutes, and used as competent cells for transformation. [0254] One hundred (100) µl of BF1965 competent cells were mixed with twenty (20) µl of the 1st amy-1 cassette integration fragment (SEQ ID NO: 7). The cell/DNA mixture was incubated at 1200 RPM, 37°C for one and a half (1.5) hours. The mixture was then plated on TSS agar plates containing eighty-eight (88) ppm lysine and hundred (100) ppm spectinomycin. The inoculated plates were incubated at 37°C for forty- eight (48) hours. Transformed colonies were screened by PCR amplification with the “oAL069” (SEQ ID NO: 24) and “oAL076” (SEQ ID NO: 25) primer pair. This PCR product, a 3,709 bp fragment (SEQ ID NO: 26), was sequenced using the method of Sanger and the “oAL081” (SEQ ID NO: 32), “seq1” (SEQ ID NO: 33) and “seq2” (SEQ ID NO: 34) primers. A colony with the correct integration of the cassette (SEQ ID NO: 7) was stored as strain AL217. [0255] The AL217 competent cells were generated as described above. One hundred (100) µl of AL217 competent cells were mixed with twenty (20) µl of the 2nd Amy-1 cassette (SEQ ID NO: 17) integration fragment. The cell/DNA mixture was incubated at 1200 RPM, 37°C for one and a half (1.5) hours. The mixture was then plated on TSS agar plates. The inoculated plates were incubated at 37°C for forty-eight (48) hours. Transformed colonies were screened by PCR amplification with the “oAL082” (SEQ ID NO: 27) and “oAL083” (SEQ ID NO: 28) primers pair. This PCR product, 2,133 bp fragment (SEQ ID NO: 29), was sequenced using the method of Sanger and the “oAL086” (SEQ ID NO: 35), “seq1” (SEQ ID NO: 33) and “seq2” (SEQ ID NO: 34) primers. [0256] A colony with the correct integration of the 1st amy-1 cassette (SEQ ID NO: 7) and the 2nd amy-1 cassette (SEQ ID NO: 17) was passaged on L agar until the colonies were stored as strain AL223 ([serA::[p3-mod-ypuAss-amy-1]-serA) and ([lysA::[p2-mod-ypuAss-amy-1]-lysA). [0257] To test relative performance of the modified ypuA signal sequence (SEQ ID NO: 2) on Amy-1 production, a control strain with a modified B. licheniformis AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) was constructed in the BF1965 strain by integration of a 1st amy-1 cassette (SEQ ID NO: 30) at the serA locus and integration of a 2nd amy-1 cassette (SEQ ID NO: 31) at the lysA locus. In particular, the modified B. licheniformis AmyL signal sequence (mod-AmyLss) has been described PCT Publication
No.2023/023642, and is an improved signal (secretion) sequence compared to the native B. licheniformis AmyL signal sequence. For instance, as shown in FIG.1B, the modified (control) AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) as compared to the native AmyL signal sequence (AmyLss; SEQ ID NO: 3) comprises a substitution of an alanine (Ala, A) to a serine (Ser; S) residue at position 28 of SEQ ID NO: 4 (i.e., the (-2) position relative to the signal peptidase cleavage site of SEQ ID NO: 4). More particularly, the 1st amy-1 cassette with mod-AmyLss (SEQ ID NO: 30; FIG. 2E/FIG. 2F) comprises an upstream homology arm (up) for the serA locus (serA.up; SEQ ID NO: 8) operably linked to DNA comprising a serA (ORF; SEQ ID NO: 9) operably linked to the synthetic p3 promoter (SEQ ID NO: 10) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) operably linked to DNA encoding Amy- 1 (SEQ ID NO: 6) operably linked to a B. licheniformis amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream homology arm (down) for the serA locus (serA.down; SEQ ID NO: 13). The 2nd amy-1 cassette with mod-AmyLss (SEQ ID NO: 31; FIG.2G/FIG.2H) comprises an upstream homology arm (up) to the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to DNA comprising a lysA ORF (SEQ ID NO: 19) operably linked to the synthetic p2 promoter (SEQ ID NO: 20) operably linked to DNA comprising the B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the mod-AmyLss (SEQ ID NO: 4) operably linked to DNA encoding Amy-1 (SEQ ID NO: 6) operably linked to the B. licheniformis [0258] amyL transcriptional terminator (amyL term; SEQ ID NO: 12) operably linked to the downstream homology arm (down) to the lysA locus (lysA.down; SEQ ID NO: 21). EXAMPLE 4 EFFECT OF THE MODIFIED YPUA SIGNAL SEQUENCE ON AMYLASE-1 PRODUCTION [0259] In the present example, the B. licheniformis AL223 strain comprising two integrated copies of the amylase-1 expression cassettes with the modified B. licheniformis ypuA signal sequence (SEQ ID NO: 2) were assayed for production of the Amy-1 reporter and compared to the control AL207 strain comprising two integrated copies of the amylase-1 expression cassettes with the modified (control) AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) using standard small-scale conditions, as described in PCT Publication No. WO2019/055261 (incorporated herein by reference). For example, the Amy-1 protein production was quantified using the method of Bradford assay, wherein the relative improvement in production of Amy-1 from the AL223 strain is compared to the production of the Amy-1 from the AL207 (control) strain, as presented below in TABLE 1
Strain Name Signal Sequence SEQ ID NO Relative Amy1 Production ± CV AL223 mod-ypuAss 2 1.50 ± 0.07 R
ELATIVE PERFORMANCE OF MODIFIED YPUA SIGNAL SEQUENCE VERSUS MODIFIED AMYL SIGNAL SEQUENCES ON AMY-1 PRODUCTION [0260] Thus, as shown above in TABLE 1, the modified ypuA signal sequence (SEQ ID NO: 2) demonstrates an approximately 50% improvement in Amy-1 protein production (strain AL223) relative to the Amy-1 protein production (control strain AL207) with the modified AmyL (control) signal sequence (SEQ ID NO: 4). EXAMPLE 5 CONSTRUCTION OF A TEMPLATE PLASMID FOR A SECOND COPY OF AN AMYLASE-2 INTEGRATION CASSETTE [0261] Applicant identified a native B. licheniformis ypuA protein signal sequence (SEQ ID NO: 1) that is particularly useful for enhancing/increasing production (secretion) of Amylase-2 (Amy-2; SEQ ID NO: 37). The instant example describes construction of a template plasmid named “pWS704” (SEQ ID NO: 36) for the integration of second copy of Amylase-2 (Amy-2; SEQ ID NO: 37) expression cassette. For example, the 2nd copy amy-2 cassette with a native ypuA signal sequence (FIG.3C/FIG.3D, amy-2 cassette with ypuAss; SEQ ID NO: 38) comprises an upstream (5′) homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to a downstream DNA comprising a B. licheniformis lysA ORF (SEQ ID NO: 19) operably linked to a synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA comprising a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the native B. licheniformis YpuA signal sequence (ypuAss; SEQ ID NO: 1) operably linked to DNA encoding the Amy- 2 protein (SEQ ID NO: 37) operably linked to a B. licheniformis amyL transcriptional terminator (AmyL term; SEQ ID NO: 12) operably linked to a downstream (3′) homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21). For example, the amino acid positions of the native B. licheniformis ypuA
signal sequence (ypuAss) comprise amino acid (residue) positions M1- A25, as presented in FIG.1A (SEQ ID NO: 1). [0262] The DNA fragments were amplified using Q5 DNA polymerase per the manufacturer’s instructions. The PCR products were purified using Zymo clean and concentrate five (5) columns per manufacturer’s instructions. The DNA fragments were assembled into a plasmid named “pRS426” purchased from ATCC (ATCC Catalogue No. 7710) by using yeast gap-repair cloning method (Joska et al., 2014), generating plasmid pWS704 (SEQ ID NO: 36). EXAMPLE 6 CONSTRUCTION OF A BACILLUS LICHENIFORMIS STRAIN EXPRESSING TWO COPIES OF AMYLASE-2 [0263] In the instant example, the 2nd copy amy-2 expression cassette (SEQ ID NO: 38) constructed in Example 5 was integrated into the lysA locus of the B. licheniformis strain BF780, resulting in the modified B. licheniformis host strain named “WS2756”. More particularly, the WS2756 host was constructed by integrating the 2nd copy of the amy-2 cassette (SEQ ID NO: 38) into the lysA locus of the BF780 strain. For example, a fragment of the amy-2 expression cassette (with native ypuAss) was generated by PCR amplification from the plasmid template pWS704 (SEQ ID NO: 36) with the ws683 (SEQ ID NO: 41) and ws688 (SEQ ID NO: 42) primer pair. The amy-2 cassette integration fragment was transformed into the BF780 strain (∆serA::[p3-mod-AmyLss-amy-2]serA-∆lysA) using the methods described in PCT Publication No. WO2023/023642. Briefly, the BF780 competent cells were generated by growing the strain overnight in L broth containing one hundred (100) ppm spectinomycin at 37°C with 250 RPM shaking. The culture was diluted the next day to OD600 of 0.7 of fresh L broth containing one hundred (100) ppm spectinomycin. This new culture was grown for one (1) hour at 37°C, 250 RPM shaking. D-xylose was added to 0.1% w·v-1. The culture was grown for an additional four (4) hours at 37°C and 250 RPM shaking. The cells were harvested at 1700·g for seven (7) minutes, and used as competent cells for transformation. [0264] One hundred (100) µl of BF780 competent cells were mixed with twenty (20) µl of the 2nd copy amy-2 integration fragment (SEQ ID NO: 38). The cell/DNA mixture was incubated at 1200 RPM, 37°C for one and a half (1.5) hours. The mixture was then plated on TSS agar plates. The inoculated plates were incubated at 37°C for forty-eight (48) hours. Transformed colonies were screened by PCR amplification with the ws775 (SEQ ID NO: 43) and ws776 (SEQ ID NO: 44) primer pair. This PCR product, a 1,892 bp fragment (SEQ ID NO: 45), was sequenced using the method of Sanger and the ws775 and ws776 primers. A colony with the correct integration of the 2nd copy amy-2 cassette (FIG. 3C/FIG. 3D; SEQ ID NO: 38) was stored as strain WS2756 (serA::[p3-mod-AmyLss-amylase 2] serA lysA::[p2-ypuAss-amy-2] lysA). [0265] To test relative performance of the ypuA signal sequence (ypuAss; SEQ ID NO: 1) on the Amy-2 (SEQ ID NO: 37) reporter production, a control strain with the modified B. licheniformis AmyL signal
sequence (SEQ ID NO: 4) was constructed in the BF780 parent by integration of a 2nd copy amy-2 cassette with mod-AmyLss (SEQ ID NO: 40) at the lysA locus. For instance, the 2nd copy amy-2 cassette with mod- AmyLss (FIG.3E/FIG.3F; SEQ ID NO: 40) comprises an upstream homology arm (up) for the lysA locus (lysA.up; SEQ ID NO: 18) operably linked to a DNA sequence encoding lysA (ORF; SEQ ID NO: 19) operably linked to the synthetic p2 promoter (p2 pro; SEQ ID NO: 20) operably linked to DNA encoding a B. subtilis aprE 5′-UTR (SEQ ID NO: 11) operably linked to DNA encoding the modified B. licheniformis AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) operably linked to DNA encoding the Amy-2 reporter protein (SEQ ID NO: 37) operably linked to the B. licheniformis amyL transcriptional terminator (SEQ ID NO: 12) operably linked to the downstream homology arm (down) for the lysA locus (lysA.down; SEQ ID NO: 21). A colony with the correct integration of the amy-2 cassette (SEQ ID NO: 38) was stored as a control strain BF822 (serA::[p3-mod-AmyLss-amylase 2] serA lysA::[p2-AmyLss-amy-2] lysA). [0266] For example, as shown in FIG. 1B, the modified AmyL signal sequence (mod-AmyLss; SEQ ID NO: 4) relative to the native AmyL signal sequence (AmyLss; SEQ ID NO: 3) comprises a substitution of alanine (A) to serine (S) at the minus (-) 2 position, relative to the signal peptidase cleavage site. EXAMPLE 7 EFFECT OF THE YPUA SIGNAL SEQUENCE ON AMYLASE-2 PRODUCTION [0267] In the present example, the WS2756 strain and the BF822 (control) strain constructed and described in Example 6 were assayed for production of the Amy-2 (SEQ ID NO: 37) reporter using standard small-scale conditions, as described in PCT Publication No. WO2018/156705 and WO2019/055261 (each Strain Name Signal Sequence (ss) SEQ ID NO Relative Amy-2 Production WS2756 A 1 118 incor re Amy-2
(SEQ ID NO: 37) production was quantified using the method of Bradford or the Ceralpha assay, wherein the relative improvement in Amy-2 production from the WS2756 strain is compared to Amy-2 production from the BF822 (control) strain. TABLE 2 RELATIVE PERFORMANCE OF YPUA SIGNAL SEQUENCE VERSUS MODIFIED AMYL SIGNAL SEQUENCES ON AMY-2 PRODUCTION
[0268] Thus, as shown in TABLE 2, the native B. licheniformis ypuA signal sequence (ypuAss) demonstrates an improvement in Amy-2 reporter protein production in the WS2756 strain relative to the B. licheniformis (control) strain BF822, comprising the modified B. licheniformis AmyL signal sequence (modAmyLss).
REFERENCES PCT Publication No. WO2002/14490 PCT Publication No. WO2003/083125 PCT Publication No. WO2008/112459 PCT Publication No. WO2023/023642 Armenteros et al., “SignalP 5.0 improves signal peptide predictions using deep neural networks”, Nature Biotechnology, 37: 420-423, 2019. Ausubel et al., “Current Protocols in Molecular Biology”, published by Greene Publishing Assoc. and Wiley-Interscience (1987). Brode et al., “Subtilisin BPN' variants: increased hydrolytic activity on surface-bound substrates via decreased surface activity”, Biochemistry, 35(10):3162-3169, 1996. Caspers et al., “Improvement of Sec-dependent secretion of a heterologous model protein in Bacillus subtilis by saturation mutagenesis of the N-domain of the AmyE signal peptide”, Appl. Microbiol. Biotechnol., 86(6):1877-1885, 2010. Chen et al., “Combinatorial Sec pathway analysis for improved heterologous protein secretion in Bacillus subtilis: identification of bottlenecks by systematic gene overexpression”, Microbial Cell Factories, 14(1):92, 2015. Earl et al., “Ecology and genomics of Bacillus subtilis”, Trends in Microbiology.,16(6):269-275, 2008. Olempska-Beer et al., “Food-processing enzymes from recombinant microorganisms--a review”’ Regul. Toxicol. Pharmacol., 45(2):144-158, 2006. Raul et al., “Production and partial purification of alpha amylase from Bacillus subtilis (MTCC 121) using solid state fermentation”, Biochemistry Research International, 2014. Sambrook et al., “Molecular Cloning: A Laboratory Manual” Cold Spring Harbor Laboratory: Cold Spring Harbor, N.Y. (1989), (2001) and (2012). Tjalsma et al., “Signal Peptide-Dependent Protein Transport in Bacillus subtilis: a Genome-Based Survey of the Secretome”, Microbiology and Molecular Biology Reviews, 64: 515-547, 2000. Van Dijl and Hecker, “Bacillus subtilis: from soil bacterium to super-secreting cell factory”, Microbial Cell Factories, 12(3).2013.
Claims
CLAIMS 1. An isolated nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) comprising SEQ ID NO: 2.
2. A polynucleotide comprising an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) comprising SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI).
3. A polynucleotide comprising an upstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) comprising SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI).
4. An expression cassette comprising an upstream promoter sequence operably linked to a downstream 5′-UTR sequence operably linked to a downstream polynucleotide of claim 2 or claim 3, wherein the promoter and 5′-UTR sequences are functional in Bacillus sp. cells.
5. The polynucleotide of claim 2 or claim 3, wherein the heterologous POI is an amylase.
6. A modified Bacillus sp. cell comprising an introduced polynucleotide comprising an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI), or comprising an introduced polynucleotide comprising an upstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous POI.
7. A modified Bacillus sp. cell comprising at least two introduced polynucleotides, wherein the first polynucleotide comprises an upstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding a heterologous protein of interest (POI), or comprises an upstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding a heterologous POI, and the second polynucleotide comprises an upstream nucleic acid encoding a ypuAss of SEQ ID NO: 1 operably linked to a downstream nucleic acid encoding the same heterologous POI, or comprising an upstream nucleic acid encoding a mod-ypuAss of SEQ ID NO: 2 operably linked to a downstream nucleic acid encoding the same heterologous POI.
8. A modified Bacillus sp. cell comprising an introduced expression cassette of claim 4.
9. A modified Bacillus sp. cell comprising at least two introduced expression cassettes of claim 4.
10. The modified cell of claim 6 or claim 7, wherein the heterologous POI is an amylase.
11. A method for producing a heterologous protein of interest (POI) in a Bacillus sp. cell comprising: (a) introducing into a Bacillus sp. cell an expression cassette comprising an upstream (5′) promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding a heterologous POI, and (b) fermenting the modified cell under suitable conditions for the production of the POI.
12. The method of claim 11, wherein the modified cell produces an increased amount of POI relative to a control cell producing the same POI when fermented under the same conditions, wherein the control cell comprises an introduced expression cassette comprising the upstream promoter operably linked to the downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a modified AmyL signal sequence (mod-AmyLss) of SEQ ID NO: 4 operably linked to a downstream nucleic acid encoding the heterologous POI.
13. A method for producing a heterologous protein of interest (POI) in a Bacillus sp. cell comprising: (a) introducing into a Bacillus sp. cell at least two expression cassettes, wherein (i) the first cassette comprises an upstream (5′) promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a native ypuA signal sequence (ypuAss) of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a modified ypuA signal sequence (mod-ypuAss) of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI, and (ii) the second cassette comprises an upstream promoter operably linked to a downstream 5′-UTR sequence operably linked to a downstream nucleic acid encoding a ypuAss of SEQ ID NO: 1, or operably linked to a downstream nucleic acid encoding a mod-ypuAss of SEQ ID NO: 2, operably linked to a downstream nucleic acid encoding the heterologous POI, and (b) fermenting the modified cell under suitable conditions for the production of the POI.
14. The method of claim 13, wherein the modified cell produces an increased amount of POI relative to a control cell producing the same POI when fermented under the same conditions, wherein the control cell comprises at least two introduced expression cassettes, wherein (i) the first cassette comprises the upstream promoter operably linked to the downstream 5′- UTR sequence operably linked to a downstream nucleic acid encoding a modified AmyL
signal sequence (mod-AmyLss) of SEQ ID NO: 4 sequence operably linked to a downstream nucleic acid encoding the heterologous POI, and (ii) the second cassette comprises an upstream promoter operably linked to the downstream 5′- UTR sequence operably linked to a downstream nucleic acid encoding a modified AmyL signal sequence (mod-AmyLss) of SEQ ID NO: 4 operably linked to a downstream nucleic acid encoding the heterologous POI.
15. The method of claim 11 or claim 13, wherein the POI is secreted into the fermentation broth when fermented under suitable conditions.
16. The method of claim 11 or claim 13, wherein the POI is an amylase.
17. The method of claim 11, wherein the increased amount of POI produced by the modified cell is at least about 10% increased as compared the same POI produced by the control cell.
18. The method of claim 13, wherein the increased amount of POI produced by the modified cell is at least about 40% increased as compared the same POI produced by the control cell.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363597527P | 2023-11-09 | 2023-11-09 | |
| US63/597,527 | 2023-11-09 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025101486A1 true WO2025101486A1 (en) | 2025-05-15 |
Family
ID=93648791
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/054516 Pending WO2025101486A1 (en) | 2023-11-09 | 2024-11-05 | Methods and compositions for enhanced protein production in bacillus cells |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2025101486A1 (en) |
Citations (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2002014490A2 (en) | 2000-08-11 | 2002-02-21 | Genencor International, Inc. | Bacillus transformation, transformants and mutant libraries |
| WO2003083125A1 (en) | 2002-03-29 | 2003-10-09 | Genencor International, Inc. | Ehanced protein expression in bacillus |
| WO2005069762A2 (en) * | 2004-01-09 | 2005-08-04 | Novozymes Inc. | Bacillus licheniformis chromosome |
| WO2008112459A2 (en) | 2007-03-09 | 2008-09-18 | Danisco Us Inc., Genencor Division | Alkaliphilic bacillus species a-amylase variants, compositions comprising a-amylase variants, and methods of use |
| WO2018156705A1 (en) | 2017-02-24 | 2018-08-30 | Danisco Us Inc. | Compositions and methods for increased protein production in bacillus licheniformis |
| WO2019040412A1 (en) | 2017-08-23 | 2019-02-28 | Danisco Us Inc | Methods and compositions for efficient genetic modifications of bacillus licheniformis strains |
| WO2019055261A1 (en) | 2017-09-13 | 2019-03-21 | Danisco Us Inc | Modified 5'-untranslated region (utr) sequences for increased protein production in bacillus |
| WO2021146411A1 (en) * | 2020-01-15 | 2021-07-22 | Danisco Us Inc | Compositions and methods for enhanced protein production in bacillus licheniformis |
| WO2023023642A2 (en) | 2021-08-20 | 2023-02-23 | Danisco Us Inc. | Methods and compositions for enhanced protein production in bacillus cells |
-
2024
- 2024-11-05 WO PCT/US2024/054516 patent/WO2025101486A1/en active Pending
Patent Citations (9)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| WO2002014490A2 (en) | 2000-08-11 | 2002-02-21 | Genencor International, Inc. | Bacillus transformation, transformants and mutant libraries |
| WO2003083125A1 (en) | 2002-03-29 | 2003-10-09 | Genencor International, Inc. | Ehanced protein expression in bacillus |
| WO2005069762A2 (en) * | 2004-01-09 | 2005-08-04 | Novozymes Inc. | Bacillus licheniformis chromosome |
| WO2008112459A2 (en) | 2007-03-09 | 2008-09-18 | Danisco Us Inc., Genencor Division | Alkaliphilic bacillus species a-amylase variants, compositions comprising a-amylase variants, and methods of use |
| WO2018156705A1 (en) | 2017-02-24 | 2018-08-30 | Danisco Us Inc. | Compositions and methods for increased protein production in bacillus licheniformis |
| WO2019040412A1 (en) | 2017-08-23 | 2019-02-28 | Danisco Us Inc | Methods and compositions for efficient genetic modifications of bacillus licheniformis strains |
| WO2019055261A1 (en) | 2017-09-13 | 2019-03-21 | Danisco Us Inc | Modified 5'-untranslated region (utr) sequences for increased protein production in bacillus |
| WO2021146411A1 (en) * | 2020-01-15 | 2021-07-22 | Danisco Us Inc | Compositions and methods for enhanced protein production in bacillus licheniformis |
| WO2023023642A2 (en) | 2021-08-20 | 2023-02-23 | Danisco Us Inc. | Methods and compositions for enhanced protein production in bacillus cells |
Non-Patent Citations (12)
| Title |
|---|
| ARMENTEROS ET AL.: "SignalP 5.0 improves signal peptide predictions using deep neural networks", NATURE BIOTECHNOLOGY, vol. 37, 2019, pages 420 - 423, XP036900634, DOI: 10.1038/s41587-019-0036-z |
| AUSUBEL ET AL.: "Current Protocols in Molecular Biology", 1987, GREENE PUBLISHING ASSOC. AND WILEY-INTERSCIENCE |
| BRODE ET AL.: "Subtilisin BPN' variants: increased hydrolytic activity on surface-bound substrates via decreased surface activity", BIOCHEMISTRY, vol. 35, no. 10, 1996, pages 3162 - 3169, XP002380291, DOI: 10.1021/bi951990h |
| CASPERS ET AL.: "Improvement of Sec-dependent secretion of a heterologous model protein in Bacillus subtilis by saturation mutagenesis of the N-domain of the AmyE signal peptide", APPL. MICROBIOL. BIOTECHNOL., vol. 86, no. 6, 2010, pages 1877 - 1885, XP019799937 |
| CHEN ET AL.: "Combinatorial Sec pathway analysis for improved heterologous protein secretion in Bacillus subtilis: identification of bottlenecks by systematic gene overexpression", MICROBIAL CELL FACTORIES, vol. 14, no. 1, 2015, pages 92, XP055639249, DOI: 10.1186/s12934-015-0282-9 |
| EARL ET AL.: "Ecology and genomics of Bacillus subtilis", TRENDS IN MICROBIOLOGY., vol. 16, no. 6, 2008, pages 269 - 275, XP022711144, DOI: 10.1016/j.tim.2008.03.004 |
| GU ZIQIANG ET AL: "High-efficiency heterologous expression of nattokinase based on a combinatorial strategy", PROCESS BIOCHEMISTRY, ELSEVIER LTD, GB, vol. 133, 19 August 2023 (2023-08-19), pages 65 - 74, XP087411310, ISSN: 1359-5113, [retrieved on 20230819], DOI: 10.1016/J.PROCBIO.2023.08.008 * |
| OLEMPSKA-BEER ET AL.: "Food-processing enzymes from recombinant microorganisms--a review", REGUL. TOXICOL. PHARMACOL., vol. 45, no. 2, 2006, pages 144 - 158, XP024915279, DOI: 10.1016/j.yrtph.2006.05.001 |
| RAUL ET AL.: "Production and partial purification of alpha amylase from Bacillus subtilis (MTCC 121) using solid state fermentation", BIOCHEMISTRY RESEARCH INTERNATIONAL, 2014 |
| SAMBROOK ET AL.: "Molecular Cloning: A Laboratory Manual", 1989, COLD SPRING HARBOR LABORATORY: COLD SPRING HARBOR |
| TJALSMA ET AL.: "Signal Peptide-Dependent Protein Transport in Bacillus subtilis: a Genome-Based Survey of the Secretome", MICROBIOLOGY AND MOLECULAR BIOLOGY REVIEWS, vol. 64, 2000, pages 515 - 547, XP002227136, DOI: 10.1128/MMBR.64.3.515-547.2000 |
| VAN DIJLHECKER: "Bacillus subtilis: from soil bacterium to super-secreting cell factory", MICROBIAL CELL FACTORIES, vol. 12, no. 3, 2013 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11781147B2 (en) | Promoter sequences and methods thereof for enhanced protein production in Bacillus cells | |
| CN110520520B (en) | Compositions and methods for increasing protein production in bacillus licheniformis | |
| US12534716B2 (en) | Compositions and methods for enhanced protein production in bacillus licheniformis | |
| US11414643B2 (en) | Mutant and genetically modified Bacillus cells and methods thereof for increased protein production | |
| JP2025072450A (en) | Compositions and methods for increased protein production in Bacillus licheniformis | |
| US20240360430A1 (en) | Methods and compositions for enhanced protein production in bacillus cells | |
| US20250223623A1 (en) | Pro-region mutations enhancing protein production in gram-positive bacterial cells | |
| US20260092297A1 (en) | Novel promoter and 5'-untranslated region mutations enhancing protein production in gram-positive cells | |
| JP7656603B2 (en) | Compositions and methods for enhanced protein production in Bacillus cells | |
| EP4294823A1 (en) | Methods and compositions for producing proteins of interest in pigment deficient bacillus cells | |
| US20260132387A1 (en) | Compositions and methods for enhanced protein production in bacillus cells | |
| WO2026107442A1 (en) | Compositions and methods for enhanced protein production in bacillus cells | |
| US20240263185A1 (en) | Compositions and methods for enhanced protein production in bacillus cells | |
| US20250002925A1 (en) | Compositions and methods for enhanced protein production in bacillus cells |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24812292 Country of ref document: EP Kind code of ref document: A1 |

