EP4655395A2 - Mb2cas12a variants with flexible pam spectrum - Google Patents

Mb2cas12a variants with flexible pam spectrum

Info

Publication number
EP4655395A2
EP4655395A2 EP24747710.2A EP24747710A EP4655395A2 EP 4655395 A2 EP4655395 A2 EP 4655395A2 EP 24747710 A EP24747710 A EP 24747710A EP 4655395 A2 EP4655395 A2 EP 4655395A2
Authority
EP
European Patent Office
Prior art keywords
polypeptide
sequence
mb2casl2a
seq
mutant
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24747710.2A
Other languages
German (de)
French (fr)
Inventor
Feng Xiong
Jiang Li
Jianping Xu
Ingrid Udranszky
Yanli Wang
James J. ENGLISH
Robert Gerald Whalen
Brett Konstantin MORRIS
Claire Marie BAZYK
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Syngenta Crop Protection AG Switzerland
Original Assignee
Syngenta Crop Protection AG Switzerland
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Syngenta Crop Protection AG Switzerland filed Critical Syngenta Crop Protection AG Switzerland
Publication of EP4655395A2 publication Critical patent/EP4655395A2/en
Pending legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/102Mutagenizing nucleic acids
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/11DNA or RNA fragments; Modified forms thereof; Non-coding nucleic acids having a biological activity
    • C12N15/113Non-coding nucleic acids modulating the expression of genes, e.g. antisense oligonucleotides; Antisense DNA or RNA; Triplex- forming oligonucleotides; Catalytic nucleic acids, e.g. ribozymes; Nucleic acids used in co-suppression or gene silencing
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/79Vectors or expression systems specially adapted for eukaryotic hosts
    • C12N15/82Vectors or expression systems specially adapted for eukaryotic hosts for plant cells, e.g. plant artificial chromosomes (PACs)
    • C12N15/8201Methods for introducing genetic material into plant cells, e.g. DNA, RNA, stable or transient incorporation, tissue culture methods adapted for transformation
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/16Hydrolases (3) acting on ester bonds (3.1)
    • C12N9/22Ribonucleases [RNase]; Deoxyribonucleases [DNase]
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/16Hydrolases (3) acting on ester bonds (3.1)
    • C12N9/22Ribonucleases [RNase]; Deoxyribonucleases [DNase]
    • C12N9/222Clustered regularly interspaced short palindromic repeats [CRISPR]-associated [CAS] enzymes
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2310/00Structure or type of the nucleic acid
    • C12N2310/10Type of nucleic acid
    • C12N2310/20Type of nucleic acid involving clustered regularly interspaced short palindromic repeats [CRISPR]

Definitions

  • Casl2a mutants from Moraxella bovoculi AAX08 and methods for use thereof. These mutants have a broader PAM spectrum than the wildtype enzyme.
  • Mb2Casl2a from Moraxella bovoculi AAX08 has demonstrated in planta genome editing capability (Zhang et al., 2021), however its PAM site is less flexible compared to Cas9. Therefore, fewer targets in eukaryotic genomes can be edited with Mb2Casl2a compared to Cas9. With its distinct property of high performance at lower temperature, Mb2Casl2a needs to have a flexible PAM.
  • mutant Mb2Casl2a polypeptide variant of the wildtype that recognizes an alternative PAM site.
  • This variant may comprise at least one amino acid substitution introduced into a wild-type Mb2Casl2a polypeptide sequence of SEQ ID NO: 1. This substitution can occur at the following positions: D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
  • a mutant Mb2Casl2a polypeptide comprising at least one amino acid substitution introduced into a wild-type Mb2Casl2a polypeptide sequence of SEQ ID NO: 1.
  • the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
  • the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q.
  • the mutant Mb2Casl2a polypeptide comprises a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
  • Described herein is a method of editing a plant genome, comprising contacting said plant genome with the mutant Mb2Casl2a polypeptide recited above.
  • the method further comprises a guide RNA.
  • the guide RNA is encoded by a sequence comprising SEQ ID NOs: 16-60.
  • Described herein is a construct or plasmid comprising a polynucleotide sequence encoding for the mutant Mb2Casl2a polypeptides recited above.
  • One embodiment is a non-human cell comprising the construct or plasmid.
  • RNP complex comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
  • SEQ ID NO: 1 is the amino acid sequence of wildtype Mb2Casl2a, also referred to as “wildtype” or “WT” throughout.
  • SEQ ID NO: 2 is the amino acid sequence of Mb2Casl2a variant named “46(172R)” comprising D172R, N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
  • SEQ ID NO: 3 is the amino acid sequence of Mb2Casl2a variant named “46(172Y)” comprising D172Y, N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
  • SEQ ID NO: 4 is the amino acid sequence of Mb2Casl2a variant named “46” comprising N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
  • SEQ ID NO: 5 is the amino acid sequence of Mb2Casl2a variant named "51(172R)" comprising D172R, N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
  • SEQ ID NO: 6 is the amino acid sequence of Mb2Casl2a variant named "51(172Y)" comprising D172Y, N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
  • SEQ ID NO: 7 is the amino acid sequence of Mb2Casl2a variant named "51" comprising N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
  • SEQ ID NO: 8 is the amino acid sequence of Mb2Casl2a variant named "18" comprising D172R, G672S, E678D, and Q725E amino acid substitutions.
  • SEQ ID NO: 9 is the amino acid sequence of Mb2Casl2a variant named "21" comprising D172R, N573D, A629S, L633I, Y636F, L642I, L695I, Q708K, Q725E, F753W, S758D, L782I, and L826F amino acid substitutions.
  • SEQ ID NO: 10 is the amino acid sequence of Mb2Casl2a variant named "28” comprising K569T, E570R, D572R, N573D, S758E, E795C, D783Q, and F824Y amino acid substitutions.
  • SEQ ID NO: 11 is the amino acid sequence of Mb2Casl2a variant named "28(172Y)" comprising D172Y, K569T, E570R, D572R, N573D, S758E, E795C, D783Q, and F824Y amino acid substitutions.
  • SEQ ID NO: 12 is the amino acid sequence of Mb2Casl2a variant named "44" comprising N563R, K569V, N573R, S758E, E795C, T788F, and F824M amino acid substitutions.
  • SEQ ID NO: 13 is the amino acid sequence of Mb2Casl2a variant named “44 (172R)” comprising D172R, N563R, K569V, N573R, S758E, E759C, T788F, F824M amino acid substitutions.
  • SEQ ID NO: 14 is the amino acid sequence of Mb2Casl2a variant named “44 (172Y)” comprising D172Y, N563R, K569V, N573R, S758E, E759C, T788F, F824M amino acid substitutions.
  • SEQ ID NO: 15 is the amino acid sequence of wildtype LbCasl2a.
  • SEQ ID NOs: 16-60 are the gRNA sequences used. See Table 5
  • SEQ ID NO: 61 is the direct repeat of Mb2Casl2a crRNA.
  • SEQ ID NO: 62 is a synthetic ZmGL2 DNA substrate.
  • SEQ ID NO: 63 is a synthetic ZmmiR528 DNA substrate.
  • SEQ ID NO: 64 is the amino acid sequence of Mb2Casl2a variant named “27” comprising K569T, E570R, D572R, N573D, A671M, Q705H, S758E and E795C amino acid substitutions.
  • SEQ ID NO: 65 is the amino acid sequence of Mb2Casl2a variant named “43” comprising N563R, K569V, N573R, S758C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
  • SEQ ID NO: 66 is the amino acid sequence of Mb2Casl2a variant named “45” comprising N563R, K569V, N573R, K625R, S696F, S758E, E759T, and T831N amino acid substitutions.
  • nucleic acid or amino acid sequence
  • nucleic acid sequence or an amino acid sequence that includes the subject sequence as a part or as its entire sequence.
  • the transitional phrase “consisting essentially of’ means that the scope of a claim is to be interpreted to encompass the specified materials or steps recited in the claim and those that do not materially affect the basic and novel characteristic(s) of the claimed matter.
  • the term “consisting essentially of’ when used in a claim of this disclosure is not intended to be interpreted to be equivalent to “comprising.”
  • a “plurality” refers to more than one entity.
  • a “plurality of individuals” refers to at least two individuals.
  • the term plurality refers to more than half of the whole.
  • a “plurality of a population” refers to more than half the members of that population.
  • plant refers to any plant at any stage of development, particularly a seed plant.
  • plant cell refers to a structural and physiological unit of a plant, comprising a protoplast and a cell wall.
  • the plant cell may be in form of an isolated single cell or a cultured cell, or as a part of higher organized unit such as, for example, plant tissue, a plant organ, or a whole plant.
  • the plant cell may be derived from or part of an angiosperm or gymnosperm.
  • the plant cell may be a monocotyledonous plant cell (e.g., a maize cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a turf grass cell, or an ornamental grass cell) or a dicotyledonous plant cell (e.g., a tobacco cell, a pepper cell, an eggplant cell, a sunflower cell, a crucifer cell, a flax cell, a potato cell, a cotton cell, a soybean cell, a sugar beet cell, or an oilseed rape cell.
  • a monocotyledonous plant cell e.g., a maize cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a turf grass cell, or an ornamental grass cell
  • a dicotyledonous plant cell e.g., a tobacco cell,
  • plant cell culture refers to cultures of plant units such as, for example, protoplasts, cell culture cells, cells in plant tissues, pollen, pollen tubes, ovules, embryo sacs, zygotes and embryos at various stages of development.
  • plant tissue refers to a group of plant cells organized into a structural and functional unit. Any tissue of a plant in planta or in culture is included. This term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue culture and any group of plant cells organized into structural and/or functional units. The use of this term in conjunction with, or in the absence of, any specific type of plant tissue as listed above or otherwise embraced by this definition is not intended to be exclusive of any other type of plant tissue.
  • plant part refers to a part of a plant, including single cells and cell tissues such as plant cells that are intact in plants, cell clumps and tissue cultures from which plants can be regenerated.
  • plant parts include, but are not limited to, single cells and tissues from pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, flower parts, fruits, stems, shoots, cuttings, and seeds; as well as pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, flower parts, fruits, stems, shoots, cuttings, scions, rootstocks, seeds, protoplasts, calli, and the like.
  • polypeptide “peptide,” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, the terms encompass amino acid chains of any length, including full-length proteins, wherein the amino acid residues are linked by covalent peptide bonds.
  • nucleic acid and “polynucleotide” are used interchangeably and as used herein refer to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single- or double-stranded form, as well as to both sense and anti-sense strands of RNA, cDNA, genomic DNA, mitochondrial DNA, and synthetic forms and mixed polymers of the above.
  • DNA is the genetic material while RNA is involved in the transfer of information contained within DNA into proteins.
  • a “genome” is the entire body of genetic material contained in each cell of an organism. It is understood that when an RNA is described, its corresponding cDNA is also described, wherein uridine is represented as thymidine.
  • a nucleotide refers to a ribonucleotide, deoxynucleotide or a modified form of either type of nucleotide, and combinations thereof.
  • a polynucleotide disclosed herein may include either or both naturally occurring and modified nucleotides linked together by naturally occurring and/or non-naturally occurring nucleotide linkages.
  • the nucleic acid molecules may be modified chemically or biochemically or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those of skill in the art.
  • Such modifications include, for example, labels, methylation, substitution of one or more of the naturally occurring nucleotides with an analogue, internucleotide modifications such as uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, and the like), charged linkages (e.g., phosphorothioates, phosphorodithioates, and the like), pendent moi eties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, and the like), chelators, alkylators, and modified linkages (e.g., alpha anomeric nucleic acids, and the like).
  • uncharged linkages e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, and the like
  • charged linkages e.g., phosphorothioates, phosphorodithioates, and the like
  • nucleic acid sequence encompasses its complement unless otherwise specified.
  • a reference to a nucleic acid molecule having a particular sequence should be understood to encompass its complementary strand, with its complementary sequence.
  • Nucleotide sequences are “complementary” when they specifically hybridize in solution (e.g., according to Watson-Crick base pairing rules).
  • the term also includes codon-optimized nucleic acids that encode the same polypeptide sequence. It is also understood that nucleic acids can be unpurified, purified, or attached, for example, to a synthetic material such as a bead or column matrix.
  • Nucleic acid residues can be referred to by individual letters: “A” is adenine; “C” is cytosine; “G” is guanine; “T” is thymine; “N” is any nucleotide; “V” is nucleotide A, C, or G.
  • nucleic acid sequences in the context of nucleic acid sequences means that when the nucleic acid sequences of certain sequences are aligned with each other, the nucleic acids that “correspond to” certain enumerated positions in the present invention are those that align with these positions in a reference sequence, but that are not necessarily in these exact numerical positions relative to a particular nucleic acid sequence of the invention.
  • Optimal alignment of sequences for comparison can be conducted by computerized implementations of known algorithms, or by visual inspection. Readily available sequence comparison and multiple sequence alignment algorithms are, respectively, the Basic Local Alignment Search Tool (BLAST) and ClustalW/ClustalW2/Clustal Omega programs available on the Internet (e g., the website of the EMBL-EBI).
  • BLAST Basic Local Alignment Search Tool
  • ClustalW/ClustalW2/Clustal Omega programs available on the Internet (e g., the website of the EMBL-EBI).
  • nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated.
  • degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed- base and/or deoxyinosine residues. See Batzer et al., Nucleic Acid Res. 19:5081 (1991);
  • identity refers to a sequence that has at least 60% sequence identity to a reference sequence.
  • percent identity can be any integer from 60% to 100%.
  • Exemplary embodiments include at least: 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, as compared to a reference sequence using the programs described herein; preferably BLAST using standard parameters, as described below.
  • sequence comparison typically one sequence acts as a reference sequence to which test sequences are compared.
  • test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated.
  • sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters.
  • a “comparison window,” as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of from 20 to 600, usually about 50 to about 200, more usually about 100 to about 150 in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned.
  • Methods of alignment of sequences for comparison are well-known in the art.
  • Optimal alignment of sequences for comparison may be conducted by the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981), by the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson and Lipman Proc. Natl. Acad. Sci. (U.S.A.) 85: 2444 (1988), by computerized implementations of these algorithms (e g., BLAST), or by manual alignment and visual inspection.
  • a “gene” is a defined region that is located within a genome and that, besides the aforementioned coding nucleic acid sequence, comprises other, primarily regulatory, nucleic acid sequences responsible for the control of the expression, that is to say the transcription and translation, of the coding portion. Genes can include both coding and non-coding regions (e g., introns, regulatory elements, promoters, enhancers, termination sequences and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or specific protein, including regulatory sequences. Genes may or may not be capable of being used to produce a functional protein. In some embodiments, a gene refers to only the coding region.
  • a chimeric gene refers to any gene that contains 1) DNA sequences, including regulatory and coding sequences that are not found together in nature, or 2) sequences encoding parts of proteins not naturally adjoined, or 3) parts of promoters that are not naturally adjoined. Accordingly, a chimeric gene may comprise regulatory sequences and coding sequences that are derived from different sources, or comprise regulatory sequences and coding sequences derived from the same source, but arranged in a manner different from that found in nature.
  • a gene may be “isolated” by which is meant a nucleic acid molecule that is substantially or essentially free from components normally found in association with the nucleic acid molecule in its natural state. Such components include other cellular material, culture medium from recombinant production, and/or various chemicals used in chemically synthesizing the nucleic acid molecule.
  • nucleic acid molecule or nucleotide sequence or an “isolated” polypeptide is a nucleic acid molecule, nucleotide sequence or polypeptide that, by the hand of man, exists apart from its native environment and/or has a function that is different, modified, modulated and/or altered as compared to its function in its native environment and is therefore not a product of nature.
  • An isolated nucleic acid molecule or isolated polypeptide may exist in a purified form or may exist in a non-native environment such as, for example, a recombinant host cell.
  • the term isolated means that it is separated from the chromosome and/or cell in which it naturally occurs.
  • a polynucleotide is also isolated if it is separated from the chromosome and/or cell in which it naturally occurs and is then inserted into a genetic context, a chromosome, a chromosome location, and/or a cell in which it does not naturally occur.
  • the recombinant nucleic acid molecules and nucleotide sequences of the invention can be considered to be “isolated” as defined above.
  • an “isolated nucleic acid molecule” or “isolated nucleotide sequence” is a nucleic acid molecule or nucleotide sequence that is not immediately contiguous with nucleotide sequences with which it is immediately contiguous (one on the 5' end and one on the 3' end) in the naturally occurring genome of the organism from which it is derived. Accordingly, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences that are immediately contiguous to a coding sequence.
  • 5' non-coding e.g., promoter
  • the term therefore includes, for example, a recombinant nucleic acid that is incorporated into a vector, into an autonomously replicating plasmid or virus, or into the genomic DNA of a prokaryote or eukaryote, or which exists as a separate molecule (e.g., a cDNA or a genomic DNA fragment produced by PCR or restriction endonuclease treatment), independent of other sequences. It also includes a recombinant nucleic acid that is part of a hybrid nucleic acid molecule encoding an additional polypeptide or peptide sequence.
  • isolated nucleic acid molecule or “isolated nucleotide sequence” can also include a nucleotide sequence derived from and inserted into the same natural, original cell type, but which is present in a nonnatural state, e.g., present in a different copy number, and/or under the control of different regulatory sequences than that found in the native state of the nucleic acid molecule.
  • isolated can further refer to a nucleic acid molecule, nucleotide sequence, polypeptide, peptide or fragment that is substantially free of cellular material, viral material, and/or culture medium (e.g., when produced by recombinant DNA techniques), or chemical precursors or other chemicals (e.g., when chemically synthesized).
  • an “isolated fragment” is a fragment of a nucleic acid molecule, nucleotide sequence or polypeptide that is not naturally occurring as a fragment and would not be found as such in the natural state. “Isolated” does not necessarily mean that the preparation is technically pure (homogeneous), but it is sufficiently pure to provide the polypeptide or nucleic acid in a form in which it can be used for the intended purpose.
  • “Homology dependent repair” or “homology directed repair” or “HDR” refers to a mechanism for repairing ssDNA and double stranded DNA (dsDNA) damage in cells. This repair mechanism can be used by the cell when there is an HDR template with a sequence with significant homology to the injury site.
  • the term “perfect HDR” refers to a situation in which genomic-homology junctions in the replaced allele underwent complete HDR and “imperfect HDR” refers to a situation in which genomic-homology junctions in the replaced allele underwent partial or incomplete HDR.
  • a donor DNA molecule with homology to the cleaved target DNA sequence is used as a template for repair of the cleaved target DNA sequence, resulting in the transfer of genetic information from the donor polynucleotide to the target DNA.
  • new nucleic acid material may be inserted/copied into the site.
  • a target DNA is contacted with a donor molecule, for example a donor DNA molecule.
  • a donor DNA molecule is introduced into a cell.
  • at least a segment of a donor DNA molecule integrates into the genome of the cell.
  • MMEJ Microhomology-mediated end joining
  • Alt-NHEJ alternative nonhomologous endjoining
  • This repair mechanism utilizes microhomologous sequences to align the broken strands
  • Nonhomologous end joining or “NHEJ” refers to a form of repairing double-stranded breaks in DNA.
  • the double-strand breaks are repaired by direct ligation of the break ends to one another. Generally, no new nucleic acid material is inserted into the site, although some nucleic acid material may be lost or added, resulting in a small deletion or a small insertion.
  • the proteins provided herein comprise a site-directed polypeptide.
  • a site-directed modifying polypeptide modifies target DNA (e.g., via cleavage or methylation of target DNA) and/or a polypeptide associated with target DNA (e.g., methylation or acetylation of a histone tail).
  • a site-directed modifying polypeptide interacts with a guide RNA, which is either a single RNA molecule or a RNA duplex of at least two RNA molecules, and is guided to a DNA sequence (e.g. a chromosomal sequence or an extrachromosomal sequence, e.g.
  • the site-directed polypeptide is a site-directed nuclease, which is able to cleave one or both strands of DNA at a specified target sequence.
  • cleavage refers to breaking of the covalent phosphodiester linkage in the ribosylphosphodiester backbone of a polynucleotide and encompass both singlestranded breaks and double-stranded breaks. Double-stranded cleavage can occur as a result of two distinct single-stranded cleavage events. Cleavage can result in the production of either blunt ends or staggered ends (also known as sticky ends).
  • a “nuclease cleavage site” or “genomic nuclease cleavage site” is a region of nucleotides within which a site-directed nuclease cleaves (e.g., when bound to a proximal binding site).
  • the polynucleotide is DNA (e.g., genomic DNA)
  • one or both strands can be cleaved at the nuclease cleavage site.
  • Such cleavage by the nuclease enzyme initiates DNA repair mechanisms within the cell, which establishes an environment for homologous recombination to occur.
  • a site-directed nuclease can be a naturally-occurring site-directed nuclease.
  • Exemplary naturally-occurring site-directed nucleases are known in the art (see for example, Makarova et al., 2017, Cell 168: 328-328. el, and Shmakov et al., 2017, Nat Rev Microbiol 15(3): 169- 182, both herein incorporated by reference).
  • a site-directed nuclease binds a DNA-targeting polynucleotide (e.g., a guide RNA) and is thereby directed to a specific sequence within a target DNA and cleaves the target DNA.
  • a DNA-targeting polynucleotide e.g., a guide RNA
  • the site-directed nuclease is modified from its natural sequence (e.g., via mutation or one or more amino acid residues) to change its function.
  • the site-directed nuclease may be modified to be enzymatically inactive.
  • the term “enzymatically inactive” can refer to a site-directed nuclease that can bind to a nucleic acid sequence in a polynucleotide in a sequence-specific manner, but may not cleave a target polynucleotide.
  • An enzymatically inactive site-directed polypeptide can comprise an enzymatically inactive domain (e.g., a nuclease domain). Enzymatically inactive can refer to no activity.
  • Enzymatically inactive can refer to substantially no activity. Enzymatically inactive can refer to essentially no activity. Enzymatically inactive can refer to an activity no more than 1%, no more than 2%, no more than 3%, no more than 4%, no more than 5%, no more than 6%, no more than 7%, no more than 8%, no more than 9%, or no more than 10% activity compared to a wild-type exemplary activity.
  • the site-directed nuclease comprises a CRISPR-associated (Cas) protein or a Cas nuclease that functions in a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)/Cas system.
  • CRISPR Clustered Regularly Interspaced Short Palindromic Repeats
  • this system can provide adaptive immunity against foreign DNA (Barrangou, R , et al, “CRISPR provides acquired resistance against viruses in prokaryotes, “Science (2007) 315: 1709-1712; Makarova, K.S., et al, “Evolution and classification of the CRISPR-Cas systems,” Nat Rev Microbiol (2011) 9:467- 477; Garneau, J.
  • CRISPR/Cas bacterial immune system cleaves bacteriophage and plasmid DNA,” Nature (2010) 468:67-71; Sapranauskas, R., et al, “The Streptococcus thermophilus CRISPR/Cas system provides immunity in Escherichia coli,” Nucleic Acids Res (2011) 39: 9275-9282).
  • a CRISPR/Cas system e.g., modified and/or unmodified
  • a CRISPR/Cas system can comprise a guide nucleic acid such as a guide RNA (gRNA) complexed with a Cas protein for targeted regulation of gene expression and/or activity or nucleic acid editing.
  • a guide nucleic acid such as a guide RNA (gRNA) complexed with a Cas protein for targeted regulation of gene expression and/or activity or nucleic acid editing.
  • An RNA- guided Cas protein e.g., a Cas nuclease such as a Cas9 nuclease
  • the Cas protein if possessing nuclease activity, can cleave the DNA (Gasiunas, G., et al, “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” Proc Natl Acad Sci USA (2012) 109: E2579-E2 86; Jinek, M., et al, “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012) 337:816-821; Sternberg, S.
  • DNA cleavage e.g., double-strand breaks
  • DNA break repair can occur via non-homologous end joining (NHEJ), microhomology -mediated end joining (MMEJ), or homology -directed repair (HDR).
  • NHEJ non-homologous end joining
  • MMEJ microhomology -mediated end joining
  • HDR homology -directed repair
  • donor nucleic acids are used to promote HDR, as detailed below in the “Systems” section.
  • CRISPR-Cas systems have been widely used for programmable genome editing in a variety of organisms and model systems (Cong, L., et al, “Multiplex genome engineering using CRISPR Cas systems,” Science (2013) 339:819-823; liang, W., et al, “RNA-guided editing of bacterial genomes using CRISPR-Cas systems,” Nat. Biotechnol. (2013) 31 : 233-239; Sander, J. D. & loung, J. K, “CRISPR-Cas systems for editing, regulating and targeting genomes,” Nature Biotechnol. (2014) 32:347-355).
  • the site-directed nuclease described herein comprises a Cas protein that forms a complex with a guide nucleic acid, such as a guide RNA (described further below in the “Systems” section).
  • the site-directed nuclease comprises a Cas protein that forms a complex with a single guide nucleic acid, such as a single guide RNA (sgRNA).
  • the site-directed nuclease comprises a RNA-binding protein (RBP) optionally complexed with a guide nucleic acid, such as a guide RNA (e.g., sgRNA), which is able to form a complex with a Cas protein.
  • RBP RNA-binding protein
  • RNA-guided Cas proteins recognize DNA targets that are complementary to a portion of the gRNA known as a CRISPR RNA (crRNA) sequence.
  • the target sequence is often referred to as a protospacer, and the part of the crRNA sequence that is complementary to the protospacer is often referred to as a spacer.
  • many Cas nucleases also require a specific protospacer adjacent motif (PAM), an approximately 2 to 6 base pair DNA sequence immediately following the protospacer sequence.
  • a Cas protein used herein can be an active variant, inactive variant, or fragment of a wildtype or modified Cas protein.
  • a Cas protein can comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof relative to a wild-type version of the Cas protein.
  • a Cas protein can be a polypeptide with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild-type exemplary Cas protein.
  • a Cas protein can be a polypeptide with at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% sequence identity and/or sequence similarity to a wild-type exemplary Cas protein.
  • Variants or fragments can comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild-type or modified Cas protein or a portion thereof. Variants or fragments can be targeted to a nucleic acid locus in complex with a guide nucleic acid while lacking nucleic acid cleavage activity.
  • a Cas protein can be modified to optimize regulation of gene expression.
  • a Cas protein can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and/or enzymatic activity.
  • Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the function of the protein or to optimize (e.g., enhance or reduce) the activity of the Cas protein for regulating gene expression.
  • variants of the polypeptides of this disclosure retain their respective biological activity, unless explicitly noted otherwise.
  • variants of a site-directed nuclease polypeptide retain the biological function of the full length, native sequence site directed nuclease.
  • variants of the nonspecific end-processing enzyme retain the biological function of the full length, native sequence nonspecific end-processing enzyme.
  • Modifications to any of the polypeptides or proteins provided herein are made by known methods.
  • modifications are made by site specific mutagenesis of nucleotides in a nucleic acid encoding the polypeptide, thereby producing a DNA encoding the modification, and thereafter expressing the DNA in recombinant cell culture to produce the encoded polypeptide.
  • Techniques for making substitution mutations at predetermined sites in DNA having a known sequence are well known. For example, Ml 3 primer mutagenesis and PCR-based mutagenesis methods can be used to make one or more substitution mutations.
  • Any of the nucleic acid sequences provided herein can be codon- optimized to alter, for example, maximize expression, in a host cell or organism.
  • the amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D-stereoisomers of the naturally occurring amino acids, unnatural amino acids and chemically modified amino acids.
  • Unnatural amino acids that is, those that are not naturally found in proteins
  • Zhang et al. Protein engineering with unnatural amino acids,” Curr. Opin. Struct. Biol. 23(4): 581-587 (2013); Xie et la. “Adding amino acids to the genetic repertoire,” 9(6): 548-54 (2005)); and all references cited therein.
  • B and y amino acids are known in the art and are also contemplated herein as unnatural amino acids.
  • a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified.
  • a side chain can be modified to comprise a signaling moiety, such as a fluorophore or a radiolabel.
  • a side chain can also be modified to comprise a new functional group, such as a thiol, carboxylic acid, or amino group.
  • Post- translationally modified amino acids are also included in the definition of chemically modified amino acids.
  • conservative amino acid substitutions can be made in one or more of the amino acid residues, for example, in one or more lysine residues of any of the polypeptides provided herein.
  • conservative amino acid substitutions can be made in one or more of the amino acid residues, for example, in one or more lysine residues of any of the polypeptides provided herein.
  • One of skill in the art would know that a conservative substitution is the replacement of one amino acid residue with another that is biologically and/or chemically similar.
  • the following eight groups each contain amino acids that are conservative substitutions for one another:
  • a DNA construct comprising a promoter operably linked to a recombinant nucleic acid encoding a fusion protein or domains thereof as described herein.
  • a nucleic acid is “operably linked” when it is placed into a functional relationship with another nucleic acid sequence.
  • Numerous promoters can be used in the constructs described herein.
  • a promoter is a region or a sequence located upstream and/or downstream from the start of transcription that is involved in recognition and binding of RNA polymerase and other proteins to initiate transcription.
  • promoter refers to a nucleotide sequence, usually upstream (5’) to its coding sequence, which controls the expression of the coding sequence by providing the recognition for RNA polymerase and other factors required for proper transcription.
  • Promoter regulatory sequences consist of proximal and more distal upstream elements. Promoter regulatory sequences influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, untranslated leader sequences, introns, and polyadenylation signal sequences. They include natural and synthetic sequences as well as sequences that may be a combination of synthetic and natural sequences.
  • An “enhancer” is a DNA sequence that can stimulate promoter activity and may be an innate element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of a promoter. It is capable of operating in both orientations (e.g., forward or reverse) and is capable of functioning even when moved either upstream or downstream from the promoter.
  • promoter includes “promoter regulatory sequences.”
  • promoters The choice of promoters to be included depends upon several factors, including, but not limited to, efficiency, selectability, inducibility, desired expression level, and cell- or tissue- preferential expression. It is a routine matter for one of skill in the art to modulate the expression of a sequence by appropriately selecting and positioning promoters and other regulatory regions relative to that sequence.
  • tissue specific promoters or tissue-preferred promoters
  • RNA synthesis may occur in other tissues at reduced levels. Since patterns of expression of a chimeric gene (or genes) introduced into a plant are controlled using promoters, there is an ongoing interest in the isolation of novel promoters that are capable of controlling the expression of a chimeric gene (or genes) at certain levels in specific tissue types or at specific plant developmental stages.
  • Certain promoters are able to direct RNA synthesis at relatively similar levels across all tissues of a plant. These are called “constitutive promoters" or “tissue-independent” promoters. Constitutive promoters can be divided into strong, moderate, and weak categories according to their effectiveness to directing RNA synthesis. Since it is necessary in many cases to simultaneously express a chimeric gene (or genes) in different tissues of a plant to get the desired functions of the gene (or genes), constitutive promoters are especially useful in this regard.
  • the recombinant nucleic acids provided herein can be included in expression cassettes for expression in a host cell or an organism of interest.
  • the cassette will include 5' and 3' regulatory sequences operably linked to a recombinant nucleic acid provided herein that allows for expression of a fusion protein.
  • the cassette may additionally contain at least one additional gene or genetic element to be cotransformed into the cell or organism. Where additional genes or elements are included, the components are operably linked. Alternatively, the additional gene(s) or element(s) can be provided on multiple expression cassettes.
  • Such an expression cassette is provided with a plurality of restriction sites and/or recombination sites for insertion of the polynucleotides to be under the transcriptional regulation of the regulatory regions.
  • the expression cassette may additionally contain a selectable marker gene.
  • the expression cassette will include in the 5' to 3' direction of transcription: a transcriptional and translational initiation region (i.e., a promoter), a polynucleotide of the invention, and a transcriptional and translational termination region (i.e., termination region) functional in the cell or organism of interest.
  • the promoters of the invention are capable of directing or driving expression of a coding sequence (i.e., a nucleic acid sequence that is transcribed into RNA such as mRNA, rRNA, tRNA, snRNA, ncRNA, IncRNA, sense RNA, or antisense RNA, regardless of whether the RNA is then translated to produce a protein) in a host cell.
  • the regulatory regions may be endogenous or heterologous to the host cell or to each other.
  • heterologous in reference to a sequence is a sequence that originates from a foreign species, or, if from the same species, is substantially modified from its native form in composition and/or genomic locus by deliberate human intervention.
  • Additional regulatory signals include, but are not limited to, transcriptional initiation start sites, operators, activators, enhancers, other regulatory elements, ribosomal binding sites, an initiation codon, termination signals, and the like. See Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.); Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, N.Y., and the references cited therein.
  • the expression cassette can also comprise a selectable marker gene for the selection of transformed cells.
  • Marker genes include genes conferring antibiotic resistance, such as those conferring hygromycin resistance, ampicillin resistance, gentamicin resistance, neomycin resistance, to name a few. Additional selectable markers are known and any can be used.
  • the various DNA fragments may be manipulated, so as to provide for the DNA sequences in the proper orientation and, as appropriate, in the proper reading frame.
  • adapters or linkers may be employed to join the DNA fragments or other manipulations may be involved to provide for convenient restriction sites, removal of superfluous DNA, removal of restriction sites, or the like.
  • in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, e.g., transitions and transversions may be involved.
  • a vector comprising a recombinant nucleic acid or DNA construct set forth herein.
  • the vector is contemplated to have the necessary functional elements that direct and regulate transcription of the inserted nucleic acid.
  • These functional elements include, but are not limited to, a promoter, regions upstream or downstream of the promoter, such as enhancers that may regulate the transcriptional activity of the promoter, an origin of replication, appropriate restriction sites to facilitate cloning of inserts adjacent to the promoter, antibiotic resistance genes or other markers which can serve to select for cells containing the vector or the vector containing the insert, RNA splice junctions, a transcription termination region, or any other region which may serve to facilitate the expression of the inserted gene or hybrid gene. See generally, Sambrook et al.
  • Transformation of a cell may be stable or transient.
  • a transgenic cell, plant cell, plant and/or plant part of the invention can be stably transformed or transiently transformed. Transformation can refer to the transfer of a nucleic acid molecule into the genome of a host cell, resulting in genetically stable inheritance.
  • the introduction into a plant, plant part and/or plant cell is via bacterial-mediated transformation, particle bombardment transformation, calcium-phosphate-mediated transformation, cyclodextrin- mediated transformation, electroporation, liposome-mediated transformation, nanoparticle- mediated transformation, polymer-mediated transformation, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, sonication, infdtration, polyethylene glycol-mediated transformation, protoplast transformation, or any other electrical, chemical, physical and/or biological mechanism that results in the introduction of nucleic acid into the plant, plant part and/or cell thereof, or any combination thereof.
  • Procedures for transforming plants are well known and routine in the art and are described throughout the literature.
  • Non-limiting examples of methods for transformation of plants include transformation via bacterial-mediated nucleic acid delivery (e.g. via bacteria from the genus Agrobacterium), viral-mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium-phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation,, sonication, infiltration, PEG-mediated nucleic acid uptake, as well as any other electrical, chemical, physical (mechanical) and/or biological mechanism that results in the introduction of nucleic acid into the plant cell, including any combination thereof.
  • Agrobacterium-mQ ⁇ iateA transformation is a commonly used method for transforming plants because of its high efficiency of transformation and because of its broad utility with many different species.
  • Agrobacterium-mQ a ⁇ e . transformation typically involves transfer of the binary vector carrying the foreign DNA of interest to an appropriate Agrobacterium strain that may depend on the complement of vir genes carried by the host Agrobacterium strain either on a co-resident Ti plasmid or chromosomally (Uknes et al. 1993, Plant Cell 5:159- 169).
  • the transfer of the recombinant binary vector to Agrobacterium can be accomplished by a tri-parental mating procedure using Escherichia coli carrying the recombinant binary vector, a helper E.
  • the recombinant binary vector can be transferred to Agrobacterium by nucleic acid transformation (Hbfgen and Willmitzer 1988, Nucleic Acids Res 16:9877).
  • Transformation of a plant by recombinant Agrobacterium usually involves co-cultivation of the Agrobacterium with explants from the plant and follows methods well known in the art. Transformed tissue is typically regenerated on selection medium carrying an antibiotic or herbicide resistance marker between the binary plasmid T-DNA borders.
  • Another method for transforming plants, plant parts and plant cells involves propelling inert or biologically active particles at plant tissues and cells. See, e.g., US Patent Nos. 4,945,050; 5,036,006 and 5, 100,792. Generally, this method involves propelling inert or biologically active particles at the plant cells under conditions effective to penetrate the outer surface of the cell and afford incorporation within the interior thereof.
  • the vector can be introduced into the cell by coating the particles with the vector containing the nucleic acid of interest.
  • a cell or cells can be surrounded by the vector so that the vector is carried into the cell by the wake of the particle.
  • Biolistic transformation refers to a method of introducing RNA or DNA into cells (e.g., plant cells) directly, in which RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cell (e.g., the plant cell) using high speed pressure to allow the RNA or DNA to penetrate the cell (e.g., to penetrate the plant cell wall).
  • heavy metal particles e.g., tungsten or gold
  • the CRISPR/Cas system can also be used to edit the genome of a host cell or organism.
  • the “CRISPR/Cas” system refers to a widespread class of bacterial systems for defense against foreign nucleic acid. Any of the CRISPR/Cas system components described herein may be used to introduce fusion proteins, recombinant nucleic acids, or systems into the genome of a host cell or organism. Methods for CRISPR/Cas system mediated genome editing are known in the art. It will be understood that use of a CRISPR/Cas system for introduction of fusion proteins, recombinant nucleic acids, or systems described herein into the genome of a host cell or organism is different from the particular methods and systems provided herein.
  • systems useful for editing one or more nucleic acids comprise one or more of the fusion proteins (or recombinant nucleic acids, constructs, vectors, or host cells) described above.
  • the systems further comprise one or more additional elements that are useful for editing one or more nucleic acids.
  • a system provided herein can further comprise a donor polynucleotide.
  • a system comprising a fusion protein comprising a Cas nuclease may further comprise one or more guide nucleic acids and/or one or more donor polynucleotide sequences.
  • the systems and methods described herein comprise at least one guide nucleic acid polynucleotide.
  • the systems and methods described herein comprise a plurality of guide nucleic acids
  • the polynucleotide can be deoxyribonucleic acid (DNA).
  • the DNA sequence can be single-stranded or doubled-stranded.
  • the at least one guide nucleic acid polynucleotide can be ribonucleic acid (guide RNA).
  • the nuclease can be complexed with the at least one guide RNA polynucleotide.
  • the at least one guide RNA polynucleotide can comprise a nucleic-acid targeting region that comprises a complementary sequence to a nucleic acid sequence on the targeted polynucleotide such as the targeted genomic loci or genes to confer sequence specificity of nuclease targeting.
  • the guide nucleic acid is a single guide nucleic acid comprising a crRNA. In some embodiments, the guide nucleic acid is a single guide nucleic acid comprising a crRNA but lacking a tracrRNA.
  • a crRNA can comprise the nucleic acid-targeting segment (e.g., spacer region) of the guide nucleic acid and a stretch of nucleotides that can form one half of a double-stranded duplex of the Cas protein-binding segment of the guide nucleic acid.
  • nucleic acid-targeting segment e.g., spacer region
  • the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer) is 20 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 19 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 18 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 17 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 16 nucleotides in length.
  • the nucleic acid-targeting region of a guide nucleic acid is 21 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 22 nucleotides in length.
  • the nucleotide sequence of the guide nucleic acid that is complementary to a nucleotide sequence (target sequence) of the target nucleic acid can have a length of, for example, at least about 12 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt or at least about 40 nt.
  • the nucleotide sequence of the guide nucleic acid that is complementary to a nucleotide sequence (target sequence) of the target nucleic acid can have a length of from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 45 nt, from about 12 nt to about 40 nt, from about 12 nt to about 35 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about 19 nt, from about 19 nt to about 20 nt, from about 19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50
  • a protospacer sequence of a targeted polynucleotide can be identified by identifying a protospacer-adjacent motif (PAM) within a region of interest and selecting a region of a desired size upstream or downstream of the PAM as the protospacer.
  • a corresponding spacer sequence can be designed by determining the complementary sequence of the protospacer region.
  • a spacer sequence can be identified using a computer program (e.g., machine readable code).
  • the computer program can use variables such as predicted melting temperature, secondary structure formation, and predicted annealing temperature, sequence identity, genomic context, chromatin accessibility, % GC, frequency of genomic occurrence, methylation status, presence of SNPs, and the like.
  • the percent complementarity between the nucleic acid-targeting sequence (e.g., a spacer sequence of the at least one guide polynucleotide as disclosed herein) and the target nucleic acid (e.g., a protospacer sequence of the one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%.
  • the percent complementarity between the nucleic acid-targeting sequence and the target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 contiguous nucleotides.
  • Guide nucleic acids of the systems of the disclosure can include modifications or sequences that provide for additional desirable features (e g., modified or regulated stability; subcellular targeting; tracking with a fluorescent label; a binding site for a protein or protein complex; and the like).
  • modifications include, for example, a 5' cap (a 7- m ethyl guanylate cap (m7G)); a 3' polyadenylated tail (a 3' poly(A) tail); a riboswitch sequence (e.g., to allow for regulated stability and/or regulated accessibility by proteins and/or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (a hairpin)); a modification or sequence that targets the RNA to a subcellular location (e g., nucleus, mitochondria, chloroplasts, and the like); a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, and so forth); a modification or sequence that provides a binding site for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyl transferases, DNA de
  • a guide nucleic acid can comprise one or more modifications (e.g., a base modification, a backbone modification), to provide the nucleic acid with a new or enhanced feature (e.g., improved stability).
  • a guide nucleic acid can comprise a nucleic acid affinity tag.
  • a nucleoside can be a base-sugar combination. The base portion of the nucleotide can be a heterocyclic base. The two most common classes of such heterocyclic bases are the purines and the pyrimidines.
  • Nucleotides can be nucleosides that further include a phosphate group covalently linked to the sugar portion of the nucleoside.
  • the phosphate group can be linked to the 2', the 3', or the 5' hydroxyl moiety of the sugar.
  • the phosphate groups can covalently link adjacent nucleosides to one another to form a linear polymeric compound.
  • the respective ends of this linear polymeric compound can be further joined to form a circular compound; however, linear compounds can be suitable.
  • linear compounds can have internal nucleotide base complementarity and can therefore fold in a manner as to produce a fully or partially double-stranded compound.
  • the phosphate groups can commonly be referred to as forming the internucleoside backbone of the guide nucleic acid.
  • the linkage or backbone of the guide nucleic acid can be a 3' to 5' phosphodiester linkage.
  • the at least one guide RNA polynucleotide of a system or method provided herein can bind to at least a portion of a genome (e.g., a plant genome) or a gene (e g., a plant gene).
  • the at least one guide RNA polynucleotide is capable of forming a complex with a site-directed nuclease to direct the site-directed nuclease to target the portion of a target nucleic acid (e.g., a site in a genome or a gene).
  • the systems described herein comprise at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides that are able to form a complex with a site-directed nuclease.
  • a mutant Mb2Casl2a polypeptide comprising at least one amino acid substitution introduced into a wild-type Mb2Casl2a polypeptide sequence of SEQ ID NO: 1.
  • the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
  • the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q.
  • the mutant Mb2Casl2a polypeptide comprises a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG.
  • the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
  • Described herein is a method of editing a plant genome, comprising contacting said plant genome with the mutant Mb2Casl2a polypeptide recited above.
  • the method further comprises a guide RNA
  • the guide RNA is encoded by a sequence comprising SEQ ID NOs: 16-60.
  • Described herein is a construct or plasmid comprising a polynucleotide sequence encoding for the mutant Mb2Casl2a polypeptides recited above.
  • One embodiment is a non-human cell comprising the construct or plasmid.
  • RNP complex comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
  • Example 1 Altering amino acid residues can increase the PAM spectrum of Mb2Casl2a.
  • the nucleotide sequence for wildtype sequence of Mb2Casl2a (encoding the amino acid sequence of SEQ ID NO: 1) was subjected to various means of mutagenesis, including targeted mutagenesis, domain variants, and molecular breeding (also known as family DNA shuffling) in order to generate a library of diverse mutants. Mutants were first screened using E. coli plasmid clearance assay at the First Tier screening, comprising First and Second Stages. This assay uses an E. coli host with a resident plasmid containing the conditional lethal gene ccdB under control of the arabinose-inducible araC promoter.
  • the transformation mixture was plated on media containing ampicillin, kanamycin, IPTG, and arabinose. Only cells expressing active Mb2Casl2a could cut the resident plasmid, which was then degraded, thus eliminating ccdB expression and allowing those cells to survive. Negative controls should result in no colonies. Likewise, if the wild-type Mb2Casl2a had an absolute requirement for TTV, TTTV PAM, none of the VTV variant PAMs should allow cutting and no colonies would survive.
  • the Mb2Casl2a expression plasmids encoding variants that can utilize NVTV PAMs were recovered and moved to the Second Stage selection.
  • the Second Stage utilized an analogous screening pool, with three E. coll strains carrying resident plasmids comprising the degenerate PAM TTV (TTA, TTC, TTG).
  • the result of this two-step process identified variants that react with non-canonical PAMs but also reacted with the wild-type PAM.
  • the Mb2Casl2a expression plasmids were recovered from surviving colonies either as individuals or as a pool. Pooled plasmid DNA from surviving colonies was retransformed into the screening host as a further enrichment step for Mb2Casl2a variants with altered PAM preference in an iterative process. Plasmid DNA from individual colonies also was subjected to DNA sequencing to determine which amino acid positions have been changed.
  • yeast cells can grow over a wider range of temperatures (e.g., 20-37°C).
  • Yeast strains were generated in which the Ade2 gene is interrupted by a Ura3 gene adjacent to a Mb2Casl2a target site preceded by each of the 12 PAM sequences comprising NTV.
  • the Ade2 gene is split in such a way as to leave 200 bp of Ade2 sequence homology on either side of the insertion. These homology arms constitute sites for recombination.
  • yeast colonies exhibit a cell-autonomous red pigmentation.
  • the screening host consisted of a pool of strains as in the bacterial assay described above.
  • Mb2Casl2a nuclease expression was controlled by the GALI galactose-inducible promoter on a plasmid that also expresses the crRNA.
  • Mb2Casl2a variants were transformed into the screening pool. Upon Mb2Casl2a cutting, single-strand annealing repair resulted in an intact Ade2 gene and functional enzyme. A negative control, or inactive Mb2Casl2a variant, would result in 100% red colonies. Wildtype Mb2Casl2a, if it had a strict TTV PAM requirement, would result in 25% white colonies. The ideal Mb2Casl2a variant with a preference for a PAM equal to NTV would result in 100% white colonies. Sectored colonies represent Mb2Casl2a variants and/or PAM sequences that result in incomplete cutting.
  • the frequencies of red/white/sectored colonies provide an indication of the PAM preference and activity of each Mb2Casl2a variant.
  • Sequence analysis of colonies with various red/white/sectored phenotypes indicated the PAM requirements of each Mb2Casl2a variant.
  • This strategy allowed for quantification of Mb2Casl2a efficiency at different PAM sites using the same target sequence to interrogate each variant PAM.
  • This approach is based on a set of yeast strains in which variant PAM sequences comprising NTV (or NNTV) are introduced just upstream of the Kozak sequence preceding the Ade2 coding sequence.
  • the Mb2Casl2a nuclease cuts at 18/23 bases downstream of the PAM and produces indels, which will put the Ade2 coding sequence out of frame.
  • the red/white phenotypic readout and quantification would be similar to the previously described assay.
  • Example 2 In vitro assay for Mb2Casl2a screening.
  • the variants were assayed in vitro.
  • Target DNA containing randomly combined four nucleotides (NNNN) followed by a validated gRNA sequence were synthesized. DNA sequences were then incubated with Mb2Casl2a variants. With carefully designed PCR and NGS sequencing, the cut efficiency of each NNNN (total 256 combinations) were determined which indicate the recognition of Mb2Casl2a to that sequence.
  • the DNA substrates used comprised, in the 5’ to 3’ direction, a 5’ barcode sequence, a four random nucleotide residues, a validated target sequence, and a 3’ barcode sequence.
  • These DNA substrates were synthesized by Azenta Inc., and are represented by SEQ ID NO: 62 or SEQ ID NO: 63 (ZmGL2 and ZmmiR528, respectively).
  • RNP was assembled for each variant by incubating the following recipe for 30 minutes at 25°C:
  • reaction buffer and DNA substrates were added to each assembled RNP and incubated for one hour at 37°C:
  • DNA was recovered from the reaction tubes and mixed into one tube for next-generation sequencing.
  • Example 3 Assaying the PAM spectrum in prokaryotic cells.
  • the PAMDA consists of 2 components.
  • the first component is an ampicillin resistance plasmid library of gRNA target sequence contains 48 PAMs (NNTV) with same target sequence: GTGATAAGTGGAATGGCATGTGGG.
  • the 2 nd component is an E.
  • coli strain harboring a low copy number pl5A origin chloramphenicol resistant plasmid which overexpresses one of engineered Mb2Casl2a variants, wildtype Mb2Casl2a or dead dMb2Casl2a (D864A, E958A) driven by T7 promoter and a gRNA expression cassette recognizes the target sequence in the PAM library.
  • Transform the PAM plasmid library into competent E. coli cell prepared from the over expression Mb2Casl2a variant and gRNA expression cassette, further selected on Lb agar plate containing chloramphenicol and ampicillin antibiotics.
  • the gRNA recognizes the targeted sequence on the plasmid containing correct PAM
  • the plasmid will be cut by Casl2a and further be degraded.
  • the survival plasmids contain PAM sequence that can’t be recognized by co-transformed Mb2Casl2a variant. Pool the survival colonies and NGS analysis the remaining PAM abundance. More abundant PAMs mean less cutting and less abundant PAMs means more cutting.
  • the PAMs abundance of each Mb2Cas 12a variant is normalized to dMb2Casl2a. E. coli screening tests were performed as described in Example 1.
  • Example 4 Assaying the PAM spectrum in eukaryotic cells.
  • Maize mesophyll protoplasts were prepared for transfection with vectors expressing Mb2Casl2a variants and gRNAs. See, e.g., M.R. Coy, et al., Protoplast Isolation and Transfection in Maize, in PROTOPLAST TECHNOLOGY METHODS AND PROTOCOLS, 91-104 (K. Wang & F. Zhang, eds., 2022) (detailing a standard protoplast transfection protocol).
  • Mb2 Casl2a which contained the specific position mutations were maize codon optimized and synthesized commercially (GenScript, Nanjing, China), and cloned under the sugarcane Ubiquitin-4 (SoUbi4) gene promoter to generate the plant expression vector.
  • the CRISPR/Casl2a guide RNA transcript is expressed under the control of the OsU6 promoter which targets the corn native gene. It also included direct repeat of Mb2Casl2a crRNA AATTTCTACTGTTTGTAGAT (SEQ ID NO: 61) as the scaffold.
  • Corn protoplast transformation of expression vector Plasmid DNA was purified by Plasmid DNA Purification using QIAGEN Plasmid Plus Midi Kits (Catalog No. 12945). 20 pg Casl2a expression plasmid and 7.5 pg gRNA expression plasmid (2 pmol in each) were mixed into 20 pl solution, and transfected by PEG method into 5 x 10 4 corn protoplast cell. Each treatment have three transformations as well as the repeat.
  • Genomic DNA extraction after cell harvesting Cells were collected after 48 hours, centrifugated 5 min at 100g, and removed 250 pl supernatant. Added 100 pl Lysis buffer, and placed on a shaker for 5 min at room temperature (“RT”). Centrifuged for 15 min at 4000 rpm.
  • PCR using the appropriate primers to amplify the DNA fragment that contained target site, was conducted following the recipes and PCR conditions below.
  • PCR product purification protocol Add Agencourt AMPure XP (54 pl beads to 30 pl PCR) Let the mixed samples incubate for 5 minutes at room temperature for maximum recovery. Place the reaction plate onto an Magnet Plate for 2 minutes to separate beads from the solution. Dispense 200 pl of 70% ethanol to each well of the reaction plate and incubate for 30 seconds at room temperature. Aspirate out the ethanol and discard. Remove the reaction plate from the magnet plate, and then add 35 pl warm water. Incubate for 2 minutes. Transfer the eluate into another clean plate.
  • T7EI detection of Editing Anneal PCR product using below recipe and thermocycler conditions of 95C 5min; 95-85C -2C/sec; 85-25C -O.lC/sec; 4C hold.
  • gRNA sequences gRNA expression vector SEQ ID NO: PAM Target gene gRNA-A-1 SEQ ID NO: 16 TTTG Waxyl gRNA-A-6 SEQ ID NO: 17 GATA Bx9 gRNA-A-7 SEQ ID NO: 18 GATC Bx9 gRNA-A-8 SEQ ID NO: 19 GATG Bx9 gRNA- A- 12 SEQ ID NO: 20 GTTA Bx9 gRNA-A-13 SEQ ID NO: 21 GTTC Waxyl gRNA-A-14 SEQ ID NO: 22 GTTG Waxyl gRNA-A-15 SEQ ID NO: 23 TTTA MLO2c gRNA-A-19 SEQ ID NO: 24 CATA Waxyl gRNA-A-20 SEQ ID NO: 25 CATC MLO2c gRNA-A-21 SEQ ID NO: 26 CATG Bx9 gRNA-A-25 SEQ ID NO: 27 CTTA MLO2c gRNA-A-26 SEQ ID NO: 28 CTTC MLO2c

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Biomedical Technology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Organic Chemistry (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Biotechnology (AREA)
  • General Engineering & Computer Science (AREA)
  • Molecular Biology (AREA)
  • Microbiology (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Biophysics (AREA)
  • Plant Pathology (AREA)
  • Physics & Mathematics (AREA)
  • Medicinal Chemistry (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Cell Biology (AREA)
  • Micro-Organisms Or Cultivation Processes Thereof (AREA)
  • Peptides Or Proteins (AREA)
  • Enzymes And Modification Thereof (AREA)

Abstract

Disclosed are mutant Mb2Cas12a polypeptide variants of the wildtype that recognize alternative PAM sites. This variants comprise at least one amino acid substitution introduced into a wild-type Mb2Cas12a polypeptide sequence of SEQ ID NO: 1. This substitution can occur at the following positions: D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.

Description

Mb2Casl2a VARIANTS WITH FLEXIBLE PAM SPECTRUM
FIELD OF THE INVENTION
Described are Casl2a mutants from Moraxella bovoculi AAX08 and methods for use thereof. These mutants have a broader PAM spectrum than the wildtype enzyme.
CLAIM FOR PRIORITY
This application claims priority under 35 U.S.C. § 119 to PCT Application No. PCT/CN2023/073487, filed January 27, 2023, the contents of which are incorporated herein by reference in their entirety.
SEQUENCE LISTING
This application is accompanied by a sequence listing entitled 82857_PCT.xml, created January 11, 2024, which is approximately 80 kilobytes in size. This sequence listing is incorporated herein by reference in its entirety.
BACKGROUND
Mb2Casl2a from Moraxella bovoculi AAX08 has demonstrated in planta genome editing capability (Zhang et al., 2021), however its PAM site is less flexible compared to Cas9. Therefore, fewer targets in eukaryotic genomes can be edited with Mb2Casl2a compared to Cas9. With its distinct property of high performance at lower temperature, Mb2Casl2a needs to have a flexible PAM.
SUMMARY
Therefore, there is a need for a mutant Mb2Casl2a polypeptide variant of the wildtype that recognizes an alternative PAM site. This variant may comprise at least one amino acid substitution introduced into a wild-type Mb2Casl2a polypeptide sequence of SEQ ID NO: 1. This substitution can occur at the following positions: D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
Described herein is a mutant Mb2Casl2a polypeptide comprising at least one amino acid substitution introduced into a wild-type Mb2Casl2a polypeptide sequence of SEQ ID NO: 1. In one embodiment, the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834. In another embodiment, the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q. In another embodiment, the mutant Mb2Casl2a polypeptide comprises a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66. In one embodiment, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV. In another embodiment, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG. In another embodiment, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG. In another asp embodiment ect, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG. In another embodiment, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
Described herein is a method of editing a plant genome, comprising contacting said plant genome with the mutant Mb2Casl2a polypeptide recited above. In one embodiment, the method further comprises a guide RNA. In another embodiment, the guide RNA is encoded by a sequence comprising SEQ ID NOs: 16-60.
Described herein is an edited plant obtained by the methods recited above.
Described herein is a construct or plasmid comprising a polynucleotide sequence encoding for the mutant Mb2Casl2a polypeptides recited above. One embodiment is a non-human cell comprising the construct or plasmid.
Described herein is an RNP complex comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66. BRIEF DESCRIPTION OF THE SEQUENCES IN THE SEQUENCE LISTING
SEQ ID NO: 1 is the amino acid sequence of wildtype Mb2Casl2a, also referred to as “wildtype” or “WT” throughout.
SEQ ID NO: 2 is the amino acid sequence of Mb2Casl2a variant named “46(172R)” comprising D172R, N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
SEQ ID NO: 3 is the amino acid sequence of Mb2Casl2a variant named “46(172Y)” comprising D172Y, N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
SEQ ID NO: 4 is the amino acid sequence of Mb2Casl2a variant named “46” comprising N563R, K569V, N573R, K625R, S758E, E759T, D802L, F824Y, and F834Q amino acid substitutions.
SEQ ID NO: 5 is the amino acid sequence of Mb2Casl2a variant named "51(172R)" comprising D172R, N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
SEQ ID NO: 6 is the amino acid sequence of Mb2Casl2a variant named "51(172Y)" comprising D172Y, N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
SEQ ID NO: 7 is the amino acid sequence of Mb2Casl2a variant named "51" comprising N563R, K569V, N573R, S758E, E759C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
SEQ ID NO: 8 is the amino acid sequence of Mb2Casl2a variant named "18" comprising D172R, G672S, E678D, and Q725E amino acid substitutions.
SEQ ID NO: 9 is the amino acid sequence of Mb2Casl2a variant named "21" comprising D172R, N573D, A629S, L633I, Y636F, L642I, L695I, Q708K, Q725E, F753W, S758D, L782I, and L826F amino acid substitutions.
SEQ ID NO: 10 is the amino acid sequence of Mb2Casl2a variant named "28" comprising K569T, E570R, D572R, N573D, S758E, E795C, D783Q, and F824Y amino acid substitutions. SEQ ID NO: 11 is the amino acid sequence of Mb2Casl2a variant named "28(172Y)" comprising D172Y, K569T, E570R, D572R, N573D, S758E, E795C, D783Q, and F824Y amino acid substitutions.
SEQ ID NO: 12 is the amino acid sequence of Mb2Casl2a variant named "44" comprising N563R, K569V, N573R, S758E, E795C, T788F, and F824M amino acid substitutions.
SEQ ID NO: 13 is the amino acid sequence of Mb2Casl2a variant named “44 (172R)” comprising D172R, N563R, K569V, N573R, S758E, E759C, T788F, F824M amino acid substitutions.
SEQ ID NO: 14 is the amino acid sequence of Mb2Casl2a variant named “44 (172Y)” comprising D172Y, N563R, K569V, N573R, S758E, E759C, T788F, F824M amino acid substitutions.
SEQ ID NO: 15 is the amino acid sequence of wildtype LbCasl2a.
SEQ ID NOs: 16-60 are the gRNA sequences used. See Table 5
SEQ ID NO: 61 is the direct repeat of Mb2Casl2a crRNA.
SEQ ID NO: 62 is a synthetic ZmGL2 DNA substrate.
SEQ ID NO: 63 is a synthetic ZmmiR528 DNA substrate.
SEQ ID NO: 64 is the amino acid sequence of Mb2Casl2a variant named “27” comprising K569T, E570R, D572R, N573D, A671M, Q705H, S758E and E795C amino acid substitutions.
SEQ ID NO: 65 is the amino acid sequence of Mb2Casl2a variant named “43” comprising N563R, K569V, N573R, S758C, D783Q, T788F, D802L, and F824Y amino acid substitutions.
SEQ ID NO: 66 is the amino acid sequence of Mb2Casl2a variant named “45” comprising N563R, K569V, N573R, K625R, S696F, S758E, E759T, and T831N amino acid substitutions.
DEFINITIONS
All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art.
References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques and/or substitutions of equivalent techniques that would be apparent to one of skill in the art. While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject.
As used herein, the singular forms “a”, “an” and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “an enzyme” optionally includes a combination of two or more such molecules, and the like.
As used herein, “and/or” refers to and encompasses any and all possible combinations of one or more of the associated listed items.
The term “about” as used herein refers to the usual error range for the respective value readily known to the skilled person in this technical field, for example ± 20%, ± 10%, or ± 5%, are within the intended meaning of the recited value.
As used herein, the term “comprising” or “comprise” is open-ended. When used in connection with a subject nucleic acid (or amino acid sequence), it refers to a nucleic acid sequence (or an amino acid sequence) that includes the subject sequence as a part or as its entire sequence.
As used herein, the transitional phrase “consisting essentially of’ means that the scope of a claim is to be interpreted to encompass the specified materials or steps recited in the claim and those that do not materially affect the basic and novel characteristic(s) of the claimed matter. Thus, the term “consisting essentially of’ when used in a claim of this disclosure is not intended to be interpreted to be equivalent to “comprising.”
The term “plurality” refers to more than one entity. Thus, a “plurality of individuals” refers to at least two individuals. In some embodiments, the term plurality refers to more than half of the whole. For example, in some embodiments a “plurality of a population” refers to more than half the members of that population.
The term “plant” as used herein refers to any plant at any stage of development, particularly a seed plant. The term “plant cell” as used herein refers to a structural and physiological unit of a plant, comprising a protoplast and a cell wall. The plant cell may be in form of an isolated single cell or a cultured cell, or as a part of higher organized unit such as, for example, plant tissue, a plant organ, or a whole plant. The plant cell may be derived from or part of an angiosperm or gymnosperm. The plant cell may be a monocotyledonous plant cell (e.g., a maize cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a turf grass cell, or an ornamental grass cell) or a dicotyledonous plant cell (e.g., a tobacco cell, a pepper cell, an eggplant cell, a sunflower cell, a crucifer cell, a flax cell, a potato cell, a cotton cell, a soybean cell, a sugar beet cell, or an oilseed rape cell. The term “plant cell culture” as used herein refers to cultures of plant units such as, for example, protoplasts, cell culture cells, cells in plant tissues, pollen, pollen tubes, ovules, embryo sacs, zygotes and embryos at various stages of development. The term “plant tissue” as used herein refers to a group of plant cells organized into a structural and functional unit. Any tissue of a plant in planta or in culture is included. This term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue culture and any group of plant cells organized into structural and/or functional units. The use of this term in conjunction with, or in the absence of, any specific type of plant tissue as listed above or otherwise embraced by this definition is not intended to be exclusive of any other type of plant tissue. The term “plant part” as used herein refers to a part of a plant, including single cells and cell tissues such as plant cells that are intact in plants, cell clumps and tissue cultures from which plants can be regenerated. Examples of plant parts include, but are not limited to, single cells and tissues from pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, flower parts, fruits, stems, shoots, cuttings, and seeds; as well as pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, flower parts, fruits, stems, shoots, cuttings, scions, rootstocks, seeds, protoplasts, calli, and the like.
The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, the terms encompass amino acid chains of any length, including full-length proteins, wherein the amino acid residues are linked by covalent peptide bonds.
The terms “nucleic acid” and “polynucleotide” are used interchangeably and as used herein refer to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single- or double-stranded form, as well as to both sense and anti-sense strands of RNA, cDNA, genomic DNA, mitochondrial DNA, and synthetic forms and mixed polymers of the above. In higher plants, DNA is the genetic material while RNA is involved in the transfer of information contained within DNA into proteins. A “genome” is the entire body of genetic material contained in each cell of an organism. It is understood that when an RNA is described, its corresponding cDNA is also described, wherein uridine is represented as thymidine. In particular embodiments, a nucleotide refers to a ribonucleotide, deoxynucleotide or a modified form of either type of nucleotide, and combinations thereof. In addition, a polynucleotide disclosed herein may include either or both naturally occurring and modified nucleotides linked together by naturally occurring and/or non-naturally occurring nucleotide linkages. The nucleic acid molecules may be modified chemically or biochemically or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more of the naturally occurring nucleotides with an analogue, internucleotide modifications such as uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, and the like), charged linkages (e.g., phosphorothioates, phosphorodithioates, and the like), pendent moi eties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, and the like), chelators, alkylators, and modified linkages (e.g., alpha anomeric nucleic acids, and the like). The above term is also intended to include any topological conformation, including single- stranded, double-stranded, partially duplexed, triplex, hairpinned, circular and padlocked conformations. A reference to a nucleic acid sequence encompasses its complement unless otherwise specified. Thus, a reference to a nucleic acid molecule having a particular sequence should be understood to encompass its complementary strand, with its complementary sequence. Nucleotide sequences are “complementary” when they specifically hybridize in solution (e.g., according to Watson-Crick base pairing rules). The term also includes codon-optimized nucleic acids that encode the same polypeptide sequence. It is also understood that nucleic acids can be unpurified, purified, or attached, for example, to a synthetic material such as a bead or column matrix.
Nucleic acid residues can be referred to by individual letters: “A” is adenine; “C” is cytosine; “G” is guanine; “T” is thymine; “N” is any nucleotide; “V” is nucleotide A, C, or G.
The term “corresponding to” in the context of nucleic acid sequences means that when the nucleic acid sequences of certain sequences are aligned with each other, the nucleic acids that “correspond to” certain enumerated positions in the present invention are those that align with these positions in a reference sequence, but that are not necessarily in these exact numerical positions relative to a particular nucleic acid sequence of the invention. Optimal alignment of sequences for comparison can be conducted by computerized implementations of known algorithms, or by visual inspection. Readily available sequence comparison and multiple sequence alignment algorithms are, respectively, the Basic Local Alignment Search Tool (BLAST) and ClustalW/ClustalW2/Clustal Omega programs available on the Internet (e g., the website of the EMBL-EBI). Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and FASTA, which are part of the Accelrys GCG Package available from Accelrys, Inc. of San Diego, Calif., United States of America. See also Smith & Waterman, 1981; Needleman & Wunsch, 1970; Pearson & Lipman, 1988; Ausubel et al., 1988; and Sambrook & Russell, 2001.
Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, SNPs, and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions may be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed- base and/or deoxyinosine residues. See Batzer et al., Nucleic Acid Res. 19:5081 (1991);
Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994).
The terms “ identity” or “substantial identity,” as used in the context of a polynucleotide or polypeptide sequence described herein, refers to a sequence that has at least 60% sequence identity to a reference sequence. Alternatively, percent identity can be any integer from 60% to 100%. Exemplary embodiments include at least: 60%, 65%, 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, as compared to a reference sequence using the programs described herein; preferably BLAST using standard parameters, as described below. One of skill will recognize that these values can be appropriately adjusted to determine corresponding identity of proteins encoded by two nucleotide sequences by taking into account codon degeneracy, amino acid similarity, reading frame positioning and the like.
For sequence comparison, typically one sequence acts as a reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters.
A “comparison window,” as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of from 20 to 600, usually about 50 to about 200, more usually about 100 to about 150 in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well-known in the art. Optimal alignment of sequences for comparison may be conducted by the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981), by the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson and Lipman Proc. Natl. Acad. Sci. (U.S.A.) 85: 2444 (1988), by computerized implementations of these algorithms (e g., BLAST), or by manual alignment and visual inspection.
A “gene” is a defined region that is located within a genome and that, besides the aforementioned coding nucleic acid sequence, comprises other, primarily regulatory, nucleic acid sequences responsible for the control of the expression, that is to say the transcription and translation, of the coding portion. Genes can include both coding and non-coding regions (e g., introns, regulatory elements, promoters, enhancers, termination sequences and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or specific protein, including regulatory sequences. Genes may or may not be capable of being used to produce a functional protein. In some embodiments, a gene refers to only the coding region. The term “native gene” refers to a gene as found in nature The term “chimeric gene” refers to any gene that contains 1) DNA sequences, including regulatory and coding sequences that are not found together in nature, or 2) sequences encoding parts of proteins not naturally adjoined, or 3) parts of promoters that are not naturally adjoined. Accordingly, a chimeric gene may comprise regulatory sequences and coding sequences that are derived from different sources, or comprise regulatory sequences and coding sequences derived from the same source, but arranged in a manner different from that found in nature. A gene may be “isolated” by which is meant a nucleic acid molecule that is substantially or essentially free from components normally found in association with the nucleic acid molecule in its natural state. Such components include other cellular material, culture medium from recombinant production, and/or various chemicals used in chemically synthesizing the nucleic acid molecule.
An “isolated” nucleic acid molecule or nucleotide sequence or an “isolated” polypeptide is a nucleic acid molecule, nucleotide sequence or polypeptide that, by the hand of man, exists apart from its native environment and/or has a function that is different, modified, modulated and/or altered as compared to its function in its native environment and is therefore not a product of nature. An isolated nucleic acid molecule or isolated polypeptide may exist in a purified form or may exist in a non-native environment such as, for example, a recombinant host cell. Thus, for example, with respect to polynucleotides, the term isolated means that it is separated from the chromosome and/or cell in which it naturally occurs. A polynucleotide is also isolated if it is separated from the chromosome and/or cell in which it naturally occurs and is then inserted into a genetic context, a chromosome, a chromosome location, and/or a cell in which it does not naturally occur. The recombinant nucleic acid molecules and nucleotide sequences of the invention can be considered to be “isolated” as defined above.
Thus, an “isolated nucleic acid molecule” or “isolated nucleotide sequence” is a nucleic acid molecule or nucleotide sequence that is not immediately contiguous with nucleotide sequences with which it is immediately contiguous (one on the 5' end and one on the 3' end) in the naturally occurring genome of the organism from which it is derived. Accordingly, in one embodiment, an isolated nucleic acid includes some or all of the 5' non-coding (e.g., promoter) sequences that are immediately contiguous to a coding sequence. The term therefore includes, for example, a recombinant nucleic acid that is incorporated into a vector, into an autonomously replicating plasmid or virus, or into the genomic DNA of a prokaryote or eukaryote, or which exists as a separate molecule (e.g., a cDNA or a genomic DNA fragment produced by PCR or restriction endonuclease treatment), independent of other sequences. It also includes a recombinant nucleic acid that is part of a hybrid nucleic acid molecule encoding an additional polypeptide or peptide sequence. An “isolated nucleic acid molecule” or “isolated nucleotide sequence” can also include a nucleotide sequence derived from and inserted into the same natural, original cell type, but which is present in a nonnatural state, e.g., present in a different copy number, and/or under the control of different regulatory sequences than that found in the native state of the nucleic acid molecule.
The term “isolated” can further refer to a nucleic acid molecule, nucleotide sequence, polypeptide, peptide or fragment that is substantially free of cellular material, viral material, and/or culture medium (e.g., when produced by recombinant DNA techniques), or chemical precursors or other chemicals (e.g., when chemically synthesized). Moreover, an “isolated fragment” is a fragment of a nucleic acid molecule, nucleotide sequence or polypeptide that is not naturally occurring as a fragment and would not be found as such in the natural state. “Isolated” does not necessarily mean that the preparation is technically pure (homogeneous), but it is sufficiently pure to provide the polypeptide or nucleic acid in a form in which it can be used for the intended purpose.
“Homology dependent repair” or “homology directed repair” or “HDR” refers to a mechanism for repairing ssDNA and double stranded DNA (dsDNA) damage in cells. This repair mechanism can be used by the cell when there is an HDR template with a sequence with significant homology to the injury site. The term “perfect HDR” refers to a situation in which genomic-homology junctions in the replaced allele underwent complete HDR and “imperfect HDR” refers to a situation in which genomic-homology junctions in the replaced allele underwent partial or incomplete HDR. a donor DNA molecule with homology to the cleaved target DNA sequence is used as a template for repair of the cleaved target DNA sequence, resulting in the transfer of genetic information from the donor polynucleotide to the target DNA. As such, new nucleic acid material may be inserted/copied into the site. In some cases, a target DNA is contacted with a donor molecule, for example a donor DNA molecule. In some cases, a donor DNA molecule is introduced into a cell. In some cases, at least a segment of a donor DNA molecule integrates into the genome of the cell.
“Microhomology-mediated end joining” or “MMEJ” or “alternative nonhomologous endjoining” (Alt-NHEJ) refers to a form of repairing double-stranded breaks in DNA. This repair mechanism utilizes microhomologous sequences to align the broken strands “Nonhomologous end joining” or “NHEJ” refers to a form of repairing double-stranded breaks in DNA. The double-strand breaks are repaired by direct ligation of the break ends to one another. Generally, no new nucleic acid material is inserted into the site, although some nucleic acid material may be lost or added, resulting in a small deletion or a small insertion.
The proteins provided herein comprise a site-directed polypeptide. A site-directed modifying polypeptide modifies target DNA (e.g., via cleavage or methylation of target DNA) and/or a polypeptide associated with target DNA (e.g., methylation or acetylation of a histone tail). In some embodiments, a site-directed modifying polypeptide interacts with a guide RNA, which is either a single RNA molecule or a RNA duplex of at least two RNA molecules, and is guided to a DNA sequence (e.g. a chromosomal sequence or an extrachromosomal sequence, e.g. an episomal sequence, a minicircle sequence, a mitochondrial sequence, a chloroplast sequence, etc.) by virtue of its association with the guide RNA. In some embodiments, the site-directed polypeptide is a site-directed nuclease, which is able to cleave one or both strands of DNA at a specified target sequence.
The term “cleavage” or “cleaving” refers to breaking of the covalent phosphodiester linkage in the ribosylphosphodiester backbone of a polynucleotide and encompass both singlestranded breaks and double-stranded breaks. Double-stranded cleavage can occur as a result of two distinct single-stranded cleavage events. Cleavage can result in the production of either blunt ends or staggered ends (also known as sticky ends). A “nuclease cleavage site” or “genomic nuclease cleavage site” is a region of nucleotides within which a site-directed nuclease cleaves (e.g., when bound to a proximal binding site). When the polynucleotide is DNA (e.g., genomic DNA), one or both strands can be cleaved at the nuclease cleavage site. Such cleavage by the nuclease enzyme initiates DNA repair mechanisms within the cell, which establishes an environment for homologous recombination to occur.
A site-directed nuclease can be a naturally-occurring site-directed nuclease. Exemplary naturally-occurring site-directed nucleases are known in the art (see for example, Makarova et al., 2017, Cell 168: 328-328. el, and Shmakov et al., 2017, Nat Rev Microbiol 15(3): 169- 182, both herein incorporated by reference). In some embodiments, a site-directed nuclease binds a DNA-targeting polynucleotide (e.g., a guide RNA) and is thereby directed to a specific sequence within a target DNA and cleaves the target DNA.
In some embodiments, the site-directed nuclease is modified from its natural sequence (e.g., via mutation or one or more amino acid residues) to change its function. For example, the site-directed nuclease may be modified to be enzymatically inactive. The term “enzymatically inactive” can refer to a site-directed nuclease that can bind to a nucleic acid sequence in a polynucleotide in a sequence-specific manner, but may not cleave a target polynucleotide. An enzymatically inactive site-directed polypeptide can comprise an enzymatically inactive domain (e.g., a nuclease domain). Enzymatically inactive can refer to no activity. Enzymatically inactive can refer to substantially no activity. Enzymatically inactive can refer to essentially no activity. Enzymatically inactive can refer to an activity no more than 1%, no more than 2%, no more than 3%, no more than 4%, no more than 5%, no more than 6%, no more than 7%, no more than 8%, no more than 9%, or no more than 10% activity compared to a wild-type exemplary activity.
In some embodiments, the site-directed nuclease comprises a CRISPR-associated (Cas) protein or a Cas nuclease that functions in a CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats)/Cas system. In bacteria, this system can provide adaptive immunity against foreign DNA (Barrangou, R , et al, “CRISPR provides acquired resistance against viruses in prokaryotes, “Science (2007) 315: 1709-1712; Makarova, K.S., et al, “Evolution and classification of the CRISPR-Cas systems,” Nat Rev Microbiol (2011) 9:467- 477; Garneau, J. E., et al, “The CRISPR/Cas bacterial immune system cleaves bacteriophage and plasmid DNA,” Nature (2010) 468:67-71; Sapranauskas, R., et al, “The Streptococcus thermophilus CRISPR/Cas system provides immunity in Escherichia coli,” Nucleic Acids Res (2011) 39: 9275-9282). In a wide variety of organisms including diverse mammals, animals, plants, microbes, and yeast, a CRISPR/Cas system (e.g., modified and/or unmodified) can be utilized as a genome engineering tool. A CRISPR/Cas system can comprise a guide nucleic acid such as a guide RNA (gRNA) complexed with a Cas protein for targeted regulation of gene expression and/or activity or nucleic acid editing. An RNA- guided Cas protein (e.g., a Cas nuclease such as a Cas9 nuclease) can specifically bind a target polynucleotide (e.g., DNA) in a sequence-dependent manner. The Cas protein, if possessing nuclease activity, can cleave the DNA (Gasiunas, G., et al, “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” Proc Natl Acad Sci USA (2012) 109: E2579-E2 86; Jinek, M., et al, “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012) 337:816-821; Sternberg, S. EL, et al, “DNA interrogation by the CRISPR RNA-guided endonuclease Cas9,” Nature (2014) 507:62; Deltcheva, E., et al, “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Nature (201 1) 471 :602- 607). DNA cleavage (e.g., double-strand breaks) can result in DNA break repair which allows for the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ), microhomology -mediated end joining (MMEJ), or homology -directed repair (HDR). In some embodiments, donor nucleic acids are used to promote HDR, as detailed below in the “Systems” section. CRISPR-Cas systems have been widely used for programmable genome editing in a variety of organisms and model systems (Cong, L., et al, “Multiplex genome engineering using CRISPR Cas systems,” Science (2013) 339:819-823; liang, W., et al, “RNA-guided editing of bacterial genomes using CRISPR-Cas systems,” Nat. Biotechnol. (2013) 31 : 233-239; Sander, J. D. & loung, J. K, “CRISPR-Cas systems for editing, regulating and targeting genomes,” Nature Biotechnol. (2014) 32:347-355).
In some embodiments, the site-directed nuclease described herein comprises a Cas protein that forms a complex with a guide nucleic acid, such as a guide RNA (described further below in the “Systems” section). In some embodiments, the site-directed nuclease comprises a Cas protein that forms a complex with a single guide nucleic acid, such as a single guide RNA (sgRNA). In some embodiments, the site-directed nuclease comprises a RNA-binding protein (RBP) optionally complexed with a guide nucleic acid, such as a guide RNA (e.g., sgRNA), which is able to form a complex with a Cas protein. In some instances, RNA-guided Cas proteins recognize DNA targets that are complementary to a portion of the gRNA known as a CRISPR RNA (crRNA) sequence. The target sequence is often referred to as a protospacer, and the part of the crRNA sequence that is complementary to the protospacer is often referred to as a spacer. In order to function (e.g., to cleave DNA), many Cas nucleases also require a specific protospacer adjacent motif (PAM), an approximately 2 to 6 base pair DNA sequence immediately following the protospacer sequence. A Cas protein used herein can be an active variant, inactive variant, or fragment of a wildtype or modified Cas protein. A Cas protein can comprise an amino acid change such as a deletion, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof relative to a wild-type version of the Cas protein. A Cas protein can be a polypeptide with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild-type exemplary Cas protein. A Cas protein can be a polypeptide with at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% sequence identity and/or sequence similarity to a wild-type exemplary Cas protein. Variants or fragments can comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild-type or modified Cas protein or a portion thereof. Variants or fragments can be targeted to a nucleic acid locus in complex with a guide nucleic acid while lacking nucleic acid cleavage activity.
A Cas protein can be modified to optimize regulation of gene expression. A Cas protein can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and/or enzymatic activity. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the function of the protein or to optimize (e.g., enhance or reduce) the activity of the Cas protein for regulating gene expression.
Also provided herein are variants of the polypeptides of this disclosure. Polypeptide variants retain their respective biological activity, unless explicitly noted otherwise. For example, variants of a site-directed nuclease polypeptide retain the biological function of the full length, native sequence site directed nuclease. In another example, variants of the nonspecific end-processing enzyme retain the biological function of the full length, native sequence nonspecific end-processing enzyme.
Modifications to any of the polypeptides or proteins provided herein are made by known methods. By way of example, modifications are made by site specific mutagenesis of nucleotides in a nucleic acid encoding the polypeptide, thereby producing a DNA encoding the modification, and thereafter expressing the DNA in recombinant cell culture to produce the encoded polypeptide. Techniques for making substitution mutations at predetermined sites in DNA having a known sequence are well known. For example, Ml 3 primer mutagenesis and PCR-based mutagenesis methods can be used to make one or more substitution mutations. Any of the nucleic acid sequences provided herein can be codon- optimized to alter, for example, maximize expression, in a host cell or organism.
The amino acids in the polypeptides described herein can be any of the 20 naturally occurring amino acids, D-stereoisomers of the naturally occurring amino acids, unnatural amino acids and chemically modified amino acids. Unnatural amino acids (that is, those that are not naturally found in proteins) are also known in the art, as set forth in, for example, Zhang et al. “Protein engineering with unnatural amino acids,” Curr. Opin. Struct. Biol. 23(4): 581-587 (2013); Xie et la. “Adding amino acids to the genetic repertoire,” 9(6): 548-54 (2005)); and all references cited therein. B and y amino acids are known in the art and are also contemplated herein as unnatural amino acids.
As used herein, a chemically modified amino acid refers to an amino acid whose side chain has been chemically modified. For example, a side chain can be modified to comprise a signaling moiety, such as a fluorophore or a radiolabel. A side chain can also be modified to comprise a new functional group, such as a thiol, carboxylic acid, or amino group. Post- translationally modified amino acids are also included in the definition of chemically modified amino acids.
Also contemplated are conservative amino acid substitutions. By way of example, conservative amino acid substitutions can be made in one or more of the amino acid residues, for example, in one or more lysine residues of any of the polypeptides provided herein. One of skill in the art would know that a conservative substitution is the replacement of one amino acid residue with another that is biologically and/or chemically similar. The following eight groups each contain amino acids that are conservative substitutions for one another:
1) Alanine (A), Glycine (G);
2) Aspartic acid (D), Glutamic acid (E);
3) Asparagine (N), Glutamine (Q);
4) Arginine (R), Lysine (K);
5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V);
6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W);
7) Serine (S), Threonine (T); and
8) Cysteine (C), Methionine (M). By way of example, when an arginine to serine is mentioned, also contemplated is a conservative substitution for the serine (e.g., threonine). Nonconservative substitutions, for example, substituting a lysine with an asparagine, are also contemplated.
Also provided is a DNA construct comprising a promoter operably linked to a recombinant nucleic acid encoding a fusion protein or domains thereof as described herein. A nucleic acid is “operably linked” when it is placed into a functional relationship with another nucleic acid sequence. Numerous promoters can be used in the constructs described herein. A promoter is a region or a sequence located upstream and/or downstream from the start of transcription that is involved in recognition and binding of RNA polymerase and other proteins to initiate transcription.
The term “promoter” as used herein refers to a nucleotide sequence, usually upstream (5’) to its coding sequence, which controls the expression of the coding sequence by providing the recognition for RNA polymerase and other factors required for proper transcription. “Promoter regulatory sequences” consist of proximal and more distal upstream elements. Promoter regulatory sequences influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, untranslated leader sequences, introns, and polyadenylation signal sequences. They include natural and synthetic sequences as well as sequences that may be a combination of synthetic and natural sequences. An “enhancer” is a DNA sequence that can stimulate promoter activity and may be an innate element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of a promoter. It is capable of operating in both orientations (e.g., forward or reverse) and is capable of functioning even when moved either upstream or downstream from the promoter. The meaning of the term “promoter” includes “promoter regulatory sequences.”
The choice of promoters to be included depends upon several factors, including, but not limited to, efficiency, selectability, inducibility, desired expression level, and cell- or tissue- preferential expression. It is a routine matter for one of skill in the art to modulate the expression of a sequence by appropriately selecting and positioning promoters and other regulatory regions relative to that sequence.
It has been shown that certain promoters are able to direct RNA synthesis at a higher rate than others. These are called "strong promoters". Certain other promoters have been shown to direct RNA synthesis at higher levels only in particular types of cells or tissues and are often referred to as "tissue specific promoters", or "tissue-preferred promoters", if the promoters direct RNA synthesis preferentially in certain tissues (RNA synthesis may occur in other tissues at reduced levels). Since patterns of expression of a chimeric gene (or genes) introduced into a plant are controlled using promoters, there is an ongoing interest in the isolation of novel promoters that are capable of controlling the expression of a chimeric gene (or genes) at certain levels in specific tissue types or at specific plant developmental stages.
Certain promoters are able to direct RNA synthesis at relatively similar levels across all tissues of a plant. These are called "constitutive promoters" or "tissue-independent" promoters. Constitutive promoters can be divided into strong, moderate, and weak categories according to their effectiveness to directing RNA synthesis. Since it is necessary in many cases to simultaneously express a chimeric gene (or genes) in different tissues of a plant to get the desired functions of the gene (or genes), constitutive promoters are especially useful in this regard. Though many constitutive promoters have been discovered from plants and plant viruses and characterized, there is still an ongoing interest in the isolation of more novel constitutive promoters, synthetic or native, which are capable of controlling the expression of a chimeric gene (or genes) at different levels and the expression of multiple genes in the same transgenic plant for gene stacking.
The recombinant nucleic acids provided herein can be included in expression cassettes for expression in a host cell or an organism of interest. The cassette will include 5' and 3' regulatory sequences operably linked to a recombinant nucleic acid provided herein that allows for expression of a fusion protein. The cassette may additionally contain at least one additional gene or genetic element to be cotransformed into the cell or organism. Where additional genes or elements are included, the components are operably linked. Alternatively, the additional gene(s) or element(s) can be provided on multiple expression cassettes. Such an expression cassette is provided with a plurality of restriction sites and/or recombination sites for insertion of the polynucleotides to be under the transcriptional regulation of the regulatory regions. The expression cassette may additionally contain a selectable marker gene. The expression cassette will include in the 5' to 3' direction of transcription: a transcriptional and translational initiation region (i.e., a promoter), a polynucleotide of the invention, and a transcriptional and translational termination region (i.e., termination region) functional in the cell or organism of interest. The promoters of the invention are capable of directing or driving expression of a coding sequence (i.e., a nucleic acid sequence that is transcribed into RNA such as mRNA, rRNA, tRNA, snRNA, ncRNA, IncRNA, sense RNA, or antisense RNA, regardless of whether the RNA is then translated to produce a protein) in a host cell. The regulatory regions (i.e., promoters, transcriptional regulatory regions, and translational termination regions) may be endogenous or heterologous to the host cell or to each other. As used herein, “heterologous” in reference to a sequence is a sequence that originates from a foreign species, or, if from the same species, is substantially modified from its native form in composition and/or genomic locus by deliberate human intervention.
Additional regulatory signals include, but are not limited to, transcriptional initiation start sites, operators, activators, enhancers, other regulatory elements, ribosomal binding sites, an initiation codon, termination signals, and the like. See Sambrook et al. (1992) Molecular Cloning: A Laboratory Manual, ed. Maniatis et al. (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.); Davis et al., eds. (1980) Advanced Bacterial Genetics (Cold Spring Harbor Laboratory Press), Cold Spring Harbor, N.Y., and the references cited therein.
The expression cassette can also comprise a selectable marker gene for the selection of transformed cells. Marker genes include genes conferring antibiotic resistance, such as those conferring hygromycin resistance, ampicillin resistance, gentamicin resistance, neomycin resistance, to name a few. Additional selectable markers are known and any can be used.
In preparing the expression cassette, the various DNA fragments may be manipulated, so as to provide for the DNA sequences in the proper orientation and, as appropriate, in the proper reading frame. Toward this end, adapters or linkers may be employed to join the DNA fragments or other manipulations may be involved to provide for convenient restriction sites, removal of superfluous DNA, removal of restriction sites, or the like. For this purpose, in vitro mutagenesis, primer repair, restriction, annealing, resubstitutions, e.g., transitions and transversions, may be involved.
Further provided is a vector comprising a recombinant nucleic acid or DNA construct set forth herein. The vector is contemplated to have the necessary functional elements that direct and regulate transcription of the inserted nucleic acid. These functional elements include, but are not limited to, a promoter, regions upstream or downstream of the promoter, such as enhancers that may regulate the transcriptional activity of the promoter, an origin of replication, appropriate restriction sites to facilitate cloning of inserts adjacent to the promoter, antibiotic resistance genes or other markers which can serve to select for cells containing the vector or the vector containing the insert, RNA splice junctions, a transcription termination region, or any other region which may serve to facilitate the expression of the inserted gene or hybrid gene. See generally, Sambrook et al. Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, 2012. The vector, for example, can be a plasmid. Transformation of a cell may be stable or transient. Thus, a transgenic cell, plant cell, plant and/or plant part of the invention can be stably transformed or transiently transformed. Transformation can refer to the transfer of a nucleic acid molecule into the genome of a host cell, resulting in genetically stable inheritance. In some embodiments, the introduction into a plant, plant part and/or plant cell is via bacterial-mediated transformation, particle bombardment transformation, calcium-phosphate-mediated transformation, cyclodextrin- mediated transformation, electroporation, liposome-mediated transformation, nanoparticle- mediated transformation, polymer-mediated transformation, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, sonication, infdtration, polyethylene glycol-mediated transformation, protoplast transformation, or any other electrical, chemical, physical and/or biological mechanism that results in the introduction of nucleic acid into the plant, plant part and/or cell thereof, or any combination thereof.
Procedures for transforming plants are well known and routine in the art and are described throughout the literature. Non-limiting examples of methods for transformation of plants include transformation via bacterial-mediated nucleic acid delivery (e.g. via bacteria from the genus Agrobacterium), viral-mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium-phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation,, sonication, infiltration, PEG-mediated nucleic acid uptake, as well as any other electrical, chemical, physical (mechanical) and/or biological mechanism that results in the introduction of nucleic acid into the plant cell, including any combination thereof. General guides to various plant transformation methods known in the art include Miki et al. (“Procedures for Introducing Foreign DNA into Plants” in Methods in Plant Molecular Biology and Biotechnology, Glick, B. R. and Thompson, J. E., Eds. (CRC Press, Inc., Boca Raton, 1993), pages 67-88) and Rakowoczy-Trojanowska (Cell Mol Biol Lett 7:849-858 (2002)).
Agrobacterium-mQ^iateA transformation is a commonly used method for transforming plants because of its high efficiency of transformation and because of its broad utility with many different species. Agrobacterium-mQ a\e . transformation typically involves transfer of the binary vector carrying the foreign DNA of interest to an appropriate Agrobacterium strain that may depend on the complement of vir genes carried by the host Agrobacterium strain either on a co-resident Ti plasmid or chromosomally (Uknes et al. 1993, Plant Cell 5:159- 169). The transfer of the recombinant binary vector to Agrobacterium can be accomplished by a tri-parental mating procedure using Escherichia coli carrying the recombinant binary vector, a helper E. coli strain that carries a plasmid that is able to mobilize the recombinant binary vector to the target Agrobacterium strain. Alternatively, the recombinant binary vector can be transferred to Agrobacterium by nucleic acid transformation (Hbfgen and Willmitzer 1988, Nucleic Acids Res 16:9877).
Transformation of a plant by recombinant Agrobacterium usually involves co-cultivation of the Agrobacterium with explants from the plant and follows methods well known in the art. Transformed tissue is typically regenerated on selection medium carrying an antibiotic or herbicide resistance marker between the binary plasmid T-DNA borders.
Another method for transforming plants, plant parts and plant cells involves propelling inert or biologically active particles at plant tissues and cells. See, e.g., US Patent Nos. 4,945,050; 5,036,006 and 5, 100,792. Generally, this method involves propelling inert or biologically active particles at the plant cells under conditions effective to penetrate the outer surface of the cell and afford incorporation within the interior thereof. When inert particles are utilized, the vector can be introduced into the cell by coating the particles with the vector containing the nucleic acid of interest. Alternatively, a cell or cells can be surrounded by the vector so that the vector is carried into the cell by the wake of the particle. Biologically active particles (e.g., dried yeast cells, dried bacteria or a bacteriophage, each containing one or more nucleic acids sought to be introduced) also can be propelled into plant tissue. As used herein, the phrase “biolistic transformation” refers to a method of introducing RNA or DNA into cells (e.g., plant cells) directly, in which RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cell (e.g., the plant cell) using high speed pressure to allow the RNA or DNA to penetrate the cell (e.g., to penetrate the plant cell wall).
The CRISPR/Cas system can also be used to edit the genome of a host cell or organism. As detailed above, the “CRISPR/Cas” system refers to a widespread class of bacterial systems for defense against foreign nucleic acid. Any of the CRISPR/Cas system components described herein may be used to introduce fusion proteins, recombinant nucleic acids, or systems into the genome of a host cell or organism. Methods for CRISPR/Cas system mediated genome editing are known in the art. It will be understood that use of a CRISPR/Cas system for introduction of fusion proteins, recombinant nucleic acids, or systems described herein into the genome of a host cell or organism is different from the particular methods and systems provided herein. In another aspect, provided herein are systems useful for editing one or more nucleic acids. The systems comprise one or more of the fusion proteins (or recombinant nucleic acids, constructs, vectors, or host cells) described above. In some embodiments, the systems further comprise one or more additional elements that are useful for editing one or more nucleic acids. For example, a system provided herein can further comprise a donor polynucleotide. As another example, a system comprising a fusion protein comprising a Cas nuclease may further comprise one or more guide nucleic acids and/or one or more donor polynucleotide sequences.
In some cases, the systems and methods described herein comprise at least one guide nucleic acid polynucleotide. In some cases, the systems and methods described herein comprise a plurality of guide nucleic acids In some embodiments, the polynucleotide can be deoxyribonucleic acid (DNA). In some cases, the DNA sequence can be single-stranded or doubled-stranded. In some embodiments, the at least one guide nucleic acid polynucleotide can be ribonucleic acid (guide RNA).
In some embodiments, the nuclease can be complexed with the at least one guide RNA polynucleotide. The at least one guide RNA polynucleotide can comprise a nucleic-acid targeting region that comprises a complementary sequence to a nucleic acid sequence on the targeted polynucleotide such as the targeted genomic loci or genes to confer sequence specificity of nuclease targeting. In some embodiments, the guide nucleic acid is a single guide nucleic acid comprising a crRNA. In some embodiments, the guide nucleic acid is a single guide nucleic acid comprising a crRNA but lacking a tracrRNA. A crRNA can comprise the nucleic acid-targeting segment (e.g., spacer region) of the guide nucleic acid and a stretch of nucleotides that can form one half of a double-stranded duplex of the Cas protein-binding segment of the guide nucleic acid.
In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer) is 20 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 19 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 18 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 17 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 16 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 21 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid is 22 nucleotides in length. The nucleotide sequence of the guide nucleic acid that is complementary to a nucleotide sequence (target sequence) of the target nucleic acid can have a length of, for example, at least about 12 nt, at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt or at least about 40 nt. The nucleotide sequence of the guide nucleic acid that is complementary to a nucleotide sequence (target sequence) of the target nucleic acid can have a length of from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 45 nt, from about 12 nt to about 40 nt, from about 12 nt to about 35 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about 19 nt, from about 19 nt to about 20 nt, from about 19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about 19 nt to about 60 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, or from about 20 nt to about 60 nt.
A protospacer sequence of a targeted polynucleotide can be identified by identifying a protospacer-adjacent motif (PAM) within a region of interest and selecting a region of a desired size upstream or downstream of the PAM as the protospacer. A corresponding spacer sequence can be designed by determining the complementary sequence of the protospacer region.
A spacer sequence can be identified using a computer program (e.g., machine readable code). The computer program can use variables such as predicted melting temperature, secondary structure formation, and predicted annealing temperature, sequence identity, genomic context, chromatin accessibility, % GC, frequency of genomic occurrence, methylation status, presence of SNPs, and the like.
The percent complementarity between the nucleic acid-targeting sequence (e.g., a spacer sequence of the at least one guide polynucleotide as disclosed herein) and the target nucleic acid (e.g., a protospacer sequence of the one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percent complementarity between the nucleic acid-targeting sequence and the target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 contiguous nucleotides. Guide nucleic acids of the systems of the disclosure can include modifications or sequences that provide for additional desirable features (e g., modified or regulated stability; subcellular targeting; tracking with a fluorescent label; a binding site for a protein or protein complex; and the like). Examples of such modifications include, for example, a 5' cap (a 7- m ethyl guanylate cap (m7G)); a 3' polyadenylated tail (a 3' poly(A) tail); a riboswitch sequence (e.g., to allow for regulated stability and/or regulated accessibility by proteins and/or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (a hairpin)); a modification or sequence that targets the RNA to a subcellular location (e g., nucleus, mitochondria, chloroplasts, and the like); a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, and so forth); a modification or sequence that provides a binding site for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyl transferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof.
A guide nucleic acid can comprise one or more modifications (e.g., a base modification, a backbone modification), to provide the nucleic acid with a new or enhanced feature (e.g., improved stability). A guide nucleic acid can comprise a nucleic acid affinity tag. A nucleoside can be a base-sugar combination. The base portion of the nucleotide can be a heterocyclic base. The two most common classes of such heterocyclic bases are the purines and the pyrimidines. Nucleotides can be nucleosides that further include a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides that include a pentofuranosyl sugar, the phosphate group can be linked to the 2', the 3', or the 5' hydroxyl moiety of the sugar. In forming guide nucleic acids, the phosphate groups can covalently link adjacent nucleosides to one another to form a linear polymeric compound. In turn, the respective ends of this linear polymeric compound can be further joined to form a circular compound; however, linear compounds can be suitable. In addition, linear compounds can have internal nucleotide base complementarity and can therefore fold in a manner as to produce a fully or partially double-stranded compound. Further, within guide nucleic acids, the phosphate groups can commonly be referred to as forming the internucleoside backbone of the guide nucleic acid. The linkage or backbone of the guide nucleic acid can be a 3' to 5' phosphodiester linkage.
In some embodiments, the at least one guide RNA polynucleotide of a system or method provided herein can bind to at least a portion of a genome (e.g., a plant genome) or a gene (e g., a plant gene). In some cases, the at least one guide RNA polynucleotide is capable of forming a complex with a site-directed nuclease to direct the site-directed nuclease to target the portion of a target nucleic acid (e.g., a site in a genome or a gene).
In some embodiments, the systems described herein comprise at least two (e.g., at least three, at least four, at least five, or at least six) different guide RNA polynucleotides that are able to form a complex with a site-directed nuclease.
DETAILED DESCRIPTION
The following description and examples recite various aspects and embodiments of the present compositions and methods. No particular embodiment is intended to define the scope of the compositions and methods. Rather, the embodiments merely provide non-limiting examples of various compositions and methods that are at least included within the scope of the disclosed compositions and methods. The description is to be read from the perspective of one of ordinary skill in the art; therefore, information well known to the skilled artisan is not necessarily included.
Described herein is a mutant Mb2Casl2a polypeptide comprising at least one amino acid substitution introduced into a wild-type Mb2Casl2a polypeptide sequence of SEQ ID NO: 1. In one embodiment, the at least one amino acid substitution occurs at a position selected from the group consisting of D172, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834. In another embodiment, the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q. In another embodiment, the mutant Mb2Casl2a polypeptide comprises a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66. In one embodiment, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV. In another embodiment, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CATA, CATC, CATG, GATA, GATC, GATG, TATA, TATC, and TATG. In another embodiment, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG. In another asp embodiment ect, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG. In another embodiment, the mutant Mc2Casl2a polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
Described herein is a method of editing a plant genome, comprising contacting said plant genome with the mutant Mb2Casl2a polypeptide recited above. In one embodiment, the method further comprises a guide RNA In another embodiment, the guide RNA is encoded by a sequence comprising SEQ ID NOs: 16-60.
Described herein is an edited plant obtained by the methods recited above.
Described herein is a construct or plasmid comprising a polynucleotide sequence encoding for the mutant Mb2Casl2a polypeptides recited above. One embodiment is a non-human cell comprising the construct or plasmid.
Described herein is an RNP complex comprising a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
EXAMPLES
Example 1. Altering amino acid residues can increase the PAM spectrum of Mb2Casl2a.
The nucleotide sequence for wildtype sequence of Mb2Casl2a (encoding the amino acid sequence of SEQ ID NO: 1) was subjected to various means of mutagenesis, including targeted mutagenesis, domain variants, and molecular breeding (also known as family DNA shuffling) in order to generate a library of diverse mutants. Mutants were first screened using E. coli plasmid clearance assay at the First Tier screening, comprising First and Second Stages. This assay uses an E. coli host with a resident plasmid containing the conditional lethal gene ccdB under control of the arabinose-inducible araC promoter. In the First Stage, nine versions of the resident plasmid were needed to encompass all variations of the VTV PAM (i.e., ATA, ATC, ATG, CTA, CTC, CTG, GTA, GTC, GTG) which flanks a common target sequence. The corresponding crRNA was expressed from the same plasmid. Each of the nine E. coli strains will contain one of the PAM-target resident plasmids. The library of Mb2Casl2a variants, under control of an IPTG-inducible T7 promoter, was transformed into the pool of the nine strains. In parallel, negative (no Mb2Casl2a) and positive (wild-type Mb2Casl2a) controls were also transformed. The transformation mixture was plated on media containing ampicillin, kanamycin, IPTG, and arabinose. Only cells expressing active Mb2Casl2a could cut the resident plasmid, which was then degraded, thus eliminating ccdB expression and allowing those cells to survive. Negative controls should result in no colonies. Likewise, if the wild-type Mb2Casl2a had an absolute requirement for TTV, TTTV PAM, none of the VTV variant PAMs should allow cutting and no colonies would survive. The Mb2Casl2a expression plasmids encoding variants that can utilize NVTV PAMs were recovered and moved to the Second Stage selection.
The Second Stage utilized an analogous screening pool, with three E. coll strains carrying resident plasmids comprising the degenerate PAM TTV (TTA, TTC, TTG). The result of this two-step process identified variants that react with non-canonical PAMs but also reacted with the wild-type PAM. The Mb2Casl2a expression plasmids were recovered from surviving colonies either as individuals or as a pool. Pooled plasmid DNA from surviving colonies was retransformed into the screening host as a further enrichment step for Mb2Casl2a variants with altered PAM preference in an iterative process. Plasmid DNA from individual colonies also was subjected to DNA sequencing to determine which amino acid positions have been changed.
In the Second Tier screen, a yeast system was used, in part because yeast cells can grow over a wider range of temperatures (e.g., 20-37°C). Yeast strains were generated in which the Ade2 gene is interrupted by a Ura3 gene adjacent to a Mb2Casl2a target site preceded by each of the 12 PAM sequences comprising NTV. The Ade2 gene is split in such a way as to leave 200 bp of Ade2 sequence homology on either side of the insertion. These homology arms constitute sites for recombination. In the absence of functional Ade2, yeast colonies exhibit a cell-autonomous red pigmentation. The screening host consisted of a pool of strains as in the bacterial assay described above. Mb2Casl2a nuclease expression was controlled by the GALI galactose-inducible promoter on a plasmid that also expresses the crRNA.
Individual Mb2Casl2a variants were transformed into the screening pool. Upon Mb2Casl2a cutting, single-strand annealing repair resulted in an intact Ade2 gene and functional enzyme. A negative control, or inactive Mb2Casl2a variant, would result in 100% red colonies. Wildtype Mb2Casl2a, if it had a strict TTV PAM requirement, would result in 25% white colonies. The ideal Mb2Casl2a variant with a preference for a PAM equal to NTV would result in 100% white colonies. Sectored colonies represent Mb2Casl2a variants and/or PAM sequences that result in incomplete cutting. Thus, the frequencies of red/white/sectored colonies provide an indication of the PAM preference and activity of each Mb2Casl2a variant. Sequence analysis of colonies with various red/white/sectored phenotypes indicated the PAM requirements of each Mb2Casl2a variant.
A Third Tier assay quantified PAM preference of Mb2Casl2a variants using a single crRNA. This strategy allowed for quantification of Mb2Casl2a efficiency at different PAM sites using the same target sequence to interrogate each variant PAM. This approach is based on a set of yeast strains in which variant PAM sequences comprising NTV (or NNTV) are introduced just upstream of the Kozak sequence preceding the Ade2 coding sequence. The Mb2Casl2a nuclease cuts at 18/23 bases downstream of the PAM and produces indels, which will put the Ade2 coding sequence out of frame. The red/white phenotypic readout and quantification would be similar to the previously described assay.
Table 1. Amino acid residue differences between variants and wildtype Mb2Casl2a.
NCXOI:D 1 2 3 4 5 6 7 8 9 10 11 12
WT (172R) (172Y> 46 (172R) (172Y) 51 18 21 28 (172Y) 44
172 D R Y D R Y D R R D Y D
563 N R R R R R R N N N N R
569 K V V V V V V K K T T V
570 E E E E E E E E E R R E
572 D D D D D D D D D R R D
573 N R R R R R R D D D D R
625 K R R R K K K K K K K K
629 A A A A A A A A S A A A
633 L L L L L L L L I L L L
636 Y Y Y Y Y Y Y Y F Y Y Y
642 L L L L L L L L I L L L
671 A A A A A A A A A A A A
672 G G G G G G G S G G G G
678 E E E E E E E D E E E E
695 L L L L L L L L I L L L
696 S S S S S S S S S S S S
705 Q Q Q Q Q Q Q Q Q Q Q Q
708 Q Q Q Q Q Q Q Q K Q Q Q
725 Q Q Q Q Q Q Q E E Q Q Q
753 F F F F F F F F W F F F
758 S E E E E E E S D E E E
759 E T T T C C C E E C C C
782 L L L L L L L L I L L L
783 D D D D Q Q Q D D Q Q D
788 T T T T F F F T T T T F
802 D L L L L L L D D D D D
824 F Y Y Y Y Y Y F F Y Y M 826 L L L L L L L L F L L L
831 T T T T T T T T T T T T
834 F Q Q Q F F F F F F F F
Example 2. In vitro assay for Mb2Casl2a screening.
To comprehensively screen the activity of Mb2Casl2a on different PAMs, the variants were assayed in vitro. Target DNA containing randomly combined four nucleotides (NNNN) followed by a validated gRNA sequence were synthesized. DNA sequences were then incubated with Mb2Casl2a variants. With carefully designed PCR and NGS sequencing, the cut efficiency of each NNNN (total 256 combinations) were determined which indicate the recognition of Mb2Casl2a to that sequence.
The DNA substrates used comprised, in the 5’ to 3’ direction, a 5’ barcode sequence, a four random nucleotide residues, a validated target sequence, and a 3’ barcode sequence. These DNA substrates were synthesized by Azenta Inc., and are represented by SEQ ID NO: 62 or SEQ ID NO: 63 (ZmGL2 and ZmmiR528, respectively).
RNP was assembled for each variant by incubating the following recipe for 30 minutes at 25°C:
The reaction buffer and DNA substrates were added to each assembled RNP and incubated for one hour at 37°C:
DNA was recovered from the reaction tubes and mixed into one tube for next-generation sequencing.
Example 3. Assaying the PAM spectrum in prokaryotic cells.
To determine the PAM recognition spectrum of engineered Mb2Casl2a variants with wildtype and dead Mb2Casl2a, a PAM determine assay (PAMDA) was performed. The PAMDA consists of 2 components. The first component is an ampicillin resistance plasmid library of gRNA target sequence contains 48 PAMs (NNTV) with same target sequence: GTGATAAGTGGAATGGCATGTGGG. The 2nd component is an E. coli strain harboring a low copy number pl5A origin chloramphenicol resistant plasmid which overexpresses one of engineered Mb2Casl2a variants, wildtype Mb2Casl2a or dead dMb2Casl2a (D864A, E958A) driven by T7 promoter and a gRNA expression cassette recognizes the target sequence in the PAM library. Transform the PAM plasmid library into competent E. coli cell prepared from the over expression Mb2Casl2a variant and gRNA expression cassette, further selected on Lb agar plate containing chloramphenicol and ampicillin antibiotics. Thus, if the gRNA recognizes the targeted sequence on the plasmid containing correct PAM, the plasmid will be cut by Casl2a and further be degraded. The survival plasmids contain PAM sequence that can’t be recognized by co-transformed Mb2Casl2a variant. Pool the survival colonies and NGS analysis the remaining PAM abundance. More abundant PAMs mean less cutting and less abundant PAMs means more cutting. The PAMs abundance of each Mb2Cas 12a variant is normalized to dMb2Casl2a. E. coli screening tests were performed as described in Example 1.
Table 2. Mb2Casl2a variants having expanded PAM spectrum in E. coli assay.
++ indicates good editing activity at that PAM; + indicated some editing activity observed at that PAM; - indicates no editing observed at that PAM.
Example 4. Assaying the PAM spectrum in eukaryotic cells.
Maize mesophyll protoplasts were prepared for transfection with vectors expressing Mb2Casl2a variants and gRNAs. See, e.g., M.R. Coy, et al., Protoplast Isolation and Transfection in Maize, in PROTOPLAST TECHNOLOGY METHODS AND PROTOCOLS, 91-104 (K. Wang & F. Zhang, eds., 2022) (detailing a standard protoplast transfection protocol).
Constructing the Mb2 Casl2a expression vector: Mb2 Casl2a which contained the specific position mutations were maize codon optimized and synthesized commercially (GenScript, Nanjing, China), and cloned under the sugarcane Ubiquitin-4 (SoUbi4) gene promoter to generate the plant expression vector.
Constructing the guide RNA expression vector: The CRISPR/Casl2a guide RNA transcript is expressed under the control of the OsU6 promoter which targets the corn native gene. It also included direct repeat of Mb2Casl2a crRNA AATTTCTACTGTTTGTAGAT (SEQ ID NO: 61) as the scaffold.
Corn protoplast transformation of expression vector: Plasmid DNA was purified by Plasmid DNA Purification using QIAGEN Plasmid Plus Midi Kits (Catalog No. 12945). 20 pg Casl2a expression plasmid and 7.5 pg gRNA expression plasmid (2 pmol in each) were mixed into 20 pl solution, and transfected by PEG method into 5 x 104 corn protoplast cell. Each treatment have three transformations as well as the repeat.
Genomic DNA extraction after cell harvesting: Cells were collected after 48 hours, centrifugated 5 min at 100g, and removed 250 pl supernatant. Added 100 pl Lysis buffer, and placed on a shaker for 5 min at room temperature (“RT”). Centrifuged for 15 min at 4000 rpm.
Transferred 100 pl supernatant to a new round bottom well, and added 5 pl Beads. Specimens were placed on shaker for 5 min at RT. Washed twice and air dried 5 min. Eluted DNA in 100 pl low Tris-EDTA (“TE”).
PCR amplification of target sites DNA fragment:
PCR, using the appropriate primers to amplify the DNA fragment that contained target site, was conducted following the recipes and PCR conditions below.
PCR recipe.
PCR thermocycler conditions.
PCR product purification protocol: Add Agencourt AMPure XP (54 pl beads to 30 pl PCR) Let the mixed samples incubate for 5 minutes at room temperature for maximum recovery. Place the reaction plate onto an Magnet Plate for 2 minutes to separate beads from the solution. Dispense 200 pl of 70% ethanol to each well of the reaction plate and incubate for 30 seconds at room temperature. Aspirate out the ethanol and discard. Remove the reaction plate from the magnet plate, and then add 35 pl warm water. Incubate for 2 minutes. Transfer the eluate into another clean plate.
T7EI detection of Editing: Anneal PCR product using below recipe and thermocycler conditions of 95C 5min; 95-85C -2C/sec; 85-25C -O.lC/sec; 4C hold.
5 lOx NEB buffer 2: 1.5 pl purified PCR product: 11 pl
Then Add 2.5 ul diluted lU/ul T7EI (10X dilution) to the above annealed PCR, 37 degree 60 min.
15 ul to do 2% Agrose gel electrophoresis or 2 ul to Agilent 5300 Fragment analyzer.
10
Table 3. PAM spectrum of Mb2Casl2a variants in corn protoplasts.
Cells with numerical entries show data normalized activity compared to LbCasl2a (TTTG).
Empty cells reflect no editing observed. “NT” indicates not tested. Example 5. List of enzymes used and their mutations, and gRNA sequences.
Table 4. Mutant Mb2Cas 12a variants.
Table 5. gRNA sequences. gRNA expression vector SEQ ID NO: PAM Target gene gRNA-A-1 SEQ ID NO: 16 TTTG Waxyl gRNA-A-6 SEQ ID NO: 17 GATA Bx9 gRNA-A-7 SEQ ID NO: 18 GATC Bx9 gRNA-A-8 SEQ ID NO: 19 GATG Bx9 gRNA- A- 12 SEQ ID NO: 20 GTTA Bx9 gRNA-A-13 SEQ ID NO: 21 GTTC Waxyl gRNA-A-14 SEQ ID NO: 22 GTTG Waxyl gRNA-A-15 SEQ ID NO: 23 TTTA MLO2c gRNA-A-19 SEQ ID NO: 24 CATA Waxyl gRNA-A-20 SEQ ID NO: 25 CATC MLO2c gRNA-A-21 SEQ ID NO: 26 CATG Bx9 gRNA-A-25 SEQ ID NO: 27 CTTA MLO2c gRNA-A-26 SEQ ID NO: 28 CTTC MLO2c gRNA-A-27 SEQ ID NO: 29 CTTG Bx9 gRNA-A-31 SEQ ID NO: 30 AATA MLO2c gRNA-A-32 SEQ ID NO: 31 AATC MLO2c gRNA-A-33 SEQ ID NO: 32 AATG MLO2c gRNA-A-36 SEQ ID NO: 33 ATTA Waxyl gRNA-A-37 SEQ ID NO: 34 ATTC MLO2c gRNA-A-38 SEQ ID NO: 35 ATTG MLO2c gRNA-A-42 SEQ ID NO: 36 TATA Waxyl gRNA-A-43 SEQ ID NO: 37 TATC MLO2c gRNA-A-44 SEQ ID NO: 38 TATG Waxyl gRNA-B-1 SEQ ID NO: 39 TTTG Glossy2 gRNA-B-6 SEQ ID NO: 40 GATA BINa gRNA-B-7 SEQ ID NO: 41 GATC Glossy2 gRNA-B-8 SEQ ID NO: 42 GATG MLO2b gRNA-B-12 SEQ ID NO: 43 GTTA MLO2b gRNA-B-13 SEQ ID NO: 44 GTTC MLO2b gRNA-B-14 SEQ ID NO: 45 GTTG Glossy2 gRNA-B-19 SEQ ID NO: 46 CATA DMR6 gRNA-B-20 SEQ ID NO: 47 CATC DMR6 gRNA-B-21 SEQ ID NO: 48 CATG DMR6 gRNA-B-25 SEQ ID NO: 49 CTTA DMR6 gRNA-B-26 SEQ ID NO: 50 CTTC Glossy2 gRNA-B-27 SEQ ID NO: 51 CTTG Glossy2 gRNA-B-31 SEQ ID NO: 52 AATA MLO2b gRNA-B-32 SEQ ID NO: 53 AATC MLO2b gRNA-B-33 SEQ ID NO: 54 AATG MLO2b gRNA-B-36 SEQ ID NO: 55 ATTA MLO2b gRNA-B-37 SEQ ID NO: 56 ATTC DMR6 gRNA-B-38 SEQ ID NO: 57 ATTG Glossy2 gRNA-B-42 SEQ ID NO: 58 TATA Glossy2 gRNA-B-43 SEQ ID NO: 59 TATC Glossy2 gRNA-B-44 SEQ ID NO: 60 TATG Glossy2

Claims

What is claimed is:
1. A mutant Mb2Casl2a polypeptide comprising at least one amino acid substitution introduced into a wild-type Mb2Casl2a polypeptide sequence of SEQ ID NO: 1
2. The mutant Mb2Casl2a polypeptide of claim 1, wherein the at least one amino acid substitution occurs at a position selected from the group consisting of DI 72, N563, K569, E570, D572, N573, K625, A629, L633, Y636, L642, A671, G672, E678, L695, Q705, Q708, Q725, F753, S758, E759, L782, D783, T788, D802, F824, L826, T831, and F834.
3. The mutant Mb2Casl2a polypeptide of claim 2, wherein the at least one amino acid substitution is selected from the group consisting of D172R, D172Y, N563R, K569V, K569T, E570R, D572R, N573R, N573D, K625R, A629S, L633I, Y636F, L642I, G672S, E678D, L695I, Q708K, Q725E, F753W, S758E, S758D, E759T, E795C, L782I, D783Q, T788F, D802L, F824Y, F824M, L826F, and F834Q
4. The mutant Mb2Casl2a polypeptide of claim 3, wherein the polypeptide comprises a sequence selected from the group comprising SEQ ID NOs: 2-14 and 64-66.
5. The mutant Mc2Casl2a polypeptide of claim 4, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of NATV, NCTV, NGTV, and NTTV.
6. The mutant Mc2Casl2a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of AATA, AATC, AATG, CAT A, CATC, CATG, GAT A, GATC, GATG, TATA, TATC, and TATG.
7. The mutant Mc2Casl2a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of ACTA, ACTC, ACTG, CCTA, CCTC, CCTG, GCTA, GCTC, GCTG, TCTA, TCTC, and TCTG.
8. The mutant Mc2Casl2a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of AGTA, AGTC, AGTG, CGTA, CGTC, CGTG, GGTA, GGTC, GGTG, TGTA, TGTC, and TGTG.
9. The mutant Mc2Casl2a polypeptide of claim 5, wherein the polypeptide has a PAM recognition sequence selected from the group consisting of ATTA, ATTC, ATTG, CTTA, CTTC, CTTG, GTTA, GTTC, GTTG, TTTA, TTTC, and TTTG.
10. A method of editing a plant genome, comprising contacting said plant genome with the mutant Mb2Cas 12a polypeptide of claims 1-5.
11. The method of claim 10, further comprising a guide RNA.
12. The method of claim 11, wherein the guide RNA is encoded by a sequence comprising SEQ ID NOs: 16-60.
13. An edited plant obtained by the method of claims 10-12.
14. A construct or plasmid comprising a polynucleotide sequence encoding for the mutant Mb2Casl2a polypeptide of claims 1-5.
15. A non-human cell comprising the construct or plasmid of claim 10.
16. An RNP complex comprising a sequence selected from the group comprising SEQ ID
NOs: 2-14 and 64-66.
EP24747710.2A 2023-01-27 2024-01-24 Mb2cas12a variants with flexible pam spectrum Pending EP4655395A2 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN2023073487 2023-01-27
PCT/US2024/012682 WO2024158857A2 (en) 2023-01-27 2024-01-24 Mb2cas12a variants with flexible pam spectrum

Publications (1)

Publication Number Publication Date
EP4655395A2 true EP4655395A2 (en) 2025-12-03

Family

ID=91971025

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24747710.2A Pending EP4655395A2 (en) 2023-01-27 2024-01-24 Mb2cas12a variants with flexible pam spectrum

Country Status (7)

Country Link
EP (1) EP4655395A2 (en)
JP (1) JP2026503701A (en)
KR (1) KR20250139286A (en)
CN (1) CN120677235A (en)
AR (1) AR131693A1 (en)
CL (1) CL2025002126A1 (en)
WO (1) WO2024158857A2 (en)

Family Cites Families (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
AU2017257274B2 (en) * 2016-04-19 2023-07-13 Massachusetts Institute Of Technology Novel CRISPR enzymes and systems
WO2021056302A1 (en) * 2019-09-26 2021-04-01 Syngenta Crop Protection Ag Methods and compositions for dna base editing

Also Published As

Publication number Publication date
WO2024158857A3 (en) 2024-10-03
WO2024158857A2 (en) 2024-08-02
CN120677235A (en) 2025-09-19
KR20250139286A (en) 2025-09-23
CL2025002126A1 (en) 2025-10-24
JP2026503701A (en) 2026-01-29
AR131693A1 (en) 2025-04-23

Similar Documents

Publication Publication Date Title
EP3110945B1 (en) Compositions and methods for site directed genomic modification
JP2020508046A (en) Genome editing system and method
KR102626503B1 (en) Target sequence-specific modification technology using nucleotide target recognition
JP2024166200A (en) Gene silencing by genome editing
US20250129367A1 (en) Methods and compositions for targeted genomic insertion
CN116286742B (en) CasD protein, CRISPR/CasD gene editing system and application thereof in plant gene editing
JP2025168346A (en) Methods and compositions for DNA base editing
EP4655396A1 (en) Mb2cas12a variants with enhanced efficiency
JP2022522823A (en) Suppression of target gene expression by genome editing of natural miRNA
EP4655395A2 (en) Mb2cas12a variants with flexible pam spectrum
WO2019213910A1 (en) Methods and compositions for targeted editing of polynucleotides
US20250043295A1 (en) Modified agrobacteria for editing plants
WO2024080067A1 (en) Genome editing method and composition for genome editing
HK40030403A (en) Target sequence specific alteration technology using nucleotide target recognition

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250827

AK Designated contracting states

Kind code of ref document: A2

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR