4239-111427-02 MODIFIED PIGGYBAT TRANSPOSITION SYSTEM AND USES THEREOF CROSS REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Application No.63/632,275, filed April 10, 2024, which is herein incorporated by reference in its entirety. FIELD This disclosure concerns a piggyBat transposition system with enhanced transposition activity resulting from modification of the terminal inverted repeats of piggyBat transposon DNA and modification of the piggyBat transposase protein. Use of the improved piggyBat transposition system, such as for editing genomic DNA, is also described. ACKNOWLEDGMENT OF GOVERNMENT SUPPORT This invention was made with government support under project number DK036153-16 awarded by the National Institutes of Health. The government has certain rights in the invention. INCORPORATION OF ELECTRONIC SEQUENCE LISTING The electronic sequence listing, submitted herewith as an XML file named 4239-111427- 02.xml (14,868 bytes), created on April 4, 2025, is herein incorporated by reference in its entirety. BACKGROUND DNA transposons are mobile genetic elements that can move from location to location in the genome of their host. There are two major classes of transposable elements. Class 1 elements are retrotransposons that use an RNA intermediate that is converted to DNA in order to be inserted into the host genome. Class 2 elements are DNA transposons that use only DNA intermediates to integrate their DNA directly. The vast majority of eukaryotic DNA transposons are “cut and paste" type. Transposon superfamilies can be further defined and categorized based on the shared genetic makeup and structural organization of the transposons and their encoded transposases (Feschotte and Pritham, Annu. Rev. Genet.41, 331–368, 2007). In many prokaryotes, transposable elements play dynamic ecological roles due to their ability to carry exogenous genes, notably those responsible for antibiotic resistance. However, in higher organisms, the potential genotoxic effects of unregulated transposition have resulted in severe restriction of their mobility, either through mutation or by host-encoded systems that silence transposon activity (Ozata et al., Nature Rev. Genet.20, 89-108, 2019). For example, while about 45% of the human genome originated from transposable elements, only a small subclass of retrotransposons remains active. Among eukaryotic transposons, the eponymous member of its superfamily, piggyBac from the moth Tricoplusia ni, has been extensively studied (Fraser et al., J.
4239-111427-02 Virol.47, 287–300,1983) and, along with Sleeping Beauty from the Tc1/Mariner superfamily (Ivics et al., Cell 91, 501-510,1997), has been used for genome manipulation of human cells (Wilson et al., Mol. Ther.15, 139–145, 2007). Key to the usefulness of piggyBac and Sleeping Beauty as tools in genome engineering and therapeutic applications has been the development of hyperactive versions, selected either by concerted changes of conserved amino acids or random mutagenetic screens (Mátés et al., Nat. Genet.41, 753–761, 2009; Yusa et al., Proc. Natl. Acad. Sci.108, 1531–1536, 2011; Tipanee et al., Hum. Gene Ther.28, 1087–1104, 2017; Ptáčková et al., Cytotherapy 20, 507–520, 2018). Another member of the piggyBac superfamily, piggyBat (Ray et al., Genome Res.18, 717- 728, 2008), is active in bat, yeast, and mammalian cells in culture (Mitra et al., Proc. Natl. Acad. Sci. USA 110, 234–239, 2013), yet its overall transposition activity is much lower than that of T. ni piggyBac, particularly when compared to piggyBac hyperactive variants (Yusa et al., Proc. Natl. Acad. Sci.108, 1531–1536, 2011). DNA transposon ends typically have sequences organized as terminal inverted repeats (TIRs) that are specifically recognized by the transposase in order to bring them together and carry out DNA cleavage and joining reactions necessary to accomplish transposition. The sequences at the piggyBat termini do not contain the same pattern of short repeated subterminal motifs as observed for piggyBac (Bouallègue et al., Genome Biol. Evol.9, 323-339, 2017; Morellet et al., Nucleic Acids Res.46, 2660-2677, 2018) despite the 28.7% amino acid identity between the two transposases. As transposon end recognition is fundamental for organizing the nucleoprotein assemblies (called “transpososomes”) that are needed to carry out transposition in an orderly fashion, it is unknown how the piggyBat transposase recognizes its ends. DNA sequence changes accumulated through evolution that have obscured binding motifs could be responsible for the limited activity of the wild type piggyBat transposon. In view of the limited activity of the piggyBat transposon in mammalian cells, a need exists for the development of improved piggyBat transpositions systems with enhanced transposition activity. SUMMARY Most of the currently existing transposition systems capable of functioning in human cells exhibit low transposition activity. To address the need for transposition systems with enhanced activity, the present disclosure describes modified piggyBat transposon DNA sequences and modified piggyBat transposase proteins that exhibit significantly increased (up to 100-fold) transposition activity in human cell cultures. Provided herein is a modified piggyBat transposon that includes a left end (LE) terminal inverted repeat (TIR) and a right end (RE) TIR, both of which are truncated relative to the wild-type piggyBat LE and RE (and relative to LE153 and RE208, shorter versions of these sequences that retain transposition activity). The LE TIR includes a 3’ deletion that removes the third repeat sequence and third palindromic sequence (G3LE+P3LE), starting at nucleotide 99 of the LE sequence
4239-111427-02 set forth as SEQ ID NO: 4. It is demonstrated herein that deletion of the G3LE+P3LE motifs of the LE led to a substantial increase in transposition activity (FIG.5C). The RE TIR includes a 3’ deletion that retains all four repeat sequences (G1RE, G2RE, G3RE and G4RE) and up to nucleotide 100 to nucleotide 160 of SEQ ID NO: 5. The combination of the truncated LE TIR and truncated RE TIR led to an additional ~2-fold increase in transposition activity, compared to activity using only the truncated LE TIR. Thus, in some aspects, the modified piggyBat transposon includes an LE TIR sequence including a 3’ deletion following any one of nucleotide 88 to nucleotide 98 of SEQ ID NO: 4 (i.e., following nucleotide 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, or 98 of SEQ ID NO: 4); and/or the RE TIR includes a 3' deletion following any one of nucleotide 100 to nucleotide 160 of SEQ ID NO: 5 (i.e., following nucleotide 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159 or 160 of SEQ ID NO: 5). In some examples, the nucleotide sequence of the LE TIR is at least 80% identical to LE88 (SEQ ID NO: 1) and/or the nucleotide sequence of the RE TIR is at least 80% identical to RE100 (SEQ ID NO: 2). In specific non-limiting examples, the nucleotide sequence of the LE TIR includes or consists of LE88 (SEQ ID NO: 1) and/or the nucleotide sequence of the RE TIR includes or consists of RE100 (SEQ ID NO: 2). In some aspects, the modified piggyBat transposon further includes a transgene. Vectors that include a modified piggyBat transposon disclosed herein are also provided. Further provided herein is a modified piggyBat transposase protein with enhanced transposition activity, particularly when combined with the truncated LE TIR and RE TIR. In some aspects, the modified piggyBat transposase includes a duplication of the C-terminal cysteine-rich domain (CRD), such as a duplication of residues 573-665 of SEQ ID NO: 3. In some aspects, the modified piggyBat transposase includes or further includes serine to alanine substitutions at residues 8, 24, 32 and 37 corresponding to wild-type piggyBat transposase of SEQ ID NO: 11. In some examples, the amino acid sequence of the piggyBat transposase is at least 80% identical to SEQ ID NO: 3 (pBat-4StoA-2xCRD). In specific non-limiting examples, the amino acid sequence of the piggyBat transposase includes or consists of SEQ ID NO: 3. Also provided herein are nucleic acid molecules (such as DNA or RNA) that encode a modified piggyBat transposase disclosed herein. In some aspects, the nucleic acid molecule is codon- optimized for expression in mammalian cells (such as human cells). Vectors that include a disclosed nucleic acid molecule are also provided. Further provided are kits, such as kits for editing genomic DNA. In some aspects, the kit includes a modified piggyBat transposon disclosed herein, or a vector containing a modified piggyBat transposon, and further includes a nucleic acid molecule (such as a DNA or RNA) encoding a modified piggyBat transposase disclosed herein, or a vector containing such a nucleic acid molecule.
4239-111427-02 Also provided are methods of editing genomic DNA in an isolated cell. In some aspects, the method includes contacting the isolated cell with a modified piggyBat transposon disclosed herein, or a vector containing such a modified piggyBat transposon, and further contacting the isolated cell with a nucleic acid molecule (such as a DNA or RNA) encoding a modified piggyBat transposase disclosed herein, or a vector containing such a nucleic acid molecule. Further provided are methods of editing genomic DNA in a subject, for example to treat a disease in a subject. In some aspects, the method includes administering to the subject a modified piggyBat transposon disclosed herein, or a vector containing such a modified piggyBat transposon, and further administering to the subject a nucleic acid molecule (such as a DNA or RNA) encoding a modified piggyBat transposase disclosed herein, or a vector containing such a nucleic acid molecule. The foregoing and other features of this disclosure will become more apparent from the following detailed description of several aspects which proceeds with reference to the accompanying figures. BRIEF DESCRIPTION OF THE DRAWINGS FIG.1: Schematic of cut-and-paste transposition by piggyBac-like elements. FIGS.2A-2C: Comparison of piggyBat and piggyBac transposons. (FIG.2A) Schematic of the piggyBat transposon. The intact transposon (Ray et al., Genome Res.18, 717-728, 2008) and an active form with shorter ends comprised of 153 bp of Left End (LE) sequence and 208 bp of Right End (RE) sequence (Mitra et al., Proc. Natl. Acad. Sci. USA 110, 234–239, 2013) have been described. Box: The DNA sequence corresponding to LE153 is shown on top (SEQ ID NO: 4), with repeats at nucleotides 11-16 (G1LE), nucleotides 55-60 (G2LE), nucleotides 99-104 (G3LE), nucleotides 19, 20, 23, 25, 26 and 28-31 (P1LE), nucleotides 58, 59, 62, 64, 65 and 67-70 (P2LE) and nucleotides 107, 108, 111, 113, 114 and 116-119 (P3LE). The P1LE, P2LE and P3LE repeats may correspond to imperfect palindromes, as indicated by the arrows. The DNA sequence corresponding to RE208 is shown on the bottom (SEQ ID NO: 5), with possible repeats indicated (nucleotides 12-15, G1RE; nucleotides 22-27, G1RE; nucleotides 44-48, G3RE; and nucleotides 63-68, G4RE). Nucleotides 1-4 of LE153 and RE108 are identical nucleotides at the transposon tips. (FIG.2B) Schematic of the piggyBac transposon from Trichoplusia ni (GenBank J04364.2; Cary et al., Virology 172(1):156-169, 1989). Box: Minimal transposon ends required for activity (LE35/RE63, SEQ ID NO: 6/SEQ ID NO: 7) are indicated. (FIG.2C) Schematic representations of (left) the cryo-EM structure of the pB transposase bound to two LE35 TIRs (from PDB 6X68); (middle) proposed model for the pB synaptic complex; and (right) the redesigned hyperactive piggyBac system. FIGS.3A-3E: Cryo-EM structure of pBat transposase bound to LE44. (FIG.3A). The two monomers of the pBat dimer are shown. The motifs on LE44 are indicated as in FIG.2A. bp 31-35 are also labelled. (FIG.3B) Comparison of the structures of the C-terminal domains (CRD) of pBat
4239-111427-02 and pB, and the C1 domain. (FIG.3C) Close-up of the recognition of the GCGGGA motif. (FIG.3D) Close-up of CRD recognition. (FIG.3E) Close-up of minor groove interactions involving pBat residues 495-506. FIGS.4A-4C: piggyBat transposition in cultured human cells. (FIG.4A) Schematic of the plasmid-to-chromosome transposition assay in HEK293F cells. (FIG.4B) Transposition activity (as indicated by colony count) for active LE and RE of the piggyBat transposon with and without protein in HEK293T cells. No activity was detected with two LEs (LE/LE) or two REs (RE/RE). (FIG.4C) The effect on transposition activity in HEK293T cells of truncating the piggyBat transposon ends, using 50X initial cell dilution (left) or 400X initial cell dilution (right). LE/RE indicates LE153- RE208. FIGS.5A-5E: Mutation of predicted CKII phosphorylation sites on the N-terminal domain leads to hyperactivity in cells. (FIG.5A) pBat (top; SEQ ID NO: 8) and pB (bottom; SEQ ID NO: 9) N-terminal amino acid sequences highlighting the SDX(D/E) CKII phosphorylation motifs (underlined). Numbering above the sequences corresponds to the amino acid number of pBat. (FIG. 5B) Transposition activity for WT and phosphorylation mutant transposases. (FIG.5C) Transposition activity for WT and 4StoA point mutant transposase on truncated ends. (FIG.5D) Comparison of transposition activity of WT and pBat 4StoA and pBat 4StoA on truncated ends. (FIG.5E) Schematic of pBat-4StoA-2xCRD (left) and transposition activities of pBat-4StoA and pBat-4StoA-2xCRD on truncated piggyBat transposon ends (right). FIG.6: Flowchart of the single particle cryo-electron microscopy structure determination of the pBat-LE44 pre-synaptic complex. FIGS.7A-7D: (FIG.7A) Fourier shell correlation (SC) curves. (FIG.7B) Final sharpened reconstruction using RELION (left) and DeepEMhancer (right). (FIG.7C) Resolution distribution of the final reconstruction from RELION. (FIG.7D) Distribution of particle views used for the final reconstruction. FIGS.8A-8B: (FIG.8A) Comparison of topologies of the pBat (left) and pB (right) CRDs. (FIG.8B) Topological relatives of pBat identified by Dali. FIG.9: Comparison of the pBat-LE44 and pB-LE35 donor DNAs. FIG.10: Schematic of the observed interactions between pBat protein and the bound LE44 in the pre-synaptic complex. FIG.11: Structure determination statistics. FIGS.12A-12B: (FIG.12A) Comparison of design of (top) modified pBat transposase protein (Version 1; SEQ ID NO: 10) and (bottom) modified pBat transposase protein (pBat-4StoA- 2xCRD; SEQ ID NO: 3). (FIG.12B) Comparison of transposition activity (measured as colony counts) of WT pBat protein, modified pBat transposase protein (Version 1), and modified pBat transposase protein pBat-4StoA-2xCRD (SEQ ID NO: 3).
4239-111427-02 SEQUENCES The nucleic acid and amino acid sequences listed in the accompanying sequence listing are shown using standard letter abbreviations for nucleotide bases, and single letter code for amino acids, as defined in 37 C.F.R.1.822. Only one strand of each nucleic acid sequence is shown, but the complementary strand is understood as included by any reference to the displayed strand. In the accompanying sequence listing: SEQ ID NO: 1 is the nucleotide sequence of piggyBat LE88. SEQ ID NO: 2 is the nucleotide sequence of piggyBat RE100. SEQ ID NO: 3 is the amino acid sequence of modified transposase pBat-4StoA-2xCRD. SEQ ID NO: 4 is the nucleotide sequence of piggyBat LE153. SEQ ID NO: 5 is the nucleotide sequence of piggyBat RE208. SEQ ID NO: 6 is the nucleotide sequence of piggyBac LE35. SEQ ID NO: 7 is the nucleotide sequence of piggyBac RE63. SEQ ID NO: 8 is the N-terminal amino acid sequence of the piggyBat transposase that contains four CKII phosphorylation motifs. SEQ ID NO: 9 is the N-terminal amino acid sequence of the piggyBac transposase that contains four CKII phosphorylation motifs. SEQ ID NO: 10 is the amino acid sequence of a modified pBat transposase protein (version 1). SEQ ID NO: 11 is the amino acid sequence of the wild-type piggyBat transposase. SEQ ID NO: 12 is a nucleotide sequence of a portion of the piggyBat LE88. SEQ ID NO: 13 is the amino acid sequence of piggyBat transposase CRD. DETAILED DESCRIPTION I. Introduction Members of the piggyBac superfamily of DNA transposons are widely distributed and have been identified in a variety of host genomes ranging from insects to mammals. In the human genome, five piggyBac elements have been retained as domesticated elements but are no longer mobile. The present disclosure describes the transposition properties of piggyBat, a member of the piggyBac transposon superfamily found in Myotis lucifugus. piggyBat is the only active mammalian DNA transposon identified to date. Despite its low activity in human cells, piggyBat has been used as a genomic tool for studies in human cells (Sutrave et al., Mol Ther Methods Clin Dev 25:250-263, 2022). The data disclosed herein demonstrate that transposition activity of piggyBat in vivo is severely restricted by an internal transposase binding site on its Left End (LE), a previously unobserved mechanism for down-regulation of transposon activity. The cryo-electron microscopy structure of the piggyBat transposase pre-synaptic complex (FIG.3A) showed an unexpected DNA configuration and recognition using C-terminal domains topologically different from piggyBac. The
4239-111427-02 structural data disclosed herein allowed for the rational re-engineering of piggyBat, including modification of both the LE and the Right End (RE), elimination of predicted N-terminal phosphorylation sites of the piggyBat transposase, and duplication of the transposase C-terminal site- specific DNA binding domain. These modifications increased transposition activity by approximately two orders of magnitude relative to wild type piggyBat. The transposon and transposase modifications disclosed herein lead to a transposition activity that is comparable to the most highly active piggyBac transposition system (Luo et al., Nucleic Acids Res 50:13128-13142, 2022). II. Abbreviations BSA bovine serum albumin CRD cysteine-rice domain cryo-EM cryo-electron microscopy CTD C-terminal domain FAM carboxyfluorescein DTT dithiothreitol EMSA electrophoretic mobility shift assay LE Left End MBP maltose binding protein pB piggyBac transposase pBat piggyBat transposase RE Right End SDS-PAGE sodium dodecyl sulfate polyacrylamide gel electrophoresis TCEP tris(2-carboxyethyl)phosphine TEV tobacco etch virus TIR terminal inverted repeat TSD target site duplication WT wild-type III. Summary of Terms Unless otherwise noted, technical terms are used according to conventional usage. Definitions of many common terms in molecular biology may be found in Krebs et al. (eds.), Lewin’s genes XII, published by Jones & Bartlett Learning, 2017. As used herein, the singular forms “a,” “an,” and “the,” refer to both the singular as well as plural, unless the context clearly indicates otherwise. For example, the term “a cell” includes singular or plural cell s and can be considered equivalent to the phrase “at least one cell.” As used herein, the term “comprises” means “includes.” It is further to be understood that any and all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for nucleic acids or polypeptides are approximate, and are provided
4239-111427-02 for descriptive purposes, unless otherwise indicated. Although many methods and materials similar or equivalent to those described herein can be used, particular suitable methods and materials are described herein. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. To facilitate review of the various aspects, the following explanations of terms are provided: Administration: The introduction of a composition (such as a nucleic acid molecule or vector) into a subject by a chosen route. Administration can be local or systemic. For example, if the chosen route is intravenous, the composition is administered by introducing the composition into a vein of the subject. Exemplary routes of administration include, but are not limited to, oral, injection (such as subcutaneous, intramuscular, intradermal, intraperitoneal, intratumoral, and intravenous), infusion, sublingual, rectal, transdermal (for example, topical), intranasal, vaginal, and inhalation routes. Codon-optimized: A nucleic acid molecule encoding a protein can be codon-optimized for expression of the protein in a particular organism by including the codon most likely to encode a particular amino acid at each position of the sequence. Codon usage bias is the difference in the frequency of occurrence of synonymous codons (encoding the same amino acid) in coding DNA. A codon is a series of three nucleotides (a triplet) that encodes a specific amino acid residue in a polypeptide chain or for the termination of translation. There are 20 different naturally-occurring amino acids, but 64 different codons (61 codons encoding for amino acids plus 3 stop codons). Thus, there is degeneracy because one amino acid can be encoded by more than one codon. A nucleic acid sequence can be optimized for expression in a particular organism (such as a human) by evaluating the codon usage bias in that organism and selecting the codon most likely to encode a particular amino acid. Multivariate statistical methods, such as correspondence analysis and principal component analysis, are widely used to analyze variations in codon usage. Computer programs are available to implement the statistical analyses related to codon usage, such as Codon W, GCUA, and INCA. Conservative amino acid substitutions: Amino acid substitutions that do not substantially affect a function of a protein (e.g., a transposase), such as the native activity of the protein, and/or do not substantially change the chemical and stereochemical properties of the amino acids. In some aspects, a conservative amino acid substitution in a piggyBat transposase is one that does not reduce activity of the transposase by more than 10% (such as by more than 5% or more than 1%) compared to the activity of the corresponding transposase lacking the conservative amino acid substitution. In some aspects, the piggyBat transposase can include up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or up to 10 conservative substitutions compared to the wild-type transposase and retain transposition activity. Typically, individual substitutions, deletions or additions which alter, add or delete a single amino acid or a small percentage of amino acids (for instance less than 5%, in some aspects less than
4239-111427-02 1%) in an encoded sequence are conservative variations where the alterations result in the substitution of an amino acid with a chemically similar amino acid. The following six groups are examples of amino acids that are considered to be conservative substitutions for one another: 1) Alanine (A), Serine (S), Threonine (T); 2) Aspartic acid (D), Glutamic acid (E); 3) Asparagine (N), Glutamine (Q); 4) Arginine (R), Lysine (K); 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V); and 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W). Contacting: Placement in direct physical association; includes both in solid and liquid form, which can take place either in vivo or in vitro. Contacting includes contact between one molecule and another molecule. Contacting can also include contact between a cell and a nucleic acid molecule or protein. Degenerate variant: A polynucleotide encoding a protein (for example, a piggyBat transposase) that includes a sequence that is degenerate as a result of the genetic code. There are twenty natural amino acids, most of which are specified by more than one codon. Therefore, all degenerate nucleotide sequences are included as long as the amino acid sequence of the transposase encoded by the nucleotide sequence is unchanged. Expression vector: A vector containing a recombinant polynucleotide that includes expression control sequences operatively linked to a nucleotide sequence to be expressed. An expression vector includes sufficient cis-acting elements for expression; other elements for expression can be supplied by the host cell or in an in vitro expression system. Expression vectors include, for example, cosmids, plasmids (for example, naked or contained in liposomes) and viruses (for example, lentiviruses, retroviruses, adenoviruses, and adeno-associated viruses) that incorporate the recombinant polynucleotide. Induced pluripotent stem cell (iPSC): A type of pluripotent stem cell that is generated from a somatic cell. iPSCs are genetically reprogrammed to an embryonic stem cell-like status through forced expression of specific genes and factors. In some aspects herein, the iPSC is a human iPSC. Isolated: An “isolated” or “purified” biological component (such as a nucleic acid, peptide, protein, protein complex, cell or virus) has been substantially separated, produced apart from, or purified away from other biological components in the cell or organism in which the component occurs, that is, other chromosomal and extrachromosomal DNA and RNA, proteins, and cells. In some aspects herein, an “isolated cell” refers to a cell that is not part of an organism (e.g., an isolated primary cell or a cultured cell). The term “isolated” or “purified” does not require absolute purity; rather, it is intended as a relative term. Thus, for example, an isolated biological component (such as a cell) is one in which the biological component is more enriched than the biological component is in its natural environment. Preferably, a preparation is purified such that the biological component
4239-111427-02 represents at least 50%, such as at least 70%, at least 90%, at least 95%, or greater, of the total biological component content of the preparation. Modified (protein or nucleic acid): A protein or nucleic acid having at least one change in its amino acid or nucleotide sequence, respectively, relative to a wild-type sequence. Nucleotide and amino acid sequence modifications include, for example, substitutions, insertions and deletions, or combinations thereof. For proteins, insertions include amino and/or carboxyl terminal fusions as well as intrasequence insertions of single or multiple amino acid residues. Deletions are characterized by the removal of one or more amino acid residues from the protein sequence at any position. Substitutional modifications are those in which at least one residue has been removed and a different residue inserted in its place. For nucleic acid molecules, insertions include 5' or 3' extensions, as well as intrasequence insertions of single or multiple nucleotides. Similarly, deletions in a nucleotide sequence include deletion of one or more nucleotides at the 5' end or 3' end, or a deletion of an internal sequence. Nucleotide substitutions refer to replacement of a nucleotide at a particular position with a different nucleotide. In some aspects herein, the modification (such as a substitution, insertion or deletion) of a transposase protein or a transposon DNA sequence results in a change in its activity, such as an enhancement in transposition activity. Operably linked: A first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter, such as the CMV promoter, is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence. Generally, operably linked DNA sequences are contiguous and, where necessary to join two protein- coding regions, in the same reading frame. Pharmaceutically acceptable carriers: The pharmaceutically acceptable carriers useful in this invention are conventional. Remington: The Science and Practice of Pharmacy, 22nd ed., London, UK: Pharmaceutical Press, 2013, describes compositions and formulations suitable for pharmaceutical delivery of a therapeutic agent, such as a transposon or transposase disclosed herein. In general, the nature of the carrier will depend on the particular mode of administration being employed. For instance, parenteral formulations usually comprise injectable fluids that include pharmaceutically and physiologically acceptable fluids such as water, physiological saline, balanced salt solutions, aqueous dextrose, glycerol or the like as a vehicle. In addition to biologically-neutral carriers, pharmaceutical compositions to be administered can contain minor amounts of non-toxic auxiliary substances, such as wetting or emulsifying agents, preservatives, and pH buffering agents and the like, for example sodium acetate or sorbitan monolaurate. PiggyBat transposase: A transposase found in Myotis lucifugus, a species of little brown bat. The amino acid sequence of the wild-type piggyBat transposase protein is set forth herein as SEQ ID NO: 11. In some aspects herein, the piggyBat transposase is a modified transposase having a duplication of the C-terminal cysteine-rich domain (CRD) and four serine to alanine substitutions in
4239-111427-02 the N-terminal portion of the protein (referred to herein as “pBat-4StoA-2xCRD” which has an amino acid sequence set forth as SEQ ID NO: 3). It is demonstrated herein that the pBat-4StoA-2xCRD transposase has significantly improve transposition activity, particularly when used in combination with the modified transposons described herein. PiggyBat transposon: A transposon found in Myotis lucifugus, a species of little brown bat. The piggyBat transposon includes a left end (LE) TIR and a right end (RE) TIR. The wild-type piggyBat LE and RE are 586 bp and 324 bp, respectively. PiggyBat transposons having a 153 bp LE (LE153) and a 208 bp RE (RE208) have previously been shown to be active, albeit with low activity. In some aspects herein, the LE and RE are further truncated at their 3' ends to yield LE88 (88 bp in length) and RE100 (100 bp in length). Modified piggyBat transposons containing LE88 and RE100 are shown herein to possess significantly improved transposition activity, particularly when used in combination with a modified piggyBat transposase protein. Preventing, treating or ameliorating a disease: “Preventing” a disease refers to inhibiting the full development of a disease. “Treating” refers to a therapeutic intervention that ameliorates a sign or symptom of a disease or pathological condition (such as a genetic disease) after it has begun to develop. “Ameliorating” refers to the reduction in the number or severity of signs or symptoms of a disease. Sequence identity: The identity between two or more nucleic acid sequences, or two or more amino acid sequences, is expressed in terms of the identity between the sequences. Sequence identity can be measured in terms of percentage identity; the higher the percentage, the more identical the sequences. Homologs and variants are typically characterized by possession of at least about 75%, for example at least about 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98% or 99% sequence identity counted over the full-length alignment with the amino acid sequence of interest. Methods of alignment of sequences for comparison are well known in the art. Various programs and alignment algorithms are described in: Smith and Waterman, Adv. Appl. Math. 2(4):482-489, 1981; Needleman and Wunsch, J. Mol. Biol.48(3):443-453, 1970; Pearson and Lipman, Proc. Natl. Acad. Sci. U.S.A.85(8):2444-2448, 1988; Higgins and Sharp, Gene, 73(1):237- 244, 1988; Higgins and Sharp, Bioinformatics, 5(2):151-3, 1989; Corpet, Nucleic Acids Res. 16(22):10881-10890, 1988; Huang et al. Bioinformatics, 8(2):155-165, 1992; and Pearson, Methods Mol. Biol.24:307-331, 1994. Altschul et al., J. Mol. Biol.215(3):403-410, 1990, presents a detailed consideration of sequence alignment methods and homology calculations. The NCBI Basic Local Alignment Search Tool (BLAST) (Altschul et al., J. Mol. Biol.215(3):403-410, 1990) is available from several sources, including the National Center for Biological Information and on the Internet, for use in connection with the sequence analysis programs blastp, blastn, blastx, tblastn, and tblastx. Blastn is used to compare nucleic acid sequences, while blastp is used to compare amino acid sequences. Additional information can be found at the NCBI web site.
4239-111427-02 Once aligned, the number of matches is determined by counting the number of positions where an identical nucleotide or amino acid residue is present in both sequences. The percent sequence identity is determined by dividing the number of matches either by the length of the sequence set forth in the identified sequence, or by an articulated length (such as 100 consecutive nucleotides or amino acid residues from a sequence set forth in an identified sequence), followed by multiplying the resulting value by 100. Stem cell: An undifferentiated cell of a multicellular organism that is capable of giving rise to many different types of cells in the organism. Stem cells include both embryonic stem cells (also known as pluripotent stem cells) and adult stem cells (also known as somatic stem cells). Embryonic stem cells have the ability to differentiate into any cell type. Adult stem cells are found in a tissue or organ and can differentiate into any specialized cell type of that tissue or organ. Subject: Living multicellular vertebrate organisms, a category that includes human and non- human mammals. In some aspects, the subject is a human, such as a human with a genetic disease or disorder. In other examples, the subject is a non-human primate or mouse. T cell: A type of cell of the immune system that plays a role in protection against infectious agents and cancer. Also known as a T lymphocyte. Therapeutically effective amount: A quantity of a specific substance, such as a disclosed transposon or transposase (such as a nucleic acid molecule or vector encoding a transposase), or a composition thereof, sufficient to achieve a desired effect in a subject being treated. A “therapeutically effective amount” can be the amount necessary to inhibit, prevent or suppress one or more symptoms of a genetic disease, or to correct a genetic disease such that the subject no longer has symptoms of the disease. For example, administration of a therapeutically effective amount of composition disclosed herein can inhibit, prevent or suppress one or more signs or symptoms of a genetic disease, such as by at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or even at least 100% (elimination or prevention of detectable signs or symptoms), as compared to in the absence of treatment. Transduced and Transformed: A virus or vector “transduces” a cell when it transfers nucleic acid into the cell. A cell is “transformed” or “transfected” by a nucleic acid transduced into the cell when the nucleic acid molecule becomes stably replicated by the cell, either by incorporation of the nucleic acid into the cellular genome, or by episomal replication. Numerous methods of transfection can be used, such as: chemical methods (e.g., calcium-phosphate transfection), physical methods (e.g., electroporation, microinjection, particle bombardment), fusion (e.g., liposomes), receptor-mediated endocytosis (e.g., DNA-protein complexes, viral envelope/capsid-DNA complexes) and by biological infection by viruses such as recombinant viruses (Wolff, J. A., ed, Gene Therapeutics, Birkhauser, Boston, USA, 1994). Transgene: A gene that is not native to the cell or organism to which it is transferred. Transgenes also include artificial genes, such as genes with particular modifications. In some aspects
4239-111427-02 herein, the transgene is a therapeutic gene, such as a gene capable of treating a particular disease or disorder. In other aspects, the transgene is a reporter gene, such as a gene encoding a fluorescent protein, an enzyme or other molecule that can be detected when it is expressed in the cell or organism. Transposase: An enzyme involved in the movement of specific genetic elements from one location in a genome to another location in the genome. The term “transposase” also refers to a protein that catalyzes the excision of a transposon from a donor polynucleotide (such as a vector) and integration of the transposon into a target site of a nucleic acid (such as genomic DNA). Transposon: A mobile genetic element that can move from one location in a genome to another location in the genome. The term “transposon” also includes polynucleotides that are capable of being excised from a donor polynucleotide (such as a vector) and integrating into a target site of a nucleic acid (such as genomic DNA). Vector: A nucleic acid molecule allowing insertion of foreign nucleic acid. A vector can include nucleic acid sequences that permit it to replicate in a host cell, such as an origin of replication. A vector can also include one or more selectable marker genes and other genetic elements. An expression vector is a vector that contains the necessary regulatory sequences to allow transcription and translation of inserted gene or genes. IV. Modified piggyBat Transposition System Currently, piggyBat (from M. lucifugus, a species of little brown bat) is the only known active DNA transposon found in mammalian genomes; however, its activity is very low when compared to that of its close relative, piggyBac from T. ni. Low transposition activity is not unusual as transposons typically evolve in the absence of positive selection, and thus tend to accumulate debilitating mutations over time and eventually lose activity completely. For instance, Sleeping Beauty, a widely used DNA transposon for a variety of applications (Amberger et al., BioEssays 42, 2000136, 2020), originates from fossilized fragments of inactive elements in the genomes of fish species that were fused and mutated synthetically to restore transposition activity. Cells can also actively control and inhibit the activity of mobile genetic elements in their genomes, for example by using the well-known piRNA pathway of transposon silencing (Loubalova et al., Mobile DNA 14, 10, 2023). As robust transposition activity is often a desirable property for genetic engineering applications, substantial efforts have been put into increasing activity, and it has been demonstrated that the activity of some transposons can be dramatically increased by the introduction of only a small number of amino changes in the transposase and/or limited nucleotide changes in the transposon DNA (Mátés et al., Nat. Genet.41, 753–761, 2009; Goryshin and Reznikoff, J. Biol. Chem.273, 7367-7374, 1998; Yusa et al., Proc. Natl. Acad. Sci.108, 1531–1536, 2011; Lazarow et al., Genetics 191, 747-756, 2012). It is demonstrated herein that the very low activity of piggyBat when compared to PiggyBac is due to subterminal inhibitory sequences, and that transposition activity can be dramatically improved by their removal. The cryo-electron microscopy structure of the piggyBat transposase pre-
4239-111427-02 synaptic complex showed an unexpected DNA configuration and recognition using C-terminal domains topologically different from piggyBac. Structural data based rational re-engineering of piggyBat through the removal of putative phosphorylation sites in the transposase N-terminus and a duplication of the C-terminal domain resulted in a highly active transposition system. Provided herein is a modified piggyBat transposon that includes a left end (LE) terminal inverted repeat (TIR) and a right end (RE) TIR, both of which are truncated relative to the wild-type piggyBat LE and RE (and relative to LE153 and RE208). In some aspects, the LE TIR includes a 3' deletion that removes the third repeat sequence and third palindromic sequence (G3LE+P3LE), starting at nucleotide 99 of the LE sequence set forth as SEQ ID NO: 4. It is demonstrated herein that deletion of the G3LE and P3LE motifs of the LE led to a substantial increase in transposition activity (FIG.5C). In some aspects, the RE TIR includes a 3' deletion that retains all four repeat sequences (G1RE, G2RE, G3RE and G4RE) and up to nucleotide 100 to nucleotide 160 of SEQ ID NO: 5. The combination of the truncated LE TIR and truncated RE TIR led to an additional ~2-fold increase in transposition activity (relative to activity when only the LE TIR was truncated; FIG.5D). Thus, in some aspects, the modified piggyBat transposon includes an LE TIR sequence including a 3' deletion following any one of nucleotide 88 to nucleotide 98 of SEQ ID NO: 4 (i.e., following nucleotide 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, or 98 of SEQ ID NO: 4). In some aspects, the LE TIR is 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, or 98 nucleotides in length. In some aspects, the modified piggyBat transposon has an RE TIR that includes a 3' deletion following any one of nucleotide 100 to nucleotide 160 of SEQ ID NO: 5 (i.e., following nucleotide 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159 or 160 of SEQ ID NO: 5). In some aspects, the LE TIR is 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159 or 160 nucleotides in length. In some aspects, the modified piggyBat transposon includes combinations of these LE TIR and RE TIR sequences. In some examples, the LE TIR consists of nucleotides 1 to 88, nucleotides 1 to 89, nucleotides 1 to 90, nucleotides 1 to 91, nucleotides 1 to 92, nucleotides 1 to 93, nucleotides 1 to 94, nucleotides 1 to 95, nucleotides 1 to 96, nucleotides 1 to 97, or nucleotides 1 to 98 of SEQ ID NO: 4, or a nucleotide sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to nucleotides 1 to 88, nucleotides 1 to 89, nucleotides 1 to 90, nucleotides 1 to 91, nucleotides 1 to 92, nucleotides 1 to 93, nucleotides 1 to 94, nucleotides 1 to 95, nucleotides 1 to 96, nucleotides 1 to 97, or nucleotides 1 to 98 of SEQ ID NO: 4. In some examples, the RE TIR consists of nucleotides 1 to 100, nucleotides 1 to 101, nucleotides 1 to102, nucleotides 1 to 103, nucleotides 1 to 104, nucleotides to 105, nucleotides 1 to
4239-111427-02 106, nucleotides 1 to 107, nucleotides 1 to 108, nucleotides 1 to 109, nucleotides 1 to 110, nucleotides 1 to 111, nucleotides 1 to 112, nucleotides 1 to 113, nucleotides 1 to 114, nucleotides 1 to 115, nucleotides 1 to 116, nucleotides 1 to 117, nucleotides 1 to 118, nucleotides 1 to 119, nucleotides 1 to 120, nucleotides 1 to 121, nucleotides 1 to 122, nucleotides 1 to 123, nucleotides 1 to 124, nucleotides 1 to 125, nucleotides 1 to 126, nucleotides 1 to 127, nucleotides 1 to 128, nucleotides 1 to 129, nucleotides 1 to 130, nucleotides 1 to 131, nucleotides 1 to 132, nucleotides 1 to 133, nucleotides 1 to 134, nucleotides 1 to 135, nucleotides 1 to 136, nucleotides 1 to 137, nucleotides 1 to 138, nucleotides 1 to 139, nucleotides 1 to 140, nucleotides 1 to 141, nucleotides 1 to 142, nucleotides 1 to 143, nucleotides 1 to 144, nucleotides 1 to 145, nucleotides 1 to 146, nucleotides 1 to 147, nucleotides 1 to 148, nucleotides 1 to 149, nucleotides 1 to 150, nucleotides 1 to 151, nucleotides 1 to 152, nucleotides 1 to 153, nucleotides 1 to 154, nucleotides 1 to 155, nucleotides 1 to 156, nucleotides 1 to 157, nucleotides 1 to 158, nucleotides 1 to 159, or nucleotides 1-160 of SEQ ID NO: 5, or a nucleotide sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to nucleotides 1 to 100, nucleotides 1 to 101, nucleotides 1 to 102, nucleotides 1 to 103, nucleotides 1 to 104, nucleotides to 105, nucleotides 1 to 106, nucleotides 1 to 107, nucleotides 1 to 108, nucleotides 1 to 109, nucleotides 1 to 110, nucleotides 1 to 111, nucleotides 1 to 112, nucleotides 1 to 113, nucleotides 1 to 114, nucleotides 1 to 115, nucleotides 1 to 116, nucleotides 1 to 117, nucleotides 1 to 118, nucleotides 1 to 119, nucleotides 1 to 120, nucleotides 1 to 121, nucleotides 1 to 122, nucleotides 1 to 123, nucleotides 1 to 124, nucleotides 1 to 125, nucleotides 1 to 126, nucleotides 1 to 127, nucleotides 1 to 128, nucleotides 1 to 129, nucleotides 1 to 130, nucleotides 1 to 131, nucleotides 1 to 132, nucleotides 1 to 133, nucleotides 1 to 134, nucleotides 1 to 135, nucleotides 1 to 136, nucleotides 1 to 137, nucleotides 1 to 138, nucleotides 1 to 139, nucleotides 1 to 140, nucleotides 1 to 141, nucleotides 1 to 142, nucleotides 1 to 143, nucleotides 1 to 144, nucleotides 1 to 145, nucleotides 1 to 146, nucleotides 1 to 147, nucleotides 1 to 148, nucleotides 1 to 149, nucleotides 1 to 150, nucleotides 1 to 151, nucleotides 1 to 152, nucleotides 1 to 153, nucleotides 1 to 154, nucleotides 1 to 155, nucleotides 1 to 156, nucleotides 1 to 157, nucleotides 1 to 158, nucleotides 1 to 159, or nucleotides 1-160 of SEQ ID NO: 5. In some examples, the nucleotide sequence of the LE TIR is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to LE88 (SEQ ID NO: 1) and/or the nucleotide sequence of the RE TIR is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to RE100 (SEQ ID NO: 2). In specific non-limiting examples, the nucleotide sequence of the LE TIR includes or consists of LE88 (SEQ ID NO: 1) and/or the nucleotide sequence of the RE TIR includes or consists of RE100 (SEQ ID NO: 2). In some aspects, the modified piggyBat transposon further includes a transgene. In some examples, the transgene is a therapeutic gene or a reporter gene. The therapeutic transgene can be any gene that provides a therapeutic effect in the treatment of a disease, disorder or condition (e.g., the
4239-111427-02 gene can replace a mutated or non-functional gene in the subject). In specific examples, the therapeutic gene is capable of treating congenital deafness, such as the otoferlin (OTOF) gene (NCBI Gene ID 9831), Duchenne muscular dystrophy, such as the dystrophin gene (NCBI Gene ID 1756), or familial hypercholesterolemia, such as the low-density lipoprotein receptor (LDLR) gene (NCBI Gene ID 3949) or the apolipoprotein B (APOB) gene (NCBI Gene ID 338). In other examples, the therapeutic gene is capable of treating cystic fibrosis, hemophilia, immune deficiencies, Huntington disease, α-anti-trypsin deficiency, colon cancer, melanoma, kidney cancer, lymphoma, acute myeloid leukemia (AML), acute lymphoid leukemia (ALL), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), gastrointestinal tumors, lung cancer, gliomas, thyroid cancer, prostate tumors, hepatomas, virus-induced tumors, e.g., papillomavirus-induced carcinomas (such as cervical carcinoma), adenocarcinomas, herpesvirus-induced tumors (e.g., Burkitt's lymphoma, and Epstein- Barr virus-induced B cell lymphoma), hepatitis B virus-induced tumors, HTLV-1 and HTLV-2 induced lymphoma, lung cancer, pharyngeal cancer, anal carcinoma, glioblastoma, lymphoma, rectum carcinoma, astrocytoma, brain tumors, stomach cancer, retinoblastoma, medulloblastoma, vaginal cancer, pancreatic cancer, testis cancer, bladder cancer, meningioma, Schneeberger's disease, bronchial carcinoma, pituitary cancer, mycosis fungoides, gullet cancer, breast cancer, neurinoma, Burkitt's lymphoma, laryngeal cancer, thymoma, corpus carcinoma, bone cancer, non-Hodgkin lymphoma, urethra cancer, CUP-syndrome, oligodendroglioma, vulva cancer, intestinal cancer, esophagus carcinoma, small intestine tumors, craniopharyngioma, ovarian cancer, liver cancer, leukemia, or cancers of the skin or the eye. The reporter gene can be a gene encoding any molecule that can be detected when it is expressed in the cell, such as a gene encoding a fluorescent protein or an enzyme. Vectors that include a modified piggyBat transposon disclosed herein are also provided. In some examples, the vector is an expression vector, such as a plasmid vector. In other examples, the vector is a viral vector, such as an adenovirus vector, an adeno-associated virus (AAV) vector, or a lentivirus vector. Compositions that include a modified piggyBat transposon disclosed herein are also provided. Such compositions can also include a pharmaceutically acceptable carrier, such as a saline or water. Also provided herein are modified piggyBat transposase proteins with enhanced transposition activity. In some aspects, the modified piggyBat transposase includes a duplication of the C-terminal cysteine-rich domain (CRD), such as a duplication of at least 50, at least 60, at least 70, at least 80, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, or at least 100 amino acids of the C-terminus of the piggyBat transposase relative to the wild- type piggyBat transposase protein of SEQ ID NO: 11. In some examples, the amino acid sequence of the CRD includes DMDIVPDLQPVPSTSGMRAKPPTSDPPCRLSMDMRKHTLQAIVGSGKKKNILRRCRVCSVHK LRSETRYMCKFCNIPLHKGACFEKYHTLKNY (SEQ ID NO: 13, which corresponds to residues
4239-111427-02 573-665 of SEQ ID NO: 3), or an amino acid sequence at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to residues 573-665 of SEQ ID NO: 13. In some aspects, the modified piggyBat transposase includes or further includes one, two, three or four serine to alanine substitutions at residues selected from residues 8, 24, 32 and 37 corresponding to wild-type piggyBat transposase of SEQ ID NO: 11. In some aspects, the modified piggyBat transposase includes all four serine to alanine substitutions at residues 8, 24, 32 and 37. In some aspects, the amino acid sequence of the piggyBat transposase is at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to identical to SEQ ID NO: 3 (pBat-4StoA-2xCRD). In some examples, the amino acid sequence of the piggyBat transposase includes or consists of SEQ ID NO: 3. Further provided herein are nucleic acid molecules (such as DNA or RNA molecules) that encode a modified piggyBat transposase described herein. In some aspects, the nucleic acid molecule is codon-optimized for expression in mammalian cells, such as human cells. In some aspects, the nucleic acid molecule is operably linked to a promoter. The promoter can be, for example, a constitutive promoter, a tissue-specific promoter, or an inducible promoter. Vectors containing a transposase-encoding nucleic acid molecule are also provided. In some aspects, the vector is an expression vector, such as a plasmid vector. In other aspects, the vector is a viral vector, such as an adenovirus vector, an AAV vector, or a lentivirus vector. Compositions that include a modified piggyBat transposase (or nucleic acid molecule encoding such) disclosed herein are also provided. Such compositions can also include a pharmaceutically acceptable carrier, such as a saline or water. V. Kits and Methods for Editing Genomic DNA Also provided herein are kits, such as kits for editing genomic DNA in a cell. In some aspects, the kit includes a modified piggyBat transposon or a vector containing a modified piggyBat transposon as disclosed herein. In some examples, the vector is a plasmid vector. In some aspects, the kit further includes a nucleic acid molecule (such as a DNA or RNA molecule) or vector encoding a modified piggyBat transposase disclosed herein. In some examples, the vector is an expression vector. In other examples, the vector is a viral vector, such as an adenovirus vector, and AAV vector, or a lentivirus vector. In some examples, the kit further includes a transfection reagent, buffer, cell culture media, instructional materials, or any combination thereof. Further provided are methods of editing genomic DNA in an isolated cell. In some aspects, the method includes contacting the isolated cell with a modified piggyBat transposon or a vector containing a modified piggyBat transposon as disclosed herein, and a nucleic acid molecule (such as a DNA or RNA molecule) or vector encoding a modified piggyBat transposon disclosed herein. In some examples, “contacting” the isolated cell with a nucleic acid molecule or vector includes
4239-111427-02 transfecting the nucleic acid molecule or vector into the cells, such as by using any one of a number of commercially available transfection reagents. Transfection of the transposon and nucleic acid molecule/vector encoding the transposase can occur simultaneously or sequentially. In some examples, the transposon (or vector containing the transposon) is transfected first, followed by the nucleic acid molecule/vector encoding the transposase. In other examples, the nucleic acid molecule/vector encoding the transposase is transfected first, followed by transfection of the transposon. In some examples, one or both of the transposon and transposase nucleic acid/vector is transfected more than once, such as twice or three times. In alternative aspects, “contacting” the isolated cell with a nucleic acid molecule or vector includes electroporation or microinjection. In some aspects of the method for editing genomic DNA, the cell is a mammalian cell, such as a human cell. In other examples, the cell is a non-human primate cell, a mouse cell, or a rat cell. In some examples, the cell is a lymphocyte, such as a T cell. In other examples, the cell is a stem cell or an induced pluripotent stem cell (iPSC). In some examples, the cell is a cancer cell, a fibroblast cell, a hematopoietic cell, a primary cell or a cultured cell line. In some aspects, the piggyBat transposon further includes a transgene, such as a therapeutic transgene. In some examples, the transgene is a gene that provides a therapeutic effect in the treatment of a disease, disorder or condition (e.g., the gene can replace a mutated or non-functional gene in the subject). In specific examples, the therapeutic gene is capable of treating congenital deafness, such as the otoferlin (OTOF) gene (NCBI Gene ID 9831), Duchenne muscular dystrophy, such as the dystrophin gene (NCBI Gene ID 1756), or familial hypercholesterolemia, such as the low- density lipoprotein receptor (LDLR) gene (NCBI Gene ID 3949) or the apolipoprotein B (APOB) gene (NCBI Gene ID 338). In other examples, the therapeutic gene is capable of treating cystic fibrosis, hemophilia, immune deficiencies, Huntington disease, α-anti-trypsin deficiency, colon cancer, melanoma, kidney cancer, lymphoma, acute myeloid leukemia (AML), acute lymphoid leukemia (ALL), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), gastrointestinal tumors, lung cancer, gliomas, thyroid cancer, prostate tumors, hepatomas, virus- induced tumors, e.g., papillomavirus-induced carcinomas (such as cervical carcinoma), adenocarcinomas, herpesvirus-induced tumors (e.g., Burkitt's lymphoma, and Epstein-Barr virus- induced B cell lymphoma), hepatitis B virus-induced tumors, HTLV-1 and HTLV-2 induced lymphoma, lung cancer, pharyngeal cancer, anal carcinoma, glioblastoma, lymphoma, rectum carcinoma, astrocytoma, brain tumors, stomach cancer, retinoblastoma, medulloblastoma, vaginal cancer, pancreatic cancer, testis cancer, bladder cancer, meningioma, Schneeberger's disease, bronchial carcinoma, pituitary cancer, mycosis fungoides, gullet cancer, breast cancer, neurinoma, Burkitt's lymphoma, laryngeal cancer, thymoma, corpus carcinoma, bone cancer, non-Hodgkin lymphoma, urethra cancer, CUP-syndrome, oligodendroglioma, vulva cancer, intestinal cancer, esophagus carcinoma, small intestine tumors, craniopharyngioma, ovarian cancer, liver cancer, leukemia, or cancers of the skin or the eye.
4239-111427-02 Also provided are methods of editing genomic DNA in a subject. In some aspects, the method includes administering to the subject a modified piggyBat transposon, or a vector that includes a modified piggyBat transposon as disclosed herein, and a nucleic acid molecule (such as a DNA or RNA molecule) or vector encoding a modified piggyBat transposon disclosed herein. Administration of the transposon and nucleic acid molecule/vector encoding the transposase can occur simultaneously or sequentially. In some examples, the transposon (or vector containing the transposon) is administered to the subject first, followed by administration of the nucleic acid molecule/vector encoding the transposase. In other examples, the nucleic acid molecule/vector encoding the transposase is administered to the subject first, followed by administration of the transposon. In some examples, one or both of the transposon and transposase nucleic acid/vector is administered more than once, such as twice or three times. In some aspects, the subject is a human subject, such as a subject having a genetic disease, disorder or condition. In other aspects, the subject is a non-human primate, a mouse or a rat. In other aspects, the subject is a veterinary subject, such as a cat, dog, cow, horse or a goat. In some aspects of the method, the piggyBat transposon further includes a transgene, such as a therapeutic transgene. In some examples, the transgene is a gene that provides a therapeutic effect in the treatment of a disease, disorder or condition (e.g., the gene can replace a mutated or non-functional gene in the subject). In specific examples, the therapeutic gene is capable of treating congenital deafness, such as the otoferlin (OTOF) gene (NCBI Gene ID 9831), Duchenne muscular dystrophy, such as the dystrophin gene (NCBI Gene ID 1756), or familial hypercholesterolemia, such as the low-density lipoprotein receptor (LDLR) gene (NCBI Gene ID 3949) or the apolipoprotein B (APOB) gene (NCBI Gene ID 338). In other examples, the therapeutic gene is capable of treating cystic fibrosis, hemophilia, immune deficiencies, Huntington disease, α-anti-trypsin deficiency, colon cancer, melanoma, kidney cancer, lymphoma, acute myeloid leukemia (AML), acute lymphoid leukemia (ALL), chronic myeloid leukemia (CML), chronic lymphocytic leukemia (CLL), gastrointestinal tumors, lung cancer, gliomas, thyroid cancer, prostate tumors, hepatomas, virus-induced tumors, e.g., papillomavirus- induced carcinomas (such as cervical carcinoma), adenocarcinomas, herpesvirus-induced tumors (e.g., Burkitt's lymphoma, and Epstein-Barr virus-induced B cell lymphoma), hepatitis B virus-induced tumors, HTLV-1 and HTLV-2 induced lymphoma, lung cancer, pharyngeal cancer, anal carcinoma, glioblastoma, lymphoma, rectum carcinoma, astrocytoma, brain tumors, stomach cancer, retinoblastoma, medulloblastoma, vaginal cancer, pancreatic cancer, testis cancer, bladder cancer, meningioma, Schneeberger's disease, bronchial carcinoma, pituitary cancer, mycosis fungoides, gullet cancer, breast cancer, neurinoma, Burkitt's lymphoma, laryngeal cancer, thymoma, corpus carcinoma, bone cancer, non-Hodgkin lymphoma, urethra cancer, CUP-syndrome, oligodendroglioma, vulva cancer, intestinal cancer, esophagus carcinoma, small intestine tumors, craniopharyngioma, ovarian cancer, liver cancer, leukemia, or cancers of the skin or the eye. Thus, the disclosure provides methods of treating one or more of these genetic diseases. In one aspect, the disclosed methods
4239-111427-02 reduce one or more symptoms of a genetic disease or disorder in the treated subject, for example a reduction of at least 10%, at least 20%, at least 50%, at least 70%, at least 90%, or at least 95% (as compared to no administration of the therapeutic gene). In some aspects of the method for editing genomic DNA in a subject, the modified piggyBat transposon (or vector including same) and/or the modified piggyBat-encoding nucleic acid molecule or vector are administered as part of a pharmaceutical composition that includes a pharmaceutically- acceptable carrier, adjuvant or vehicle. In some examples, pharmaceutically-acceptable carrier, adjuvant, or vehicle is a non-toxic carrier, adjuvant or vehicle that does not impair the pharmacological activity of the component(s) with which it is formulated. Pharmaceutically acceptable carriers, adjuvants or vehicles that may be used in the compositions disclosed herein include, but are not limited to, ion exchangers, alumina, aluminum stearate, lecithin, serum proteins, such as human serum albumin, buffer substances such as phosphates, glycine, sorbic acid, potassium sorbate, partial glyceride mixtures of saturated vegetable fatty acids, water, salts or electrolytes, such as protamine sulfate, disodium hydrogen phosphate, potassium hydrogen phosphate, sodium chloride, zinc salts, colloidal silica, magnesium trisilicate, polyvinyl pyrrolidone, cellulose-based substances, polyethylene glycol, sodium carboxymethylcellulose, polyacrylates, waxes, polyethylene- polyoxypropylene-block polymers, and polyethylene glycol. The composition can be administered by any suitable route, such as orally, parenterally (e.g., by subcutaneous, intravenous, intramuscular, intra-articular, intra-synovial, intrasternal, intrathecal, intrahepatic, intralesional or intracranial injection, or via infusion techniques), via inhalation, topically, rectally, nasally, buccally, vaginally or by using an implanted reservoir. The pharmaceutical compositions can be formulated according to known techniques using suitable dispersing or wetting agents and suspending agents. A sterile injectable preparation may also be a sterile injectable solution or suspension in a non-toxic parenterally-acceptable diluent or solvent, for example as a solution in 1,3-butanediol. Among the acceptable vehicles and solvents that may be employed include but are not limited to water, Ringer's solution and isotonic sodium chloride solution. In addition, sterile, fixed oils are conventionally employed as a solvent or suspending medium. In some aspects, the pharmaceutical composition is orally administered in any orally acceptable dosage form including, but not limited to, capsules, tablets, aqueous suspensions or solutions. In the case of tablets for oral use, carriers commonly used include lactose and corn starch. Lubricating agents, such as magnesium stearate, can also be added. For oral administration in a capsule form, useful diluents include lactose and dried cornstarch. When aqueous suspensions are required for oral use, the active ingredient can be combined with emulsifying and suspending agents. If desired, certain sweetening, flavoring or coloring agents may also be added. In some aspects, the pharmaceutical composition is administered in the form of suppositories for rectal administration. These can be prepared by mixing the modified piggyBat transposon nucleic
4239-111427-02 acid/vector and/or a nucleic acid or vector encoding a modified piggyBat transposase with a suitable non-irritating excipient that is solid at room temperature but liquid at rectal temperature and therefore will melt in the rectum to release the drug. Such materials include cocoa butter, beeswax and polyethylene glycols. In other aspects, the pharmaceutical composition is administered topically, such as when the target of treatment includes areas or organs readily accessible by topical application, including diseases of the eye, the skin, or the lower intestinal tract. Suitable topical formulations are readily prepared for each of these areas or organs using known methods. In some examples, the pharmaceutical composition is formulated in a suitable ointment containing the transposition components suspended or dissolved in one or more carriers. Carriers for topical administration of the components disclosed herein include, but are not limited to, mineral oil, liquid petrolatum, white petrolatum, propylene glycol, polyoxyethylene, polyoxypropylene component, emulsifying wax and water. Alternatively, the pharmaceutical composition is formulated in a suitable lotion or cream containing the active components suspended or dissolved in one or more pharmaceutically acceptable carriers. Suitable carriers include, but are not limited to, mineral oil, sorbitan monostearate, polysorbate 60, cetyl esters wax, cetearyl alcohol, 2-octyldodecanol, benzyl alcohol and water. In some examples of ophthalmic administration, the pharmaceutical composition can be formulated as micronized suspensions in isotonic, pH adjusted sterile saline, or as solutions in isotonic, pH adjusted sterile saline, either with or without a preservative such as benzylalkonium chloride. In other examples of ophthalmic uses, the pharmaceutical composition is formulated in an ointment, such as petrolatum. In other aspects, the pharmaceutical composition is administered by nasal aerosol or inhalation. Such compositions are prepared according to well-known techniques and may be prepared as solutions in saline, employing benzyl alcohol or other suitable preservatives, absorption promoters to enhance bioavailability, fluorocarbons, and/or other conventional solubilizing or dispersing agents. The amount of each component (e.g., transposon and transposase nucleic acids/vectors) that may be combined with the carrier materials to produce a composition in a single dosage form can vary depending upon, for example, the host to be treated and the particular mode of administration. A specific dosage and treatment regimen for any particular subject can depend upon a variety of factors, including the subject’s age, body weight, general health, gender, and diet. EXAMPLES The following examples are provided to illustrate particular features of certain aspects of the disclosure, but the scope of the claims should not be limited to those features exemplified.
4239-111427-02 Example 1: Materials and Methods This example describes the materials and experimental procedures for the studies described in Examples 2-9. Plasmid constructs All helper plasmids for expression of the transposases were based on a pFV4a-piggyBat vector which was constructed from the pFV4a-RepHel plasmid (Grabundzija et al., Nat Commun 7:10716, 2016; Grabundzija et al., Nat Commun 9:1278, 2018) by cloning in a codon-optimized open reading frame of the Myotis lucifugus (M. luc) piggyBat gene using NotI and SpeI enzymes. The piggyBat-4StoA-2xCRD and 2xCRD inserts were ordered as gene blocks from GenScript and cloned into the pFV4a-piggyBat plasmid. The donor plasmids which contain the TIRs were constructed from the p2NGFPmini plasmid, which contains both a puromycin resistance cassette and a GFP gene between the transposon ends, by cloning in the left 153bps (NruI-XhoI) and right 208bps (NheI-SphI) of the TIRs of the piggyBat transposon (Mitra et al., Proc. Natl. Acad. Sci. USA 110, 234–239, 2013). All plasmid construction was performed by GenScript and confirmed by diagnostic restriction digesting and DNA sequencing. Purification of piggyBat The plasmid for pD2610-pBat was transfected into 500 ml EXPI293F cells (Thermo Fisher Scientific) for transient protein expression using PEI and harvested 3 days later and stored at −80°C until use. Cells expressing maltose binding protein (MBP)-tagged pBat were resuspended in lysis buffer containing 25 mM TrisCl pH 7.5, 500 mM NaCl, 1 mM TCEP, and protease inhibitor cocktail (Roche). The cells were lysed by sonication, and cell lysates were centrifuged at ~25,500 rpm for 45 minutes at 4°C. The supernatant was filtered and mixed with 10 ml amylose resin (New England BioLabs) equilibrated with lysis buffer, and after one hour of rotation, the mixture was loaded onto a gravity flow column and washed with lysis buffer. The protein was eluted with 50 ml elution buffer (25 mM TrisCl pH 7.5, 500 mM NaCl, 10 mM maltose, 1 mM TCEP, and protease inhibitor cocktail), and then incubated with TEV protease and dialyzed against dialysis buffer (50 mM TrisCl pH 7.5, 500 mM NaCl, and 1 mM TCEP) overnight at 4°C. MBP and TEV protease were separated from pBat using a 5 ml HiTrap Heparin HP column (GE Healthcare) by linear gradient elution from 500 mM to 1 M NaCl. Purified pBat was dialyzed overnight against storage buffer (50 mM TrisCl pH 7.5, 500 mM NaCl, 15% glycerol and 1 mM TCEP), frozen and stored at -80°C until use. Cell culture, transfection, and colony count assays HEK293T cells were cultured using standard procedures. For transfection, cells were seeded at a density of 0.5 x 106 cells per well in a six-well plate and transfected one day later, initially with 1.5 µg of total plasmid DNA, containing 1 µg of transposon (donor) and 0.5 µg of transposase
4239-111427-02 (helper) plasmid DNA using Lipofectamine 3000 (Invitrogen), according to the manufacturer’s instructions. For later experiments as indicated, the amount of DNA was decreased to 10 ng donor and 20 ng helper. At 48 hours post-transfection, cells were trypsinized and diluted into 100 mm dishes followed by selection with 2 µg/ml of puromycin for ~10 days, with media changes every three days. Either a 50-fold or 400-fold dilution was plated into selection media as indicated and subsequently changed every 3 days. Plates were then fixed using 4% formaldehyde in phosphate-buffered saline (PBS), stained with 1% methylene blue in PBS and counted. Cryo-EM specimen preparation Purified MLT transposase (5 mg/mL) and LE44 DNA were mixed in a 2:1 protein-to-DNA ratio and dialyzed overnight at 4°C against 25 mM Tris pH 7.5, 200 mM NaCl, 0.5 mM TCEP and 5 mM CaCl2. The sample was then run on a size exclusion chromatography column (Superdex 200 Increase 3.2/300, GE Healthcare) equilibrated with the sample buffer, attached to an Äkta system kept at 4°C. The fraction containing the MLT/LE44 complex was selected moving forward. The protein concentration was estimated at ~0.5 mg/mL by comparing the SDS-PAGE band intensity of the fraction with that of a dilution series of purified MLT. The MLT/LE44 complex was applied on a gold grid covered with a holey carbon film (Protochips C-Flat, R1.2/1.3, 300 mesh) freshly glow discharged for 30 seconds at 15 mA (PELCO easiGlow). The specimen was prepared with a Vitrobot Mark IV (FEI) rapid plunging device with the chamber at 16°C and 100% humidity. Three µL of sample were applied on the grid and after 5 seconds the excess of sample was blotted for 2.5 seconds (force 4) and immediately flash frozen in liquid ethane cooled by liquid nitrogen. Cryo-EM data collection Approximately 8300 movies were collected with a 300 kV Titan Krios TEM (FEI) equipped with a K3 direct electron detector camera (Gatan) and an energy filter (20 eV slit width). The movies were recorded in super resolution mode at a nominal magnification of 105 kx, corresponding to a calibrated pixel size of 0.43 Å, and a defocus range from -0.8 to -2.0 µm. The acquisition was supervised by the semi-automated program SerialEM (Mastronarde et al., J Struct Biol 152:36-51, 2005). The dose rate on the camera was set at 16.4 electron per physical pixel per second. The total exposure time for each movie was 2.2 seconds with a total exposure dose of 48.8 e−/Å2 (1.11 e−/Å2 per frame). Each movie was composed of 44 frames (50 ms exposure per frame). Cryo-EM single particle analysis The cryo-EM movies preprocessing and the single particles analysis were performed with RELION 4.0.1 ran on the NIH HPC Biowulf cluster (online at hpc.nih.gov) (Zivanov et al., Elife 7:e42166, 2018; Scheres, J Mol Biol 415:406-418, 2012; Scheres, J Struct Biol 180:519-530, 2012). The movies were motion corrected with RELION’s own implementation and binned by a factor 2,
4239-111427-02 resulting in a pixel size of 0.86 Å for further processing. The contrast transfer function (CTF) parameters were estimated with CTFFind 4.1.14 (Rohou et al., J Struct Biol 192:216-221, 2015). Nine hundred particles were manually picked with Topaz-Denoise turned on. The best particles (600) were selected through 2D classification and used to train the neural network of the particle picker Topaz 0.2.5. (Bepler et al., Nat Methods 16(11):1153-1160, 2019; Bepler et al., Nat Commun 11:1-12, 2020), to eventually automatically pick ~3.5 million particles on the whole dataset (0.5 picking threshold value). After extraction (240 pixel box size), the particles were purified through a first 2D classification. Two types of particles corresponding to the MLT/LE44 complex and the unbound MLT were identified and pooled into two separated selection jobs. Reconstruction of the MLT/LE44 complex The stack of MLT/LE44 particles was further purified by 2D classification/selection jobs. The resulting 795,701 particles were used to generate an initial reference-free 3D model that was used as reference for the 3D classification of the particles (six classes). The best class (class #1) was selected (162,244 particles) and gold-standard refined to 4.0 Å resolution at 0.143 FSC and then 3.6 Å after CTF refinement and particle polishing. The handedness of the map was corrected in ChimeraX (Pettersen et al., J Comput Chem 25:1605-1612, 2004; Goddard et al., Protein Sci 27:14-25, 2018). The final map was sharpened and denoised with DeepEMhancer and this map was used for model building (Sanchez-Garcia et al., Commun Biol 4:1-8, 2021). Reconstruction of the MLT dimer The stack of 698,327 MLT particles was used to generate an initial reference-free 3D model that was used as reference for the 3D classification of the particles (six classes). The particles of the best two classes (classes #3 and #6, 583,943 particles) were pooled and subjected to another 3D classification (4 classes). One resulting class was contaminated with MLT/LE44 particles and only pure apo-MLT classes were selected (classes#1, #2 and #4, 365,728 particles) and gold-standard refined to generate a map at 3.7 Å resolution (at 0.143 FSC). A mask of the whole map was used to run a 3D classification restricted to angular local search (4 classes) on the same pool of particles to better resolve local differences. The 3D classes #3 and #4 presented density between both MLT protomers and were selected (255,658 particles) to run a gold-standard 3D refinement that gave a map at 3.8 Å resolution (at 0.143 FSC) and then 3.6 Å after CTF refinement and particle polishing. The handedness of the map was corrected in ChimeraX (Pettersen et al., J Comput Chem 25:1605-1612, 2004; Goddard et al., Protein Sci 27:14-25, 2018). MLT/LE44 atomic model building In COOT 0.9 EL (Emsley et al., Acta Crystallogr Sect D Biol Crystallogr 66:486-501, 2010; Casañal et al., Protein Sci 29:1069-1078, 2020), the ideal B-DNA atomic model of LE44 was
4239-111427-02 generated and manually positioned into MLT/LE44 final map. All-molecule 5 Å self-restraints (Geman-McClure set at 0.1) were generated to perform an all-atom real space refinement of the DNA duplex against the MLT/LE44 cryo-EM map. The base pairs 36 to 44 and the single-strand overhang were deleted from the model due to poor or a lack of density. Two copies of the AlphaFold2 model of MLT were individually rigid body fitted into the MLT/LE44 final map. The models mostly sit in the map except the CRD domain that had to be specifically fitted (UCSF Chimera) (Pettersen et al., J Comput Chem 25:1605-1612, 2004). The MLT chains were then real space refined against the cryoEM map in COOT 0.9 EL. The regions of the proteins laying outside the map were trimmed and rebuilt when possible. The MLT and LE44 models were then joined into a single model to be real space refined in Phenix 1.19.2. The final model was obtained after several rounds of model rebuilding/fixing in COOT and Phenix refinement. Example 2: Identification of DNA sequence motifs in the pBat LE and RE When the DNA sequencing on the LE extending from the transposon end was examined, two distinct DNA motifs were identified that recur three times (FIG.2A, in box). One motif, 5'-GCGGGA (FIG.2A), is found at bp 11-16 (designated G1LE), 55-60 (G2LE), and 99-104 (G3LE). The second motif (in purple) begins three bp downstream of each repeating 5'-GCGGGA and resembles an imperfect palindrome. The most interior palindrome-like sequence extends from bp 107-120 (P3LE). There are no other occurrences of these two motifs elsewhere in the full 586-bp LE. The location and spacing of the two motifs closest to the transposon LE terminus resemble those of two repeats found on the pB LE (FIG.2B, in box). The cryo-electron microscopy (cryo-EM) structure of the pB dimer bound to two LE35 oligonucleotides (Chen et al., Nature Commun.11, 3446, 2020) (shown schematically in FIG.2C, left) demonstrated that one motif is bound predominantly through interactions with the RNase-H like catalytic domain whereas the palindromic sequence binds two C-terminal cysteine-rich domains (CRD), one from each of the monomers in the dimer. There was an apparent absence of corresponding sequence motifs on the pBat RE. The 5'- GCGGGA motif occurs twice but in the opposite orientation relative to the transposon end (bp 27-22, G2RE; and 68-63, G4RE) although there are truncated motifs at bp 12-15 (G1RE) and 44-48 (G3RE). No regions were identified on the RE that would correspond to the imperfect palindrome found three time on the LE. In this, pBat differs from pB whose LE and RE contain identical 19-bp palindromic sequences albeit different distances from the transposon tips (FIG.2B). Example 3: Cryo-EM structure of the pBat pre-synaptic complex assembled on LE44 To precisely define the sequence requirements of transposon end recognition, pBat complexes were assembled with an LE44 oligonucleotide and the three-dimensional structure was solved at 3.6Å using cryo-EM (FIGS.6, 7A-7D and 11). In the structure, a pBat dimer binds a single LE44 with
4239-111427-02 contributions from both monomers (FIG.3A). Each pBat monomer is composed of a large core domain similar to that of the piggyBac DNA transposase to which it can be aligned between residues 87 and 475 (pB residues 119-513) at an r.m.s. of 1.8Å over 329 α-carbon positions. The core consists of the catalytic subdomain with an RNaseH-like topology that contains the DDD active site (D237, D309, D413) and a predominantly β-stranded insertion subdomain, that are in turn inserted into an all α-helical subdomain. pBat has no potential density between residues 1 and 78 suggesting disorder, although residues 10 to 17 were assigned to density observed but only for one of the polypeptide chains of the dimer. A similarly disordered N-terminus has also been observed for pB (Chen et al., Nat Commun.11, 3446, 2020). The second major domain is a cysteine-rich C-terminal site-specific DNA binding domain (CRD) spanning residues S494 and Y572, with clear electrostatic potential density. The linker connecting the two domains is disordered in both monomers (residues 477-494 in one monomer and 480-493 in the other). Unlike the core domain, the pBat CRD is not structurally homologous to the CRD of pB: it has a different topological fold and recognizes a palindrome unrelated in sequence to that of pB. Each CRD has a small, central three-stranded antiparallel β-sheet and binds two Zn2+ ions, one Zn2+ using C3H1 ligands and the other with a C2H2 coordination sphere (FIG.8A). The two CRDs contributed by each transposase monomer together bind the imperfect palindrome (P1LE, FIG.2A). The symmetry axis of the CRD dimer formed on P1LE differs from that of the core domain, yielding an asymmetric assembly as previously observed in pB transpososomes (Chen et al., Nat Commun.11, 3446, 2020). The Zn finger topology of the pBat CRD is clearly unusual as a DALI search for structural homologs (Holm et al., Protein Sci.32, e4519, 2023) yielded only three similar Zn2+ binding domains (PDB codes 1X4S, 7SEK, and 3CXL), all contained within larger proteins with diverse functions apparently unrelated to DNA binding (FIG.8B). In the ZNHIT2 protein, the homologous domain (1X4S) plays a role in assembly of the U5 small nuclear ribonucleoprotein particle (He et al., Protein Sci.16, 1577-1587, 2007; Cloutier et al., Nat Commun.8, 15615, 2017) and attempts to demonstrate DNA binding were not successful (He et al., Protein Sci.16, 1577-1587, 2007). Methionine aminopeptidase 1 (7SEK) is involved in regulating zinc homeostasis (Weiss et al., Cell 185, 2148- 2163, 2022) and human chimerin 1 (3CXL) is a GTPase activating protein (Yang et al., Biochem J. 403, 1-12, 2007). Interestingly, and not detected by DALI, the topology seen in chimerin is also shared with the C1 domain of protein kinase C (Katti et al., Nat. Commun.13, 2695, 2022) which binds the second messenger diacylglycerol (7L92). Despite the topological similarity to the pBat CRD, the chimerin and PKC C1 domains have two C3H1 Zn2+ binding sites. The interactions between the pBat CRDs and the imperfect palindrome result in a ~60º bend in the DNA as it exits from the core domain to curl around the CRD dimer (see FIG.3A). The core domain is largely responsible for recognition of the GCGGGA motif and six bases are specifically recognized by amino acid side chains (G11 by R189; G-12 by R185: C-13 by D186; G14 by K134; and C- 15 and T-16 by R497 of the CTD).
4239-111427-02 Palindrome recognition is mediated exclusively by the two CRDs largely through interactions involving two loops from each (FIGS.3A-3E). In the first half of the P1LE palindrome, the loop following the first β-strand of the central three-stranded beta sheet is in the major groove where it forms the dominant share of palindrome recognition (in C1 domains, this loop is much shorter and forms the diacylglycerol binding site). As shown in FIG.3D, E545 and R543 contact C19, G-19, and C- 20 in the major groove. A loop between G495 and T502 (FIG.3C) sits in the adjacent minor groove. Importantly, this loop also mediates contacts with the core and CRD domains such that the overall effect is to pack the CRD and the core domain of the same monomer together. The interactions with the second half of the palindrome by the second CRD are less well defined in the density, with K527 and R543 contacting and A-28 and G30, respectively. Collectively, these interactions are completely different from those of the pB CRD with DNA, due to the different CRD topologies and the G495- T502/minor groove interactions that are not observed in the pB cryo-EM structures (interactions are shown schematically in FIG.10). Overall, the core domain that is linked to the first CRD forms most of the interactions with DNA; however, the trajectory of the DNA is such that the transposon end is directed to the active site of the second monomer. Perhaps because the DNA used in the experiment was blunt-ended at the transposon tip, DNA density was not observed at the active site. Thus, the core domain/transposon end interaction in the presynaptic complex are in trans, confirming the typically observed arrangement in DNA transpososomes. However, as both CRDs interact with the LE, the interactions are neither cis nor trans but both. Example 4: Both transposon ends are required and shortening either end increases transposition activity With the presynaptic complex defining the details of LE transposon end recognition, the sequence data and footprint data were further considered. The pBat LE has three sequential G1LE+P1LE-like motifs suggesting three transposase dimer binding sites. In contrast, there is no evidence of P1LE-like palindromes on the RE, and only two G1LE-like G-rich tracts were identified on the RE that could potentially be bound by pBat core domain in a manner that would promote transposon end synapsis (G1RE and G3RE). In order to investigate the role of these sequence motifs in transposition, a colony count transposition assay in human cells was used. In this assay, the pBat transposase was expressed from one plasmid and a second plasmid contained a puromycin expression cassette flanked by piggyBat transposon ends (FIG.4A). Upon co-transfection into HEK293T cells, wild-type (WT) pBat and the active LE153 and RE208 ends ("LE/RE") showed only moderate transposition (FIG.5B) consistent with previous reports (Mitra et al., Proc. Natl. Acad. Sci. USA 110, 234–239, 2013; Sutrave et al., Mol. Ther. Methods Clin. Dev.25, 250-263, 2022; Beckermann et al., Nucleic Acids Res.49, 8135-8144, 2021). When symmetrized LE/LE or RE/RE transposon donors
4239-111427-02 were used, no transposition activity was observed (FIG.4B), indicating that both the LE and RE are needed together for transposition in vivo. To test the hypothesis that the motifs identified on the LE and RE were necessary and sufficient for transposition activity in cells, the effect of using shortened versions of LE and/or RE was investigated. As shown in FIG.4C for WT pBat under standard assay conditions, when the active RE was present, there was little change in transposition activity in truncating from the active LE (LE/RE) to LE132 (LE132/RE). However, when the innermost G3LE+P3LE pair of motifs was deleted (LE88/RE) there was a substantial (~7-fold) increase in activity. Further shortening to remove G2LE (LE44/RE) reduced activity to only ~20% of that of the full-length WT LE/RE. A similar but less dramatic trend was observed upon truncating the RE: LE/RE160 showed similar activity to the active ends (LE/RE) but truncation to LE/RE100 led to a ~2.1-fold increase in transposition activity and further truncation to LE/RE37 abolished activity. To more accurately estimate the changes in transposition activity due to shortening the transposon ends, the assays were repeated using fewer cells in the selection step to avoid saturating the colony count assay (FIG.4C, right). Under these conditions, truncating the LE to LE88 increased the measured activity ~8.7-fold, and combining the LE88 truncation with RE100 increased the activity slightly more. Taken together, the data indicated that both transposon ends contain interior (subterminal) sequences that restrict transposition activity in cells, with the most significant restriction due to the region corresponding to the G3LE-P3LE transposase binding site on the LE. LE88 (SEQ ID NO: 1) CACTTGGATTGCGGGAAACGAGTTAAGTCGGCTCGCGTGAATTGCGCGTACTCCGCGGG AGCCGTCTTAACTCGGTTCATATAGATTT RE100 (SEQ ID NO: 2) CACTACGGTGTCGGGTGAATTTCCCGCCAAAATATAAGATGGCGCGGGCAATTATTTAGG TTTCCCGCGCAAAATCAGAGTTAGTTTAAAATGCCTATAT Example 5: Targeted mutation of the pBat N-terminus results in hyperactivity A defining characteristic of members of the piggyBac superfamily is that they contain N- terminal regions predicted to be intrinsically disordered (Bouallègue et al., Genome Biol. Evol.9, 323- 339, 2017). The N-termini are non-conserved, have low pI values, and contain many negatively charged amino acids rendering it unlikely that they are involved in DNA binding. In the pB transpososome structure, the first 116 N-terminal amino acids showed no density (Chen et al., Nature Commun.11, 3446, 2020). However, this region contains several predicted casein kinase II (CKII) phosphorylation motifs (S/T-D/E-X-E/D) and their removal either by N-terminal truncation or mutation results in a stimulation of transposition activity (Luo et al., Nucleic Acids Res.50, 13128-
4239-111427-02 13142, 2022). Similarly, the pBat presynaptic complex structure has no observed potential density for the first 72 residues. Although the N-termini of pB and pBat cannot be aligned, pBat contains four predicted CKII phosphorylation sites within its first 40 amino acids (FIG.5A). To determine whether the predicted CKII sites impact transposition activity, the serine residues in the first four CKII motifs (S8, S24, S32, S37) were mutated to alanine (A) or to lysine (K) individually or all together. As shown in FIG.5B, three of the four alanine mutations showed some increase in activity when assayed with LE/RE, with S8A having the largest effect (~3-fold). When all four were mutated together ("4StoA"), the increase in activity was additive. Individual mutation of the same residues to lysine (in gold) had less effect but, in combination, the mutation of all four to lysine led to a net ~6-fold stimulation of activity comparable to that of the combined serine mutations. These results suggested that phosphorylation of the N-terminus of pBat likely plays a similar role in inhibiting transposition activity as has been observed for pB. To determine if the stimulatory effects of the 4StoA mutations and truncated transposon ends could be combined, revised assay conditions were used (25X less pBat transposase helper plasmid and 100X less of the various pBat donors). Under these conditions, the activity of WT pBat with the LE/RE donor was extremely low (FIG.5C) but allowed for a more accurate measure of enhanced transposition activities of the modified piggyBat systems. Relative to the unmodified piggyBat transposon system (WT pBat and LE/RE donor), the effect of the 4StoA mutations was again on the order of an ~7-fold increase in activity, but when combined with the LE88/RE donor modification, the measured increase was greater than 40-fold. Example 6: Duplication of the piggyBat transposase C-terminal domain allows transposition using symmetrized ends The piggyBac transpososome can be rendered hyperactive by symmetrizing the TIRs to an LE35/LE35 configuration and by fusing an additional CRD domain to the transposase C-terminal (Luo et al., Nucleic Acids Res.50, 13128-13142, 2022) (FIG.2C, right). To determine if a similar effect might be observed for pBat, based on the structural information obtained by cryo-EM, a modified pBat transposase was engineered in which a second copy of the linker region and a second CRD (residues 480-572) was appended to the C-terminus ("pBat-2xCRD"). When tested in combination with the activating 4StoA mutations ("pBat-4StoA-2xCRD"; SEQ ID NO: 3), transposition activity on an LE88/RE100 donor was ~two-fold higher than that of pBat-4StoA (FIG. 5D). pBat-4StoA-2xCRD was highly active on symmetrized LE88/LE88 ends whereas pBat-4StoA was completely nonfunctional for transposition (FIG.5E). pBat-4StoA-2xCRD demonstrated only very weak activity with an LE44/LE44 donor that contains only a single transposase binding site on each end (FIG.5E). In the pBat-4StoA-2xCRD sequence below, the four serine to alanine substitutions are shown in bold font and the duplicated CRD sequence is underlined.
4239-111427-02 pBat-4StoA-2xCRD (SEQ ID NO: 3) MAQHSDYADDEFCADKLSNYSCDADLENASTADEDSADDEVMVRPRTLRRRRISSSSSDSES DIEGGREEWSHVDNPPVLEDFLGHQGLNTDAVINNIEDAVKLFIGDDFFEFLVEESNRYYNQN RNNFKLSKKSLKWKDITPQEMKKFLGLIVLMGQVRKDRRDDYWTTEPWTETPYFGKTMTR DRFRQIWKAWHFNNNADIVNESDRLCKVRPVLDYFVPKFINIYKPHQQLSLDEGIVPWRGRL FFRVYNAGKIVKYGILVRLLCESDTGYICNMEIYCGEGKRLLETIQTVVSPYTDSWYHIYMDN YYNSVANCEALMKNKFRICGTIRKNRGIPKDFQTISLKKGETKFIRKNDILLQVWQSKKPVYL ISSIHSAEMEESQNIDRTSKKKIVKPNALIDYNKHMKGVDRADQYLSYYSILRRTVKWTKRLA MYMINCALFNSYAVYKSVRQRKMGFKMFLKQTAIHWLTDDIPEDMDIVPDLQPVPSTSGMR AKPPTSDPPCRLSMDMRKHTLQAIVGSGKKKNILRRCRVCSVHKLRSETRYMCKFCNIPLHK GACFEKYHTLKNYDMDIVPDLQPVPSTSGMRAKPPTSDPPCRLSMDMRKHTLQAIVGSGKK KNILRRCRVCSVHKLRSETRYMCKFCNIPLHKGACFEKYHTLKNY Example 7: Comparison of piggyBat transposase sequences This example describes a comparison between two different modified piggyBat transposase sequences: an initial modified transposase (“version 1”) and pBat-4StoA-2xCRD. The amino acid sequences of the wild-type piggyBat transposase, modified piggyBat transposase version 1, and modified piggyBat transposase pBat-4StoA-2xCRD are provided below and set forth as SEQ ID NO: 11, SEQ ID NO: 10 and SEQ ID NO: 3, respectively. Version 1 and pBat-4StoA-2xCRD differ by an additional 13 amino acids present in pBat-4StoA-2xCRD (underlined in the sequence below). Wild-type piggyBat transposase (SEQ ID NO: 11) MAQHSDYSDDEFCADKLSNYSCDSDLENASTSDEDSSDDEVMVRPRTLRRRRISSSSSDSESD IEGGREEWSHVDNPPVLEDFLGHQGLNTDAVINNIEDAVKLFIGDDFFEFLVEESNRYYNQNR NNFKLSKKSLKWKDITPQEMKKFLGLIVLMGQVRKDRRDDYWTTEPWTETPYFGKTMTRD RFRQIWKAWHFNNNADIVNESDRLCKVRPVLDYFVPKFINIYKPHQQLSLDEGIVPWRGRLF FRVYNAGKIVKYGILVRLLCESDTGYICNMEIYCGEGKRLLETIQTVVSPYTDSWYHIYMDN YYNSVANCEALMKNKFRICGTIRKNRGIPKDFQTISLKKGETKFIRKNDILLQVWQSKKPVYL ISSIHSAEMEESQNIDRTSKKKIVKPNALIDYNKHMKGVDRADQYLSYYSILRRTVKWTKRLA MYMINCALFNSYAVYKSVRQRKMGFKMFLKQTAIHWLTDDIPEDMDIVPDLQPVPSTSGMR AKPPTSDPPCRLSMDMRKHTLQAIVGSGKKKNILRRCRVCSVHKLRSETRYMCKFCNIPLHK GACFEKYHTLKNY Modified piggyBat transposase version 1 (SEQ ID NO: 10) MAQHSDYADDEFCADKLSNYSCDADLENASTADEDSADDEVMVRPRTLRRRRISSSSSDSES DIEGGREEWSHVDNPPVLEDFLGHQGLNTDAVINNIEDAVKLFIGDDFFEFLVEESNRYYNQN RNNFKLSKKSLKWKDITPQEMKKFLGLIVLMGQVRKDRRDDYWTTEPWTETPYFGKTMTR DRFRQIWKAWHFNNNADIVNESDRLCKVRPVLDYFVPKFINIYKPHQQLSLDEGIVPWRGRL FFRVYNAGKIVKYGILVRLLCESDTGYICNMEIYCGEGKRLLETIQTVVSPYTDSWYHIYMDN YYNSVANCEALMKNKFRICGTIRKNRGIPKDFQTISLKKGETKFIRKNDILLQVWQSKKPVYL ISSIHSAEMEESQNIDRTSKKKIVKPNALIDYNKHMKGVDRADQYLSYYSILRRTVKWTKRLA MYMINCALFNSYAVYKSVRQRKMGFKMFLKQTAIHWLTDDIPEDMDIVPDLQPVPSTSGMR AKPPTSDPPCRLSMDMRKHTLQAIVGSGKKKNILRRCRVCSVHKLRSETRYMCKFCNIPLHK GACFEKYHTLKNYSTSGMRAKPPTSDPPCRLSMDMRKHTLQAIVGSGKKKNILRRCRVCSV HKLRSETRYMCKFCNIPLHKGACFEKYHTLKNY pBat-4StoA-2xCRD (SEQ ID NO: 3)
4239-111427-02 MAQHSDYADDEFCADKLSNYSCDADLENASTADEDSADDEVMVRPRTLRRRRISSSSSDSES DIEGGREEWSHVDNPPVLEDFLGHQGLNTDAVINNIEDAVKLFIGDDFFEFLVEESNRYYNQN RNNFKLSKKSLKWKDITPQEMKKFLGLIVLMGQVRKDRRDDYWTTEPWTETPYFGKTMTR DRFRQIWKAWHFNNNADIVNESDRLCKVRPVLDYFVPKFINIYKPHQQLSLDEGIVPWRGRL FFRVYNAGKIVKYGILVRLLCESDTGYICNMEIYCGEGKRLLETIQTVVSPYTDSWYHIYMDN YYNSVANCEALMKNKFRICGTIRKNRGIPKDFQTISLKKGETKFIRKNDILLQVWQSKKPVYL ISSIHSAEMEESQNIDRTSKKKIVKPNALIDYNKHMKGVDRADQYLSYYSILRRTVKWTKRLA MYMINCALFNSYAVYKSVRQRKMGFKMFLKQTAIHWLTDDIPEDMDIVPDLQPVPSTSGMR AKPPTSDPPCRLSMDMRKHTLQAIVGSGKKKNILRRCRVCSVHKLRSETRYMCKFCNIPLHK GACFEKYHTLKNYDMDIVPDLQPVPSTSGMRAKPPTSDPPCRLSMDMRKHTLQAIVGSGKK KNILRRCRVCSVHKLRSETRYMCKFCNIPLHKGACFEKYHTLKNY In the version 1 modified piggyBat transposase, the C-terminal residue Y572 of one monomer was connected directly to S492 to make a 2xCRD transposase. However, a cryo-EM structure of this mutant showed that this design would not allow for DNA binding in cis because the distance between the connection points is too long (as shown by the arrow in FIG.12A, top). In the improved piggyBat transposase pBat-4StoA-2xCRD, residue Y572 was connected directly to D480 (FIG.12A, bottom); this allowed for in cis DNA binding. Transposition activity of WT, version 1, and pBat-4StoA-2xCRD were compared using a colony count transposition assay. The three transposase proteins were tested with LE88/LE88 or LE88/RE100 as the transposon. As shown in FIG.12B, pBat-4StoA-2xCRD exhibited significantly greater activity than both WT and version 1, particularly when used in combination with LE88/RE100. The studies described herein demonstrate that the piggyBac (pB) and piggyBat (pBat) transposon systems differ in important ways. In particular, the piggyBac transposon has two distinct DNA motifs that are found on both transposon ends; once structural information revealed the details of protein-DNA interactions, the most parsimonious interpretation of the transposon end organization is that pB activity requires the assembly of a tetramer to productively pair its two ends (see FIG.2C). Binding of two pB dimers is likely driven by a palindromic CTD binding site that occurs once on the LE and once on the RE but with differential spacing relative to the transposon ends. This allows for one active dimer to carry out chemical reactions at the transposon boundaries and an interior dimer apparently required for productive synapse. It is demonstrated herein that piggyBat follows a different paradigm because its transposon LE contains three dimer binding sites and the DNA motif distribution does not follow the same pattern as that of piggyBac. Furthermore, palindromic CTD binding sites were only identified on its LE. The discovery that the most interior pBat binding site is severely inhibitory represents a previously unseen mode of transposition activity restriction through the use of an additional transposase binding site. This is somewhat paradoxical as transposase binding to transposon ends is a necessity for activity, and several transposon systems require arrays of transposase binding sites on their transposon ends to facilitate end synapsis (Baker and Mizuuchi, Genes Dev.6, 2221-2232, 1992;
4239-111427-02 Arciszewska and Craig, Nucleic Acids Res.19, 5021-5029, 1991). Among eukaryotic transposons, longer transposon sequences are often required or contribute to increased transposition activity (Urasaki et al., Genetics 174, 639-649, 2006; Liu et al., Genetics 157, 817-830, 2001; Coupland et al., Proc. Natl Acad. Sci. USA 86, 9385–9388, 1989; Lannes et al., Nat. Commun.14, 4470, 2023), a phenomenon usually attributed to the need for repeated subterminal binding sites. It will be apparent that the precise details of the methods or compositions described may be varied or modified without departing from the spirit of the described aspects of the disclosure. We claim all such modifications and variations that fall within the scope and spirit of the claims below.