WO2025259780A1 - Polypeptides and methods for modifying nucleic acids - Google Patents

Polypeptides and methods for modifying nucleic acids

Info

Publication number
WO2025259780A1
WO2025259780A1 PCT/US2025/033195 US2025033195W WO2025259780A1 WO 2025259780 A1 WO2025259780 A1 WO 2025259780A1 US 2025033195 W US2025033195 W US 2025033195W WO 2025259780 A1 WO2025259780 A1 WO 2025259780A1
Authority
WO
WIPO (PCT)
Prior art keywords
rna
polypeptide
protein
nucleic acid
editing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2025/033195
Other languages
French (fr)
Inventor
Weixin Tang
Hao Yan
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
University of Chicago
Original Assignee
University of Chicago
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by University of Chicago filed Critical University of Chicago
Publication of WO2025259780A1 publication Critical patent/WO2025259780A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12YENZYMES
    • C12Y305/00Hydrolases acting on carbon-nitrogen bonds, other than peptide bonds (3.5)
    • C12Y305/04Hydrolases acting on carbon-nitrogen bonds, other than peptide bonds (3.5) in cyclic amidines (3.5.4)
    • C12Y305/04004Adenosine deaminase (3.5.4.4)
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/16Hydrolases (3) acting on ester bonds (3.1)
    • C12N9/22Ribonucleases [RNase]; Deoxyribonucleases [DNase]
    • C12N9/222Clustered regularly interspaced short palindromic repeats [CRISPR]-associated [CAS] enzymes
    • C12N9/226Class 2 CAS enzyme complex, e.g. single CAS protein
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/78Hydrolases (3) acting on carbon to nitrogen bonds other than peptide bonds (3.5)
    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K2319/00Fusion polypeptide
    • CCHEMISTRY; METALLURGY
    • C07ORGANIC CHEMISTRY
    • C07KPEPTIDES
    • C07K2319/00Fusion polypeptide
    • C07K2319/01Fusion polypeptide containing a localisation/targetting motif
    • C07K2319/09Fusion polypeptide containing a localisation/targetting motif containing a nuclear localisation signal

Definitions

  • This invention relates to the field of molecular biology, genetic engineering, and medicine.
  • RNA editing is a natural process widely occurring in eukaryotic organisms and their associated viruses 1 . Editing events insert, delete, or convert nucleobases, resulting in RNA sequences that differ from the encoding genes.
  • the most common type of RNA editing is adenosine-to-inosine (A-to-I) conversion mediated by adenosine deaminases acting on RNA (ADARs) 2 .
  • A-to-I adenosine-to-inosine
  • ADARs adenosine deaminases acting on RNA
  • CRISPR-Cas systems through their RNA-guided targeting mechanisms, grant nucleic acid editing the ultimate programmability 5 .
  • Many effector proteins are compatible with DNA-targeting CRISPR systems for site-specific base conversion, deletion, and insertion, including different families of cytosine deaminases in cytosine base editors (CBEs) 6 ' 8 , evolved E.
  • RNA-specific adenosine deaminases in adenine base editors (ABEs) 9 ' 13
  • PEs prime editors
  • Extension of CRISPR-based editing to RNA which still follows the relatively straightforward protein engineering strategy, relies on the identification of catalytically competent, CRISPR-compatible RNA-editing enzymes. Efforts in programmed RNA editing so far center around a single family of effector proteins, ADARs, that act on double-stranded RNA (dsRNA).
  • ADARs double-stranded RNA
  • the tRNA-specific adenosine deaminase protein may be an engineered adenosine deaminase (such as an engineered TadA adenosine deaminase), which may be engineered by directed evolution or mutagenesis.
  • the tRNA-specific adenosine deaminase protein is not a wild-type TadA protein.
  • the tRNA-specific adenosine deaminase may be engineered to remove wild-type editing specificity.
  • the tRNA-specific adenosine deaminase may be any adenosine deaminase disclosed herein.
  • the disease- associated adenosine is an /A7A-A-48T mutation.
  • the patient has, is suspected of having, or has been diagnosed with having a disease associated with a mutation to an adenosine in an RNA.
  • the patient has, is suspected of having, or has been diagnosed with having Van der Woude syndrome.
  • polypeptides comprising a tRNA-specific cytosine deaminase protein.
  • RNA-targeting protein are also disclosed.
  • the tRNA-specific cytosine deaminase protein is operably linked to the RNA-targeting protein.
  • the tRNA-specific cytosine deaminase protein is a TadA.CBE3.1, TadA.CBE2.4, TadA.CBE.DLL, TadA.CBE.DL, TadA.CBE1.46, or TadA.CBE1.52 cytosine deaminase.
  • the tRNA-specific cytosine deaminase protein comprises at least one mutation from a wild-type TadA protein.
  • the tRNA-specific cytosine deaminase protein comprises a functional portion of a TadA deaminase disclosed herein.
  • the tRNA-specific cytosine deaminase protein comprises an enzymatically active portion of a TadA deaminase disclosed herein.
  • the tRNA-specific cytosine deaminase is an engineered bacterial deaminase.
  • the tRNA-specific cytosine deaminase is an engineered mammalian deaminase, such as an engineered ADAT protein.
  • the tRNA-specific cytosine deaminase is an engineered bacterial, mammalian, plant, or fungal deaminase.
  • the tRNA-specific cytosine deaminase comprises an engineered Streptococcus deaminase. In certain aspects, the tRNA-specific cytosine deaminase comprises an engineered E. coli deaminase. In certain aspects, the tRNA-specific cytosine deaminase specifically recognizes single-stranded RNA.
  • the RNA-targeting protein may be a targeting protein capable of binding to a specifically defined sequence on an RNA.
  • the RNA-targeting protein may be any RNA- targeting protein disclosed herein.
  • the RNA-targeting protein is a Cas protein.
  • the RNA-targeting protein is a protospacer-adjacent motif (PAM)- targeting protein.
  • the RNA-targeting protein is a dead Cas protein.
  • a dead Cas protein includes a Cas protein with one or more point mutations that result in a loss of nucleolytic activity of the Cas protein but, in some aspects, do not impact its binding to a target.
  • the RNA-targeting protein is a Casl3d protein.
  • the Casl3d protein is a dead Casl3d (dCasl3d) protein.
  • the dCasl3d protein is a RfxCasl3d from Ruminococcus flavefaciens XPD3002.
  • the RNA-binding comprises a functional portion of a Cas protein disclosed herein.
  • the tRNA- specific cytosine deaminase protein comprises an enzymatically active portion of a TadA cytosine deaminase disclosed herein.
  • the TadA cytosine deaminase protein is operably linked to the N-terminus of the RNA-targeting protein. In certain aspects, the TadA cytosine deaminase protein is operably linked to the C-terminus of the RNA- targeting protein. In some aspects, the polypeptide comprises at least one nuclear localization signal sequence.
  • nucleic acids encoding any of the polypeptides disclosed herein.
  • the nucleic acid encodes a short guide RNA (sgRNA).
  • the sgRNA binds to a target RNA.
  • the target RNA is a disease-causing RNA.
  • the target RNA comprises an cytosine associated with a disease.
  • the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of the cytosine associated with the disease.
  • the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of an cytosine of interest.
  • expression vectors encoding any of the nucleic acids disclosed herein.
  • the expression vector may also encode a promoter.
  • the promoter may be a promoter for a target RNA, including any target RNA disclosed herein.
  • the promoter may be a promoter for a disease-associated RNA.
  • cells comprising any of the polypeptides disclosed herein, any of the nucleic acids disclosed herein, and/or any of the expression vectors disclosed herein.
  • the cell comprises an RNA having an cytosine associated with a disease.
  • the cell is a diseased cell.
  • systems comprising a polypeptide comprising a TadA cytosine deaminase protein operably linked to an RNA-targeting protein, and a short guide RNA. Also disclosed are systems comprising any polypeptide disclosed herein and any sgRNA disclosed herein.
  • Aspect 12 A nucleic acid encoding the polypeptide of any one of aspects 1 to 11
  • Aspect 44 The method of any one of aspects 40 to 43, wherein the patient has, is suspected of having, or has been diagnosed with a disease associated with a mutation to an adenosine in an RNA
  • Aspect 45 The method of any one of aspects 40 to 44, wherein the patient has, is suspected of having, or has been diagnosed with Van der Woude syndrome
  • a polypeptide comprising a tRNA-specific cytosine deaminase protein operably linked to an RNA-targeting protein Aspect 47.
  • Aspect 48 The polypeptide of aspect 46 or 47, wherein the tRNA-specific cytosine deaminase protein has at least one mutation from a TadA protein
  • Aspect 49 The polypeptide of any one of aspects 46 to 48, wherein the RNA-targeting protein is a Casl3d protein
  • Aspect 50 The polypeptide of aspect 49, wherein the Casl3d protein is a dead Casl3d (dCasl3d) protein
  • Aspect 51 The polypeptide of aspect 50, wherein the dCasl3d protein is an RfxCasl3d from *Ruminococcus flavefaciens* XPD3002
  • Aspect 52 The polypeptide of any one of aspects 46 to 51, wherein the TadA cytosine deaminase protein is operably linked to the RNA-targeting protein by a flexible linker
  • Aspect 53 The polypeptide of aspect 52, wherein the flexible linker is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 amino acids in length
  • Aspect 54 The polypeptide of any one of aspects 46 to 53, wherein the TadA cytosine deaminase protein is operably linked to the N-terminus of the RNA-targeting protein
  • Aspect 55 The polypeptide of any one of aspects 46 to 53, wherein the TadA cytosine deaminase protein is operably linked to the C-terminus of the RNA-targeting protein
  • Aspect 56 The polypeptide of any one of aspects 46 to 55, further comprising at least one nuclear localization signal sequence
  • Aspect 58 The nucleic acid of aspect 57, further encoding a short guide RNA (sgRNA)
  • sgRNA short guide RNA
  • Aspect 59 The nucleic acid of aspect 58, wherein the sgRNA binds to a target RNA
  • Aspect 60 The nucleic acid of aspect 59, wherein the target RNA is a disease-causing RNA
  • Aspect 61 The nucleic acid of aspect 59 or 60, wherein the target RNA comprises a cytosine associated with a disease
  • Aspect 62 The nucleic acid of aspect 61, wherein the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of the cytosine associated with the disease
  • Aspect 63 An expression vector encoding the nucleic acid of any one of aspects 57 to 61
  • Aspect 64 A cell comprising the polypeptide of any one of aspects 46 to 56, the nucleic acid of any one of aspects 12 to 17, or the expression vector of aspect 63
  • Aspect 65 The cell of aspect 64, comprising an RNA having a cytosine associated with a disease
  • Aspect 66 A system comprising a polypeptide comprising a tRNA-specific cytosine deaminase protein operably linked to an RNA-targeting protein; and a short guide RNA
  • Aspect 67 The system of aspect 66, wherein the polypeptide is the polypeptide of any one of aspects 1 to 11
  • Aspect 68 The system of aspect 66 or 67, wherein the short guide RNA binds to a target RNA
  • Aspect 69 The system of aspect 68, wherein the target RNA is a disease-causing RNA
  • a method for modifying one or more cytosine bases and/or for editing one or more cytosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the polypeptide of any one of aspects 46 to 56
  • Aspect 72 The method of aspect 71, wherein at least one of the cytosine bases is associated with a disease
  • Aspect 74 The method of aspect 73, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one cytosine associated with the disease
  • Aspect 75 The method of any one of aspects 71 to 74, wherein the contacting results in a conversion of at least one of the cytosine bases in the nucleic acid molecule to a uridine base
  • Aspect 76 The method of aspect 75, wherein the conversion is at least 15 % efficient
  • Aspect 78 A method for modifying cytosine bases and/or for editing cytosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the system of any one of aspects 66 to 70
  • Aspect 79 The method of aspect 78, wherein at least one of the cytosine bases is associated with a disease
  • Aspect 80 The method of aspect 79, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one cytosine associated with the disease
  • Aspect 81 The method of any one of aspects 78 to 80, wherein the contacting results in a conversion of at least one of the cytosine bases in the nucleic acid molecule to a uridine base
  • Aspect 82 The method of aspect 81, wherein the conversion is at least 15 % efficient
  • Aspect 83 The method of any one of aspects 78 to 82, wherein at least one additional cytosine base is not converted to a uridine
  • Aspect 84 A method of treating a disease in a patient, the method comprising administering to the patient an effective amount of the polypeptide of any one of aspects 1 to 11, the nucleic acid of any one of aspects 12 to 17, the expression vector of aspect 63, the cell of aspect 19 or 20, or the system of any one of aspects 21 to 25
  • Aspect 85 The method of aspect 84, wherein the patient has at least one cell comprising a disease-associated cytosine in an RNA
  • Aspect 86 The method of aspect 85, wherein the polypeptide converts the cytosine to a uridine
  • Aspect 87 The method of aspect 85, wherein the polypeptide converts at least 15 % of the disease-associated cytosines in the cell to uridine
  • Aspect 88 The method of any one of aspects 85 to 87, wherein the patient has, is suspected of having, or has been diagnosed with a disease associated with a mutation to a cytosine in an RNA.
  • Aspect 89 A polypeptide comprising 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any range derivable therein) identity to any one of SEQ ID NOs: l-29.
  • A, B, and/or C includes: A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C.
  • A, B, and/or C includes: A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C.
  • “and/or” operates as an inclusive or.
  • compositions and methods for their use can “comprise,” “consist essentially of,” or “consist of’ any of the ingredients or steps disclosed throughout the specification. Compositions and methods “consisting essentially of’ any of the ingredients or steps disclosed limits the scope of the claim to the specified materials or steps which do not materially affect the basic and novel characteristic of the claimed invention.
  • FIGs. 1A-1H show programmed RNA editing by dCas!3-deaminase fusions, a.
  • DECOR-mediated RNA editing constitutes two functional modules — an adenosine deaminase TadA8e and a dCasl3d complex.
  • dCasl3d gRNA Guided by the dCasl3d gRNA, DECOR recognizes target RNA and deaminates specified adenosine to inosine.
  • Inosine is recognized as guanosine by polymerases and ribosomes, serving as a functional equivalent of guanosine in the coding and decoding process, b.
  • TadA8e-enabled A-to-I editing at EZH2 -Al 179 guided by LwaCasl3a, PspCasl3b, Casl3X. l, mini-Casl3X. l and RfxCasl3d. h.
  • the gRNA binding site is annotated in the heatmap.
  • FIGs. 2A-2G show DECOR-mediated editing of coding and noncoding RNA.
  • a Design and placement of DECOR, REPAIRvl, and REPAIRv2 gRNAs.
  • b A-to-I editing mediated by DECOR and REPAIRvl at designated As in seven cellular RNAs.
  • c Heatmaps of all A edits observed in seven RNAs. The target As are aligned across seven RNAs as position 0. gRNA binding regions are boxed in grey rectangles. Editing rates are normalized to the highest editing level observed in a given RNA. For a zoomed-in view of editing distributions, see FIGs. 12A-12D. d.
  • REPAIR-max represents the maximum editing level observed at a given A among all REPAIR gRNAs.
  • REPAIR-all includes all REPAIR gRNA designs.
  • the solid line within the box represents the median, while upper and lower hinges depict the first and third quartiles. Whiskers extend from the minimum to the maximum.
  • FIGs. 3A-3G show transcriptome-wide specificity of DECOR, a. Transcriptomewide A-to-I editing by DECOR and REPAIRvl paired with EZH2 -targeting and non-targeting gRNA. b. Violin plots of A-to-I conversion rates observed at off-target sites. Solid lines represent the median, whereas dashed lines denote the first and third quantiles. Statistical analysis was done using a two-tailed Student's t-test. n.s. P > 0.05; ** P ⁇ 0.01; **** P ⁇ 0.0001. c.
  • FIGs. 4A-4F show efficiency and specificity of hfDECOR.
  • b Transcriptome-wide A-to-I editing by hfDECOR-F148A.
  • c Violin plots of off-target editing rates for hfDECOR-F148A. Solid lines represent the median, whereas dashed lines denote the first and third quantiles. Statistical analysis was done using a two-tailed Student's t-test. n.s. P > 0.05. d.
  • FIGs. 1 Venn diagram showing off-target sites shared by DECOR and hfDECOR-F148A.
  • e Venn diagram showing the overlap of off-target sites between REPAIRvl and hfDECOR- F148A.
  • f Sequence logos of A sites impacted by hfDECOR-F148A in the HEK293T transcriptome. All overlap analyses present here use data obtained with non-targeting gRNA. the inventors observe similar overlapping rates with /;Z7/2-targeting gRNA. All off-target profiling experiments were carried out in duplicate in HEK293T cells. [0045] FIGs.
  • 5A-5I show removal of a disease-causing upstream open reading frame (uORF) and rescue of protein expression by DECOR, a. Schematic of RNA editing in the coding region of GFP carrying a premature stop codon W58X. Four gRNAs tiled 1-31 nt upstream of the UAG codon were designed and assayed, b. Fluorescence imaging of HEK293T cells expressing GFP-W58X, RFP (control), and non-targeting and targeting DECOR. gRNA b that binds 12-41 nt upstream of the target A resulted in strongest GFP recovery. Scale bar, 100 pm. c. A-to-I editing in GFP mRNA measured by high-throughput sequencing, d.
  • pORF Expression of the pORF was quantified by measuring luminescence.
  • a plasmid carrying a Renilla luciferase gene was co-delivered as a transfection/cell number/cell state control, e. Normalized luciferase activity with healthy, disease-causing uORF-containing, and A-to-G mutated IRF6 5’ UTRs.
  • f. Normalized luciferase activity Transcription of the uORF-bearing luciferase reporter was driven by an HSV-TK promoter, g. A-to-I editing at IRF6-A- 9.
  • h Normalized luciferase activity.
  • FIGs. 6A-6B show normalized expression levels for AD ARI (a) and ADAR2 (b) in 55 human tissues. The figure was produced with data obtained from the Human Protein Atlas (https://www.proteinatlas.org). TPM: transcript per million.
  • FIGs. 8A-8B show programmed editing of EZH2 mRNA by adenosine deaminases guided by a Casl3d complex
  • Adenosine deaminases evaluated in this assay included, coli tRNA adenosine deaminase TadA and three TadA derivatives, TadA7.10, TadA8.20, and TadA8e.
  • FIGs. 9A-9D show programmed RNA editing by dCas!3d and TadA8e fusions with different signal peptides
  • a Architectures of dCasl3d and TadA8e fusion proteins
  • NT non-targeting
  • 10A-10B show programmed editing of ACTB mRNA by dCas!3 and TadA8e fusions, a. Architectures of dCasl3 and TadA8e fusion proteins.
  • Casl3 LwaCasl3a, PspCasl3b, RfxCasl3d, Casl3X. l, and mini-Casl3X. l.
  • FIGs. 12A-12D show heatmaps of DECOR- and REPAIRvl-mediated A-to-I editing in seven cellular RNAs.
  • a positions in RNA refer to the coordinates in NCBI accession (Table 1). Black lines indicate the gRNA binding regions. Editing rates are presented as the mean of three independent replicates, a. ACTB (left) and B4GALNT1 (right), b. EZH2 (left) and KRAS (right), c. MAL ATI (left) and METTL3 (right), d. STAT3.
  • FIGs. 13A-13C show programmed RNA editing by DECOR with varying linker lengths between TadA8e and dCasl3.
  • a. Architectures of DECOR b. A-to-I editing mediated by DECOR of varying linker lengths at designated As in three cellular RNAs.
  • c. Heatmaps of all A edits observed in target RNAs. All fusion proteins were evaluated in HEK293T cells with targeting and non-targeting (NT) gRNA (mean ⁇ s.d., n 3).
  • FIGs. 14A-14C show context preferences of DECOR, a. Schematic diagram of the identification process for flanking sequences preferred by DECOR.
  • DECOR was designed with gRNA binding 11 nt upstream of the target NAN sequences in GFP mRNA to enable editing.
  • High-throughput sequencing reads of reverse-transcribed GFP mRNA were demultiplexed and analyzed for A-to-I editing, b. Editing rates across all possible NAN motifs, c. Heatmap of editing rates observed at the target A with different 5’ and 3’ flanking nucleotides.
  • FIGs. 16A-16G show transcriptome-wide specificity of DECOR and ADAR- based RNA-editing systems, a. A-to-I editing detected at EZH2-KW19 by whole- transcriptome sequencing, b. A-to-I editing detected at AXES'- A 504 and MALAT1- 2A 6 by whole-transcriptome sequencing, c. Transcriptome-wide off-target sites of different editors detected in individual replicates and those persisting across two replicates, d. Venn diagram showing the overlap of off-target sites between EZH2 -targeting and NT REPAIRvl. e-f.
  • FIGs. 17A-17J show gRNA-independent off-target DNA editing by DECOR, a. Overview of the R-loop assay adapted to evaluate the off-target DNA-editing activity of DNA base editors and DECOR.
  • DECOR or SpABE8e was delivered into HEK293T cells alongside a catalytically inactive SaCas9 (dSaCas9) and a Sa sgRNA that targets a specified genomic locus.
  • dSaCas9 catalytically inactive SaCas9
  • Sa sgRNA that targets a specified genomic locus.
  • A-to-G conversion within the R-loop created by the dSaCas9 complex serves a surrogate for gRNA-independent off-target effects in DNA.
  • HEK4-M by SpABE8e when codelivered with dSaCas9 and a Sa sgRNA.
  • Six experiments were carried out in parallel to generate R-loops at 6 separate genomic loci.
  • c On-target A-to-I RNA editing at EZH2-A1179 by DECOR when codelivered with dSaCas9 and a Sa sgRNA targeting genomic site 1.
  • d-i Off-target A-to-G editing by DECOR and SpABE8e at R-loops created by dSaCas9. A sites edited to greater than 1% are plotted, j.
  • FIGs. 18A-18G show efficiency and specificity of hfDECOR.
  • a A-to-I editing mediated by hfDECOR detected at EZH2 -Al 179 by whole-transcriptome sequencing
  • b Transcriptome-wide off-target sites of hfDECOR detected by individual replicates and those persisting across replicates
  • c Transcriptome-wide off-target A-to-I editing by hfDECOR- V106G.
  • FIG. 19 shows programmed RNA editing with an evolved bacterial adenosine deaminase.
  • FIGs. 20A-20D shows RNA editing in HEK293T cells, a. EZH2 RNA editing in HEK293T cells, b. RPS5 RNA editing in HEK293T cells, c. UBE3A RNA editing in HEK293T cells, d. XIST RNA editing in HEK293T cells.
  • RNA editing is an attractive strategy to treat genetic disease. It has so far been limited to a single family of effector proteins, adenosine deaminases acting on RNA (ADARs). Aspects herein include bacterial deaminase-enabled recoding of RNA (DECOR), a CRISPR-based, ADAR-independent RNA-editing platform that delivers robust adenosine-to- inosine editing to user-specified sites in the human transcriptome. DECOR can exploit an evolved bacterial tRNA adenosine deaminase, TadA8e, and shows transcriptome-wide off- target effects significantly lower than ADAR-overexpressing RNA-editing platforms.
  • DECOR can exploit an evolved bacterial tRNA adenosine deaminase, TadA8e, and shows transcriptome-wide off- target effects significantly lower than ADAR-overexpressing RNA-editing platforms.
  • RNA-editing technologies include 1) it is the first ADAR-independent A- to-I editing platform, DECOR expands the current portfolio of programmed RNA-editing technologies.
  • DECOR has features distinct from existing technologies, e.g., sequence preference, cellular localization, and target scope (single-stranded RNA versus doublestranded RNA). 2) DECOR is broadly compatible with RNA-targeting CRISPR systems, delivering editing activity comparable or higher than existing RNA-editing technologies.
  • DECOR has lower off-target effects than REPAIRvl, the a previous RNA base editing system.
  • programmed C-to-U editing in RNA is an unsolved problem. With deaminases described herein, programmed C-to-U editing in RNA can be achieved.
  • a “protein” “peptide” or “polypeptide” refers to a molecule comprising at least five amino acid residues.
  • wild-type refers to the endogenous version of a molecule that occurs naturally in an organism.
  • wildtype versions of a protein or polypeptide are employed, however, in many aspects of the disclosure, a modified protein or polypeptide is employed to generate an immune response.
  • a “modified protein” or “modified polypeptide” or a “variant” refers to a protein or polypeptide whose chemical structure, particularly its amino acid sequence, is altered with respect to the wild-type protein or polypeptide.
  • a modified/variant protein or polypeptide has at least one modified activity or function (recognizing that proteins or polypeptides may have multiple activities or functions). It is specifically contemplated that a modified/variant protein or polypeptide may be altered with respect to one activity or function yet retain a wild-type activity or function in other respects, such as immunogenicity.
  • a protein is specifically mentioned herein, it is in general a reference to a native (wild-type) or recombinant (modified) protein or, optionally, a protein in which any signal sequence has been removed.
  • the protein may be isolated directly from the organism of which it is native, produced by recombinant DNA/exogenous expression methods, or produced by solid-phase peptide synthesis (SPPS) or other in vitro methods.
  • SPPS solid-phase peptide synthesis
  • nucleic acid segments and recombinant vectors incorporating nucleic acid sequences that encode a polypeptide (e.g., a tRNA-specific deaminase or fragment thereof and/or an RNA-binding protein or fragment thereof).
  • recombinant may be used in conjunction with a polypeptide or the name of a specific polypeptide, and this generally refers to a polypeptide produced from a nucleic acid molecule that has been manipulated in vitro or that is a replication product of such a molecule.
  • the size of a protein or polypeptide may comprise, but is not limited to, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22,
  • polypeptides may be mutated by truncation, rendering them shorter than their corresponding wild-type form, also, they might be altered by fusing or conjugating a heterologous protein or polypeptide sequence with a particular function (e.g., for targeting or localization, for enhanced immunogenicity, for purification purposes, etc.).
  • polypeptides, proteins, or polynucleotides encoding such polypeptides or proteins of the disclosure may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 (or any derivable range therein) or more variant amino acids or nucleic acid substitutions or be at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any derivable
  • the peptide or polypeptide is or is based on a human sequence. In certain aspects, the peptide or polypeptide is not naturally occurring and/or is in a combination of peptides or polypeptides.
  • polypeptides of the disclosure may include at least, at most, or exactly 1, 2, 3,
  • the polypeptide comprises one or more substitutions at one or more amino acid positions selected from amino acid 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16,
  • each substitution is independently chosen from an amino acid selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine; and wherein the polypeptide is or is at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any derivable range therein) sequence identity to one sequence disclosed herein.
  • the protein or polypeptide may comprise amino acids 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113
  • the protein or polypeptide may comprise amino acids 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113
  • the protein may comprise, comprise at least, or comprise at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112,
  • polypeptide, protein, or nucleic acid may comprise at least, at most, or exactly 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24,
  • nucleic acid molecule or polypeptide starting at position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28,
  • nucleotide as well as the protein, polypeptide, and peptide sequences for various genes have been previously disclosed, and may be found in the recognized computerized databases.
  • Two commonly used databases are the National Center for Biotechnology Information’s Genbank and GenPept databases (on the World Wide Web at ncbi.nlm.nih.gov/) and The Universal Protein Resource (UniProt; on the World Wide Web at uniprot.org).
  • Genbank and GenPept databases on the World Wide Web at ncbi.nlm.nih.gov/
  • the Universal Protein Resource UniProt; on the World Wide Web at uniprot.org.
  • the coding regions for these genes may be amplified and/or expressed using the techniques disclosed herein or as would be known to those of ordinary skill in the art.
  • compositions of the disclosure there is between about 0.001 mg and about 10 mg of total polypeptide, peptide, and/or protein per ml.
  • concentration of protein in a composition can be about, at least about or at most about 0.001, 0.010, 0.050, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5, 10.0 mg/ml or more (or any range derivable therein).
  • amino acid subunits of a protein may be substituted for other amino acids in a protein or polypeptide sequence with or without appreciable loss of binding capacity or enzymatic activity. Since it is the interactive capacity and nature of a protein that defines that protein’s functional activity, certain amino acid substitutions can be made in a protein sequence and in its corresponding DNA coding sequence, and nevertheless produce a protein with similar or desirable properties. It is thus contemplated by the inventors that various changes may be made in the DNA sequences of genes which encode proteins without appreciable loss of their biological utility or activity.
  • the term “functionally equivalent codon” is used herein to refer to codons that encode the same amino acid, such as the six different codons for arginine. Also considered are “neutral substitutions” or “neutral mutations” which refers to a change in the codon or codons that encode biologically equivalent amino acids.
  • Amino acid sequence variants of the disclosure can be substitutional, insertional, or deletion variants.
  • a variation in a polypeptide of the disclosure may affect 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more non-contiguous or contiguous amino acids of the protein or polypeptide, as compared to wild-type (or any range derivable therein).
  • a variant can comprise an amino acid sequence that is at least 50%, 60%, 70%, 80%, or 90%, including all values and ranges there between, identical to any sequence provided or referenced herein.
  • a variant can include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more substitute amino acids.
  • amino acid and nucleic acid sequences may include additional residues, such as additional N- or C-terminal amino acids, or 5' or 3' sequences, respectively, and yet still be essentially identical as set forth in one of the sequences disclosed herein, so long as the sequence meets the criteria set forth above, including the maintenance of biological protein activity where protein expression is concerned.
  • the addition of terminal sequences particularly applies to nucleic acid sequences that may, for example, include various non-coding sequences flanking either of the 5' or 3' portions of the coding region.
  • Deletion variants typically lack one or more residues of the native or wild type protein. Individual residues can be deleted or a number of contiguous amino acids can be deleted. A stop codon may be introduced (by substitution or insertion) into an encoding nucleic acid sequence to generate a truncated protein.
  • Insertional mutants typically involve the addition of amino acid residues at a nonterminal point in the polypeptide. This may include the insertion of one or more amino acid residues. Terminal additions may also be generated and can include fusion proteins which are multimers or concatemers of one or more peptides or polypeptides described or referenced herein.
  • Substitutional variants typically contain the exchange of one amino acid for another at one or more sites within the protein or polypeptide, and may be designed to modulate one or more properties of the polypeptide, with or without the loss of other functions or properties. Substitutions may be conservative, that is, one amino acid is replaced with one of similar chemical properties. “Conservative amino acid substitutions” may involve exchange of a member of one amino acid class with another member of the same class.
  • Conservative substitutions are well known in the art and include, for example, the changes of: alanine to serine; arginine to lysine; asparagine to glutamine or histidine; aspartate to glutamate; cysteine to serine; glutamine to asparagine; glutamate to aspartate; glycine to proline; histidine to asparagine or glutamine; isoleucine to leucine or valine; leucine to valine or isoleucine; lysine to arginine; methionine to leucine or isoleucine; phenylalanine to tyrosine, leucine or methionine; serine to threonine; threonine to serine; tryptophan to tyrosine; tyrosine to tryptophan or phenylalanine; and valine to isoleucine or leucine.
  • Conservative amino acid substitutions may encompass non-naturally occurring amino acid residues, which
  • substitutions may be “non-conservative”, such that a function or activity of the polypeptide is affected.
  • Non-conservative changes typically involve substituting an amino acid residue with one that is chemically dissimilar, such as a polar or charged amino acid for a nonpolar or uncharged amino acid, and vice versa.
  • Non-conservative substitutions may involve the exchange of a member of one of the amino acid classes for a member from another class.
  • One skilled in the art can determine suitable variants of polypeptides as set forth herein using well-known techniques.
  • One skilled in the art may identify suitable areas of the molecule that may be changed without destroying activity by targeting regions not believed to be important for activity.
  • the skilled artisan will also be able to identify amino acid residues and portions of the molecules that are conserved among similar proteins or polypeptides.
  • areas that may be important for biological activity or for structure may be subject to conservative amino acid substitutions without significantly altering the biological activity or without adversely affecting the protein or polypeptide structure.
  • the hydropathy index of amino acids may be considered.
  • hydropathy profile of a protein is calculated by assigning each amino acid a numerical value (“hydropathy index”) and then repetitively averaging these values along the peptide chain.
  • Each amino acid has been assigned a value based on its hydrophobicity and charge characteristics. They are: isoleucine (+4.5); valine (+4.2); leucine (+3.8); phenylalanine (+2.8); cysteine/cysteine (+2.5); methionine (+1.9); alanine (+1.8); glycine (-0.4); threonine (-0.7); serine (-0.8); tryptophan (-0.9); tyrosine (-1.3); proline (1.6); histidine (-3.2); glutamate (-3.5); glutamine (-3.5); aspartate (-3.5); asparagine (-3.5); lysine (-3.9); and arginine (-4.5).
  • hydropathy amino acid index in conferring interactive biologic function on a protein is generally understood in the art (Kyte et al., J. Mol. Biol. 157: 105-131 (1982)). It is accepted that the relative hydropathic character of the amino acid contributes to the secondary structure of the resultant protein or polypeptide, which in turn defines the interaction of the protein or polypeptide with other molecules, for example, enzymes, substrates, receptors, DNA, antibodies, antigens, and others. It is also known that certain amino acids may be substituted for other amino acids having a similar hydropathy index or score, and still retain a similar biological activity.
  • the substitution of amino acids whose hydropathy indices are within ⁇ 2 is included.
  • those that are within ⁇ 1 are included, and in other aspects of the invention, those within ⁇ 0.5 are included.
  • hydrophilicity values have been assigned to these amino acid residues: arginine (+3.0); lysine (+3.0); aspartate (+3.0+1); glutamate (+3.0+1); serine (+0.3); asparagine (+0.2); glutamine (+0.2); glycine (0); threonine ( _ 0.4); proline (-0.5+1); alanine ( _ 0.5); histidine ( _ 0.5); cysteine (—1.0); methionine (-1.3); valine (-1.5); leucine (-1.8); isoleucine (-1.8); tyrosine (-2.3); phenylalanine (-2.5); and tryptophan (-3.4).
  • the substitution of amino acids whose hydrophilicity values are within ⁇ 2 are included, in other aspects, those which are within ⁇ 1 are included, and in still other aspects, those within ⁇ 0.5 are included.
  • One skilled in the art can also analyze the three-dimensional structure and amino acid sequence in relation to that structure in similar proteins or polypeptides. In view of such information, one skilled in the art may predict the alignment of amino acid residues of an protein with respect to its three-dimensional structure. One skilled in the art may choose not to make changes to amino acid residues predicted to be on the surface of the protein, since such residues may be involved in important interactions with other molecules. Moreover, one skilled in the art may generate test variants containing a single amino acid substitution at each desired amino acid residue.
  • amino acid substitutions are made that: (1) reduce susceptibility to proteolysis, (2) reduce susceptibility to oxidation, (3) alter binding affinity for forming protein complexes, (4) alter ligand or antigen binding affinities, and/or (5) confer or modify other physicochemical or functional properties on such polypeptides.
  • single or multiple amino acid substitutions may be made in the naturally occurring sequence.
  • conservative amino acid substitutions can be used that do not substantially change the structural characteristics of the protein or polypeptide (e.g., one or more replacement amino acids that do not disrupt the secondary structure that characterizes the native protein).
  • nucleic acid sequences can exist in a variety of instances such as: isolated segments and recombinant vectors of incorporated sequences or recombinant polynucleotides encoding a protein (including any polypeptide disclosed herein), or a fragment, derivative, mutein, or variant thereof, polynucleotides sufficient for use as hybridization probes, PCR primers or sequencing primers for identifying, analyzing, mutating or amplifying a polynucleotide encoding a polypeptide, anti-sense nucleic acids for inhibiting expression of a polynucleotide, and complementary sequences of the foregoing described herein.
  • Nucleic acids that encode one or multiple domains of the polypeptides disclosed herein are provided herein are also provided.
  • Nucleic acids encoding fusion proteins that include these peptides are also provided.
  • the nucleic acids can be single-stranded or double-stranded and can comprise RNA and/or DNA nucleotides and artificial variants thereof (e.g., peptide nucleic acids).
  • polynucleotide refers to a nucleic acid molecule that either is recombinant or has been isolated from total genomic nucleic acid. Included within the term “polynucleotide” are oligonucleotides (nucleic acids 100 residues or less in length), recombinant vectors, including, for example, plasmids, cosmids, phage, viruses, and the like. Polynucleotides include, in certain aspects, regulatory sequences, isolated substantially away from their naturally occurring genes or protein encoding sequences.
  • Polynucleotides may be single- stranded (coding or antisense) or double- stranded, and may be RNA, DNA (genomic, cDNA or synthetic), analogs thereof, or a combination thereof. Additional coding or noncoding sequences may, but need not, be present within a polynucleotide.
  • the term “gene,” “polynucleotide,” or “nucleic acid” is used to refer to a nucleic acid that encodes a protein, polypeptide, or peptide (including any sequences required for proper transcription, post-translational modification, or localization). As will be understood by those in the art, this term encompasses genomic sequences, expression cassettes, cDNA sequences, and smaller engineered nucleic acid segments that express, or may be adapted to express, proteins, polypeptides, domains, peptides, fusion proteins, and mutants.
  • a nucleic acid encoding all or part of a polypeptide may contain a contiguous nucleic acid sequence encoding all or a portion of such a polypeptide. It also is contemplated that a particular polypeptide may be encoded by nucleic acids containing variations having slightly different nucleic acid sequences but, nonetheless, encode the same or substantially similar protein.
  • polynucleotide variants having substantial identity to the sequences disclosed herein; those comprising at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity, including all values and ranges there between, compared to a polynucleotide sequence provided herein using the methods described herein (e.g., BLAST analysis using standard parameters).
  • the isolated polynucleotide will comprise a nucleotide sequence encoding a polypeptide that has at least 90%, preferably 95% and above, identity to an amino acid sequence described herein, over the entire length of the sequence; or a nucleotide sequence complementary to said isolated polynucleotide.
  • nucleic acid segments regardless of the length of the coding sequence itself, may be combined with other nucleic acid sequences, such as promoters, polyadenylation signals, additional restriction enzyme sites, multiple cloning sites, other coding segments, and the like, such that their overall length may vary considerably.
  • the nucleic acids can be any length.
  • nucleic acid fragments of almost any length may be employed, with the total length preferably being limited by the ease of preparation and use in the intended recombinant nucleic acid protocol.
  • a nucleic acid sequence may encode a polypeptide sequence with additional heterologous coding sequences, for example to allow for purification of the polypeptide, transport, secretion, post-translational modification, or for therapeutic benefits such as targeting or efficacy.
  • a tag or other heterologous polypeptide may be added to the modified polypeptide-encoding sequence, wherein “heterologous” refers to a polypeptide that is not the same as the modified polypeptide.
  • nucleic acids that hybridize to other nucleic acids under particular hybridization conditions are well known in the art. See, e.g., Current Protocols in Molecular Biology, John Wiley and Sons, N.Y. (1989), 6.3.1-6.3.6. As defined herein, a moderately stringent hybridization condition uses a prewashing solution containing 5* sodium chloride/sodium citrate (SSC), 0.5% SDS, 1.0 mM EDTA (pH 8.0), hybridization buffer of about 50% formamide, 6* SSC, and a hybridization temperature of 55° C.
  • SSC sodium chloride/sodium citrate
  • pH 8.0 0.5%
  • hybridization buffer of about 50% formamide
  • 6* SSC a hybridization temperature of 55° C.
  • a stringent hybridization condition hybridizes in 6*SSC at 45° C., followed by one or more washes in 0.1 * SSC, 0.2% SDS at 68° C.
  • nucleic acids comprising nucleotide sequence that are at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to each other typically remain hybridized to each other.
  • Changes can be introduced by mutation into a nucleic acid, thereby leading to changes in the amino acid sequence of a polypeptide that it encodes. Mutations can be introduced using any technique known in the art. In one aspect, one or more particular amino acid residues are changed using, for example, a site-directed mutagenesis protocol. In another aspect, one or more randomly selected residues are changed using, for example, a random mutagenesis protocol. However it is made, a mutant polypeptide can be expressed and screened for a desired property.
  • Mutations can be introduced into a nucleic acid without significantly altering the biological activity of a polypeptide that it encodes. For example, one can make nucleotide substitutions leading to amino acid substitutions at non-essential amino acid residues.
  • one or more mutations can be introduced into a nucleic acid that selectively changes the biological activity of a polypeptide that it encodes. See, eg., Romain Studer et al., Biochem. J. 449:581-594 (2013).
  • the mutation can quantitatively or qualitatively change the biological activity. Examples of quantitative changes include increasing, reducing or eliminating the activity.
  • C. Probes include increasing, reducing or eliminating the activity.
  • nucleic acid molecules are suitable for use as primers or hybridization probes for the detection of nucleic acid sequences.
  • a nucleic acid molecule can comprise only a portion of a nucleic acid sequence encoding a full-length polypeptide, for example, a fragment that can be used as a probe or primer or a fragment encoding an active portion of a given polypeptide.
  • the nucleic acid molecules may be used as probes or PCR primers for specific protein sequences.
  • a nucleic acid molecule probe may be used in diagnostic methods or a nucleic acid molecule PCR primer may be used to amplify regions of DNA that could be used, inter alia, to isolate nucleic acid sequences for use in producing proteins. See, eg., Gaily Kivi et al., BMC Biotechnol. 16:2 (2016).
  • the nucleic acid molecules are oligonucleotides.
  • Probes based on the desired sequence of a nucleic acid can be used to detect the nucleic acid or similar nucleic acids, for example, transcripts encoding a polypeptide of interest.
  • the probe can comprise a label group, e.g., a radioisotope, a fluorescent compound, an enzyme, or an enzyme co-factor. Such probes can be used to identify a cell that expresses the polypeptide.
  • nucleic acid molecule encoding polypeptides or peptides of the disclosure are generated by methods known in the art, e.g., expressed in any suitable recombinant expression system and allowed to assemble to form protein compositions.
  • the nucleic acid molecules may be used to express large quantities of polypeptides.
  • contemplated are expression vectors comprising a nucleic acid molecule encoding a polypeptide of the desired sequence or a portion thereof (e.g., a fragment containing one or more polypeptide or protein domains).
  • vectors and expression vectors may contain nucleic acid sequences that serve other functions as well.
  • DNAs encoding the polypeptides or peptides are inserted into expression vectors such that the gene area is operatively linked to transcriptional and translational control sequences.
  • expression vectors used in any of the host cells contain sequences for plasmid or virus maintenance and for cloning and expression of exogenous nucleotide sequences.
  • flanking sequences typically include one or more of the following operatively linked nucleotide sequences: a promoter, one or more enhancer sequences, an origin of replication, a transcriptional termination sequence, a complete intron sequence containing a donor and acceptor splice site, a sequence encoding a leader sequence for polypeptide secretion, a ribosome binding site, a polyadenylation sequence, a polylinker region for inserting the nucleic acid encoding the polypeptide to be expressed, and a selectable marker element.
  • a promoter one or more enhancer sequences
  • an origin of replication a transcriptional termination sequence
  • a complete intron sequence containing a donor and acceptor splice site a sequence encoding a leader sequence for polypeptide secretion
  • ribosome binding site a sequence encoding a leader sequence for polypeptide secretion
  • polyadenylation sequence a polylinker region for inserting the nucleic acid encoding the polypeptid
  • Prokaryote- and/or eukaryote-based systems can be employed for use with an aspect to produce nucleic acid sequences, or their cognate polypeptides, proteins and peptides.
  • Commercially and widely available systems include in but are not limited to bacterial, mammalian, yeast, and insect cell systems.
  • Different host cells have characteristic and specific mechanisms for the post-translational processing and modification of proteins. Appropriate cell lines or host systems can be chosen to ensure the correct modification and processing of the foreign protein expressed.
  • Those skilled in the art are able to express a vector to produce a nucleic acid sequence or its cognate polypeptide, protein, or peptide using an appropriate expression system.
  • nucleic acid delivery to effect expression of compositions are anticipated to include virtually any method by which a nucleic acid (e.g., DNA, including viral and nonviral vectors) can be introduced into a cell, a tissue or an organism, as described herein or as would be known to one of ordinary skill in the art.
  • a nucleic acid e.g., DNA, including viral and nonviral vectors
  • Such methods include, but are not limited to, direct delivery of DNA such as by injection (U.S. Patents 5,994,624,5,981,274, 5,945,100, 5,780,448, 5,736,524, 5,702,932, 5,656,610, 5,589,466 and 5,580,859, each incorporated herein by reference), including microinjection (Harland and Weintraub, 1985; U.S.
  • Patent 5,789,215 incorporated herein by reference
  • electroporation U.S. Patent No. 5,384,253, incorporated herein by reference
  • calcium phosphate precipitation Graham and Van Der Eb, 1973; Chen and Okayama, 1987; Rippe et al., 1990
  • DEAE dextran followed by polyethylene glycol
  • direct sonic loading Fechheimer et al., 1987
  • liposome mediated transfection Nicolau and Sene, 1982; Fraley et al., 1979; Nicolau et al., 1987; Wong et al., 1980; Kaneda et al., 1989; Kato et al., 1991
  • microprojectile bombardment PCT Application Nos.
  • Other methods include viral transduction, such as gene transfer by lentiviral or retroviral transduction.
  • contemplated are the use of host cells into which a recombinant expression vector has been introduced.
  • Proteins can be expressed in a variety of cell types.
  • An expression construct encoding an protein can be transfected into cells according to a variety of methods known in the art.
  • Vector DNA can be introduced into prokaryotic or eukaryotic cells via conventional transformation or transfection techniques. Some vectors may employ control sequences that allow it to be replicated and/or expressed in both prokaryotic and eukaryotic cells.
  • the protein expression construct can be placed under control of a promoter that is linked to a disease-associated RNA.
  • One of skill in the art would understand the conditions under which to incubate host cells to maintain them and to permit replication of a vector.
  • techniques and conditions that would allow large- scale production of vectors, as well as production of the nucleic acids encoded by vectors and their cognate polypeptides, proteins, or peptides.
  • a selectable marker e.g., for resistance to antibiotics
  • Cells stably transfected with the introduced nucleic acid can be identified by drug selection (e.g., cells that have incorporated the selectable marker gene will survive, while the other cells die), among other methods known in the arts.
  • the nucleic acid molecule encoding the polypeptides may be obtained from any source that produces proteins. Methods of isolating mRNA encoding an protein are well known in the art. See e.g., Sambrook et al., supra. The sequences of human heavy and light chain constant region genes are also known in the art. See, e.g., Kabat et al., 1991, supra.
  • kits for modifying and/or detecting modified adenosines in a target DNA may also include additional components that are useful for amplifying the nucleic acid, or sequencing the nucleic acid, or other applications of the present disclosure as described herein.
  • the kit may optionally provide additional components that are useful in the procedure. These optional components include buffers, capture reagents, developing reagents, labels, reacting surfaces, means for detection, control samples, instructions, and interpretive information.
  • the kit may also include reagents for DNA isolation and/or purification.
  • polypeptides capable of modifying nucleotides in certain RNA molecules, including tRNA-specific deaminases linked to an RNA-targeting molecule.
  • the polypeptide sequences of such polypeptides can include any of the following:
  • Example 1 Programmed RNA Editing with an Evolved Bacterial Adenosine Deaminase
  • DECOR bacteria deaminase-enabled recoding of RNA
  • FIG. 1A a new programmable RNA-editing platform
  • the inventors focus on deaminases that act on single-stranded nucleic acids, a substrate profile orthogonal to ADARs and ADAR derivatives.
  • DECOR exploits TadASe 11 , a variant of the E.coli TadA, to deposit A- to-I editing to sites programmed by CRISPR.
  • DECOR has a relaxed editing window, with the most strongly edited A positioned ⁇ 13 nucleotides (nt) downstream of the guide RNA-binding site.
  • DECOR functions robustly in a variety of human cells, delivering up to 53.1% A-to-I conversion to endogenous transcripts. Meanwhile, DECOR impacts 88% fewer off-target sites in the human transcriptome than REPAIRvl, the state-of- the-art CRISPR-based RNA-editing platform. With a prokaryotic deaminase catalyzing the A- to-I transition, DECOR shows an off-target profile orthogonal to REPAIR. Installation of high- fidelity mutations into TadA8e further reduces the off-target effects of DECOR to almost basal level while sustaining on-target activity.
  • the inventors applied DECOR to remove a diseasecausing upstream open reading frame (uORF) in the 5’ untranslated region (UTR) of interferon regulatory factor 6 (IRF6) and rescued translation of the primary ORF (pORF) from 12.3% to 36.5%.
  • uORF diseasecausing upstream open reading frame
  • IRF6 interferon regulatory factor 6
  • pORF primary ORF
  • rAPOBECl a cytosine deaminase that accepts both DNA and RNA 35 ;
  • A3 A an APOB EC family protein highly active in deaminating cytosine in single-stranded DNA (ssDNA) 8 ;
  • evoCDAl a deaminase evolved for robust deamination of deoxy cytidine 36 ;
  • RNA-targeting CRISPR system identified from Ruminococcus flavefaciens XPD3002 38, 39 , hereby referred to as Cast 3d. Effector proteins were fused to dCasl3d, a nuclease-deficient mutant of Casl3d, through a flexible linker (FIG. IB). The inventors appended two nuclear localization signal (NLS) peptides to dCasl3d to maximize its RNA- targeting efficiency 38 . Both termini of Casl3d are exposed and not intimately involved in RNA engagement based on the crystal structure 40 . The inventors therefore evaluated RNA-editing efficiency for constructs with effector proteins situated at either terminus of dCasl3d.
  • NLS nuclear localization signal
  • gRNAs processed guide RNAs
  • EZEI2 a histone lysine A-methyltransferase
  • gRNA targeting a sequence absent in the human transcriptome was applied as a negative control (non-targeting gRNA or NT gRNA).
  • the inventors harvested cells 28 h post-transfection and reverse transcribed the EZH2 mRNA into its complementary DNA (cDNA), which was subsequently amplified by polymerase chain reaction (PCR) and analyzed by next-generation sequencing (NGS). As reverse transcriptases pair U with A and I with C, C-to-U and A-to-I edits are detected as C-to-T and A-to-G mutations.
  • cDNA complementary DNA
  • PCR polymerase chain reaction
  • NGS next-generation sequencing
  • the effector orientation in the fusion protein also impacted the editing outcome, as the most strongly edited Cs for rAPOBECl-dCasl3d and dCasl3d-rAPOBECl are 29 nt apart. Little editing ( ⁇ 0.4%) was observed in the absence of the AZ//? -targeting gRNA, confirming that RNA editing was programmed.
  • the basal to modest editing level is consistent with the tight regulation of cytosine editing in RNA 42 and the known preference for DNA over RNA by some cytosine deaminases 43, 44 .
  • the tRNA adenosine deaminase TadA acts on the anti-codon loop of Arg tRNA (ACG) in E. coli 5 .
  • ACG Arg tRNA
  • wildtype TadA is unlikely to crosstalk with the cellular mRNA pool, even when guided by CRISPR.
  • Laboratory evolution has largely liberated the substrate restraint of TadA, enabling nucleic acid editing with minimal context dependence 9 .
  • the inventors therefore hypothesized that evolved TadAs are better suited than wildtype TadA as effector proteins for programmed RNA editing.
  • TadA8.20 and TadA8e The activity increment from TadA7.10 to TadA8.20 and TadA8e was anticipated, as the catalytic function of the latter two variants was improved via further directed evolution 10 ' 12 . Similar boosted editing activity has been reported for TadA8.20 and TadA8e when targeting DNA. Among different effector proteins and fusion designs, TadA8e-dCasl3d delivered the most robust editing. NLS peptides are essential for the observed activity, as their removal and replacement by nuclear export signal (NES) peptides both abolished editing (FIGs. 9A-9D). The inventors moved forward with this architecture and named the platform bacterial deaminase-enabled recoding of RNA (DECOR, FIG. 1A).
  • DECOR consists of two functional modules — TadA8e that catalyzes the A-to-I conversion and the dCasl3d/gRNA complex responsible for target search.
  • the latter function is broadly shared by RNA-targeting CRISPR systems.
  • the inventors therefore hypothesized that DECOR may be compatible with CRISPR systems beyond RfxCasl3d.
  • the inventors tethered TadA8e to the N- or C-terminus of nuclease-deficient LwaCasl3a 46 , PspCasl3b 18 , Casl3X.
  • the inventors designed DECOR and REPAIR systems to edit the same As — EZH2 -Al 179, ACTB-A122Q, B 4G ALN 7-A958, A7MS-A504, MALAT1-A2A 6, METTL3-A95Q, and 57XZ3-A1402 (Table 2).
  • REPAIRvl functions with spacer lengths ranging from 30 to 70 nt 18 . However, hybridization length beyond 30 nt increases CRISPR-independent editing and local off-target editing without benefiting on-target activity 49, 50 . Hence, the inventors chose 30 nt spacers.
  • ADARs deaminate As in dsRNA with a preference for As in A:C mismatches 2, 51 .
  • REPAIRvl shows high-level editing when the mismatch is close to the center of the spacer 18 , but screening is recommended to identify the optimal placement of the A:C mismatch.
  • the inventors designed four REPAIRvl gRNAs for each locus with the target A situated at different positions in the protospacer (13, 18, 22, and 26, FIG. 2A and Table 2). Indeed, the inventors did not find a single mismatch position that was constantly favored, and variable editing levels were observed for the four gRNAs at individual target sites (FIG. 2B).
  • Both DECOR and REPAIRvl showed minimal A-to-I conversion with the nontargeting gRNA, 0.8-1.8% for DECOR and 0.1-0.7% for REPAIRvl, with the exception of B 4G ALNT 1- ⁇ 95 for DECOR (17.4%) and A7US-A504 for REPAIRvl (8.4%, FIGs. 2b and 2c).
  • Background editing of B4GALNT1-A95 by DECOR can be partially attributed to nonspecific binding of the control gRNA, as gRNAs targeting EZH2 and KRAS both led to lower levels of editing at B4GALNT1- 95'& (7.5-7.9%, FIG. 11).
  • RNA MALAT1 was a particularly challenging substrate for REPAIRvl as all four gRNAs performed poorly at MALAT1- 2A 6 (0.6-4.1%).
  • DECOR edited this site robustly, delivering 39.8% A-to-I editing.
  • the difference in editing efficiency may be explained by the intrinsic difficulty in targeting nuclear-localized MALAT1 with cytoplasm-localized REPAIRvl . This observation also highlights the functional complementarity of REPAIRvl and DECOR.
  • both REPAIRvl and DECOR show moderate editing in regions surrounding the target A sites (FIG. 2C and FIGs. 12A-12D). Shortening the linker between TadA8e and dCasl3d had a modest impact on editing potency and did not result in a narrower editing window (FIGs. 13A-13C). Notably, editing activity persisted even when the 33 amino acid (aa) linker was truncated to two amino acids, indicating considerable flexibility in the protein termini.
  • the inventorsected DECOR’s context preferences using a high-throughput assay and found that DECOR was more potent towards YA (Y U and C) (Supplementary Note 1 and FIGs. 14A-14C), a context preference likely inherited from the parent Tad A protein 52 .
  • Both REPAIRv2 and REPAIR.tl conferred lower editing across seven target sites with 3-4 gRNAs designed to vary spacer placement: 10.1 ⁇ 12.5% and 1.2 ⁇ 2.4% considering all target site and gRNA combinations; 17.6 ⁇ 18.0% and 2.3 ⁇ 3.9% considering the best gRNA design at each target site (FIGs. 2a, 2d, and FIGs. 15A-15B).
  • dCasl3d-ADARoD-E488Q was generally more active than ADARoD-E488Q-dCasl3d (16.5 ⁇ 12.4% vs. 5.8 ⁇ 9.1% editing across seven target sites; FIG.
  • DECOR showed non-negligible editing with non-targeting gRNA in some cases, which prompted us to interrogate the off-target effects of DECOR on the whole transcriptome and compared side by side with REPAIRvl, REPAIRv2, REPAIR.tl, and dCasl3d-ADARoD- E488Q. All five RNA-editing platforms were assayed with EZH2 -targeting and non-targeting gRNAs in HEK293T cells in duplicate and yielded editing levels consistent with the previous observations (FIGs. 16A). The inventors included two additional gRNAs targeting KRAS and MALAT1 for DECOR to fully benchmark its off-target behavior (FIG.
  • Untransfected HEK293T cells were employed as a control.
  • the inventors called out A-to-I edits using REDItools, a pipeline developed to profile RNA editing in RNA-sequencing data 53 .
  • A-to-I mutations that fit the following criteria are identified as off-target edits: 1) sequencing depth > 10 in samples and untransfected controls; 2) A-to-G frequencies in transfected cells significantly higher than those in untransfected controls (Fisher’s exact test P- value ⁇ 0.05); 3) persisting in two independent replicates.
  • REPAIRv2, REPAIR.tl, and dCasl3d-ADAR DD -E488Q reduced the number of off- target sites to 11-29, 1,992-2,139, and 2,881-2,929 (FIG. 16C), respectively, consistent with their modest on-target activity.
  • the absolute number of off-target sites can vary with different sequencing depths and analysis pipelines. The numbers the inventors present here may not be directly comparable with numbers reported in other studies. Nevertheless, the results are in general agreement with previous reports that with an embedded hyperactive ADAR variant, REPAIRvl shows considerable transcriptome-wide off-target effects in human cells 18 ’ 21 ’ 22 ’ 50 .
  • hfDECOR-F148A and hfDECOR-V106G generated 23.1 ⁇ 10.5% and 24.0 ⁇ 9.2% A-to-I editing across 18 endogenous RNAs, respectively, slightly lower than what was observed for DECOR (35.4 ⁇ 12.1%; FIG. 4A and FIG. 18A).
  • hfDECOR When complexed with the non-targeting gRNA, hfDECOR showed drastically reduced off-target effects — 47 and 126 off-target sites for hfDECOR-F148A and hfDECOR- V106G, respectively (FIG. 4B, FIG. 18B and 18C), 94.7% and 85.8% lower than DECOR.
  • the median editing levels observed at these off-target sites were 17.0% for hfDECOR-F148A and 19.5% for hfDECOR-V106G (FIG. 4C and FIG. 18D).
  • RNA-editing properties of DECOR characterized, the inventors next attempted to leverage DECOR to recode GFP mRNA that harbors a premature stop codon (FIG. 5A).
  • GFP-W58X was overexpressed in HEK293T cells with RFP serving as a transfection control.
  • A-to-I editing mediated by DECOR changes the UAG codon to UIG, which is read as UGG by the ribosome for tryptophan (W) installation, restoring GFP fluorescence.
  • gRNA b that binds 11 to 40 nt upstream of the UAG codon led to strong recovery of GFP fluorescence (FIG. 5B).
  • Minimal GFP signals were detected in cells expressing non-targeting DECOR.
  • the editing level at the UAG codon was quantified by high- throughput sequencing to be 7.0% (FIG. 5C), confirming that the observed GFP restoration was indeed a consequence of programmed RNA editing.
  • VWS Van der Woude syndrome
  • A-48T mutation that creates a de novo start codon in the 5’ UTR of IRF6.
  • the new start codon cannot produce IRF6 protein owing to a shifted coding frame. Instead, translation of the uORF suppresses initiation at the native start site, leading to insufficient IRF6 expression (FIG.
  • RNA-editing technology including DECOR
  • DECOR can reverse an A-to-U transversion.
  • the inventors propose to edit A-49 to mute the uORF, with which the disease-causing start codon AUG will be converted into IUG, a much less potent translation initiation signal.
  • the inventors adapted a customized dual-luciferase reporter assay in which the IRF65’ UTR was employed to drive expression of firefly luciferase (FIG. 5D) 63 .
  • the inventors quantified the editing level at A-49 again by high-throughput sequencing and detected 10.2% A-to-I editing (FIG. 5g). While the diseased UTR confers a pORF expression strength 12.3% of the healthy UTR, RNA editing boosts the expression level to 36.5%, an impact that can likely translate in the clinic.
  • HSV herpes simplex virus
  • TK thymidine kinase
  • DECOR a new CRISPR-based RNA- editing platform.
  • DECOR employs an evolved bacterial tRNA adenosine deaminase to edit As flagged by CRISPR.
  • the inventors demonstrate that DECOR functions broadly in human cells, delivering A-to-I editing at levels comparable to ADAR-mediated RNA editing.
  • DECOR shows significantly lower transcriptome-wide off-target effects than RNA-editing systems with overexpressed ADARs, and an orthogonal off-target profile. High-fidelity DECOR further reduces the off-target effects close to basal level.
  • TadA has recently been engineered to deaminate C instead of A 68 ' 70 , with which DECOR can readily expand its territory into C-to-U editing.
  • deaminase evolution has primarily focused on DNA-editing applications 9 ' 11, 13, 36
  • similar strategies can be employed to isolate deaminases tailored for RNA editing 71 .
  • RNA editing is particularly diverse and abundant in symbiotic dinoflagellates 72 .
  • RNA editing occurs in 1.6% of all nuclear-encoded genes of Symbiodinium microadriaticum, with all 12 types of base substitutions observed 73 .
  • the involved nucleobase conversion enzymes although poorly characterized at the moment, may offer an ample pool of effector proteins to further expand the scope of programmed RNA editing.
  • DECOR has a relaxed editing window with the most strongly edited A jointly defined by the position of the gRNA, the context preference of TadA8e, and possibly intrinsic properties of the target mRNA. With all tested sites considered, the optimal distance between the target A and the gRNA binding site is 13 ⁇ 10 nt. As a start point for design, the inventors recommend to place the target A 1-30 nt downstream of the gRNA-binding site and prioritize YA targets. gRNA screening can be carried out to secure an optimal editing profile, as demonstrated by correcting a premature stop codon in GFP.
  • DECOR may leverage these context-specific deaminases to achieve clean and precise editing patterns.
  • DECOR may demonstrate its ultimate clinical utility for editing transcripts that function transiently in development. For example, temporary upregulation of IRF6, as enabled by DECOR-mediated editing of IRF6-A- 9 at the early stage of embryonic development, may be sufficient to prevent VWS without changing the patient’s genome.
  • DECOR is a new programmable RNA-editing platform with features distinct from and complementary to existing RNA-editing technologies.
  • Oligonucleotides were ordered from Integrated DNA Technologies (IDT). Sources of plasmid and synthetic DNA sequences are listed in Table 3. DNA sequences of signaling peptides, linker peptides, and //U6 5’ UTR are listed in Table 4. PCR fragments were generated using Phusion U DNA polymerase (Thermo Scientific, F555L) and assembled by USER assembly (New England Biolabs, M5505L) following manufactures’ protocols. DH5-alpha competent cells (New England Biolabs, C2992I) were used for cloning. The sequences of all plasmids were confirmed by Sanger sequencing (Azenta Life Science).
  • Cas 13 -deaminase fusions and GFP/RFP reporters are expressed under a cytomegalovirus (CMV) promotor.
  • Firefly luciferase reporters are expressed with a HSV-TK promoter or the native IRF6 promoter.
  • Guide RNAs are expressed with a U6 promoter. Key plasmids generated in this work will be deposited to Addgene.
  • HEK293T, HeLa, A549 cell lines and IMR-90 human primary fibroblasts were cultured in Dulbecco’s Modified Eagle Medium with high glucose, sodium pyruvate, and L- Glutamine (Coming) supplemented with 10% fetal bovine serum (Coming). Cells were maintained in 5% CO2 and passaged regularly at confluence of 80%.
  • RNA editing in IMR-90 human primary fibroblasts DNA was delivered by nucleofection (Lonza, V4XC-1032) with a 4D-Nucleofector. Briefly, 2 x 10 5 cells were transfected with program CM- 120 in a 20 pL Nucleocuvette containing 1000 ng of plasmid DNA expressing both the editor and the gRNA. Cells were resuspended in 500 pL of DMEM and incubated in a 24-well plate for an additional 28 h. Total RNA was extracted using a Direct- zol RNA Microprep Kit (Zymo Research) followed by reverse transcription and deep sequencing as described above.
  • HEK293T cells were transfected with jetOPTIMUS DNA transfection reagent.
  • Total RNA was extracted with TRIzol (Invitrogen, 15596026) followed by DNase I (Lucigen, QER09015) digestion. Reverse transcription, cDNA amplification, and high-throughput sequencing followed the same protocol as described above.
  • HEK293T cells were transfected with 10 ng of GFP-W58X plasmid, 10 ng of RFP plasmid, 30 ng of editor plasmid, and 60 ng of gRNA plasmid using 0.12 pL of jetOPTIMUS DNA transfection reagent in 12.5 pL of jetOPTIMUS buffer. After 28 h of incubation, cells were imaged using an Olympus BX53 microscope. Total RNA was extracted from the same cells with TRIzol followed by DNase I digestion. RNA editing was quantified by high- throughput sequencing as described above. Editing 0URF65’ UTR
  • RNA off target editing experiments were carried out in HEK293T cells.
  • 200 ng of editor plasmid and 400 ng of gRNA plasmid were delivered with 0.6 pL of jetOPTIMUS DNA transfection reagent in 62.5 pL of jetOPTIMUS buffer.
  • Cells were incubated for 28 h followed by total RNA extraction with TRIzol (Invitrogen).
  • mRNA was enriched by a NEBNext Poly(A) mRNA Magnetic Isolation Module (New England Biolabs, E7490L) and prepared for sequencing using an NEBNext Ultra II Directional RNA Library Prep Kit for Illumina (New England Biolabs, E7760L).
  • the libraries were sequenced on an Illumina NovaSeq instrument (single-end R1 121) or a NextSeq 2000 instrument (paired-end R1 25 R2 97 or paired-end R1 62 R2 160).
  • MiSeq reads were aligned to reference sequences with BWA-MEM 74 and sorted using Samtools 1.13 75 . Editing rates at given positions were calculated as G-reads/total-reads.
  • quality control was conducted using Fastqc 0.11.5 (http://www.bioinformatics.babraham.ac.uk/projects/fastqc/).
  • Fastqc 0.11.5 http://www.bioinformatics.babraham.ac.uk/projects/fastqc/.
  • R2 reads of paired-end sequencing data were used for analysis. All reads were trimmed to the same length (97 nt). Reads were aligned to the hgl9 (GRCh37) reference genome using STAR 2.6.1b 76 .
  • Sites with a P-value greater than 0.05 were removed as these are A sites edited to similar levels regardless of the treatment, i.e., sites hosting naturally occurring A-to-I editing. Off-target sites were considered significant if detected in both replicates. Mean editing rates of two replicates were used for all plots involving editing frequencies. Overlap of off-target sites among different samples was analyzed by BioVenn 77 . Sequence motif analysis was done using WebLogo 3 78 .
  • RNA-seq data have been deposited to the NCBI Gene Expression Omnibus (GEO) and can be accessed through GEO series accession number GSE251875.
  • GEO Gene Expression Omnibus
  • Example 3 - DECOR favors A proceeded by a pyrimidine
  • E. coli TadA The physiological function of E. coli TadA is to deaminate the A at the wobble position of Arg tRNA(ACG) 79 . Proceeded by a U-turn, this A is situated in a UAC motif. Both -1 U and +1 C interact with TadA during substrate engagement 80 . The tertiary structure of the anti-codon loop was proposed both essential and sufficient for TadA-catalyzed adenosine deamination 81 .
  • N A, C, G, and U
  • GFP GFP mRNA
  • FIG. 14A A, C, G, and U
  • A-to-I editing in NAN was quantified by RT-PCR followed by NGS.
  • UAN displayed the highest level of editing (6.7- 10.6%, FIGs. 14B and 14C).
  • the second most favored sequence category appeared to be CAN (2.2-3.6%).
  • Example 4 - DECOR has a tightly controlled off-target profile in the human genome
  • the dSaCas9 complex forms a constant R loop susceptible to TadA8e- mediated deamination; editing at the R loop therefore serves as a surrogate for potential genome-wide off-target editing (FIGs. 17A-17J).
  • Both DECOR and ABE8e showed anticipated on-target RNA and DNA editing, respectively, in the presence of these orthogonal R loops.
  • the lower DNA off-target editing observed with DECOR may stem from laboratory evolution of TadA8e as part of the ABE architecture 87 , which prompts specialization of TadA8e for robust deamination within the ABE context. This optimal DNA deamination is not expected to extend to DECOR, where TadA8e is fused to an RNA-targeting CRISPR protein.
  • Table 1 NCBI accession numbers for target cellular RNAs.

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Organic Chemistry (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Engineering & Computer Science (AREA)
  • Genetics & Genomics (AREA)
  • Wood Science & Technology (AREA)
  • Zoology (AREA)
  • Biochemistry (AREA)
  • Molecular Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • General Engineering & Computer Science (AREA)
  • Medicinal Chemistry (AREA)
  • Microbiology (AREA)
  • Biotechnology (AREA)
  • Biomedical Technology (AREA)
  • Pharmaceuticals Containing Other Organic And Inorganic Compounds (AREA)

Abstract

Aspects herein include RNA-base-editing proteins linked to RNA-targeting proteins, which can allow for targeted editing of an RNA molecule. Aspects herein include systems and compositions comprising the base-editing and targeting proteins along with methods of their use.

Description

POLYPEPTIDES AND METHODS FOR MODIFYING NUCLEIC ACIDS
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority of U.S. Provisional Application No. 63/659,189 filed June 12, 2024 which is hereby incorporated by reference in its entirety.
STATEMENT AS TO FEDERALLY SPONSORED RESEARCH
[0002] This invention was made with government support under EB032491 awarded by the National Institutes of Health. The government has certain rights in the invention.
SEQUENCE LISTING
[0003] The instant application contains a Sequence Listing which has been submitted in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on June 11, 2025, is named ARCD.P0839WO - SEQ LISTING.xml and is 210,928 bytes in size.
BACKGROUND OF THE INVENTION
II. Field of the Invention
[0004] This invention relates to the field of molecular biology, genetic engineering, and medicine.
III. Background
[0005] RNA editing is a natural process widely occurring in eukaryotic organisms and their associated viruses1. Editing events insert, delete, or convert nucleobases, resulting in RNA sequences that differ from the encoding genes. The most common type of RNA editing is adenosine-to-inosine (A-to-I) conversion mediated by adenosine deaminases acting on RNA (ADARs)2. As a means of altering the sequence of RNA and thereby the designated protein, RNA editing holds promise for treating genetic disease3, 4 Compared to DNA editing that recodes a cell’s genome to alter cell fate, RNA editing can change a cell’s behavior without enforcing permanent modifications to its genetic materials. The reversibility and dosedependent features of RNA editing may be particularly desirable for treating disease that demands a defined editing level or therapeutic time window.
[0006] CRISPR-Cas systems, through their RNA-guided targeting mechanisms, grant nucleic acid editing the ultimate programmability5. Many effector proteins are compatible with DNA-targeting CRISPR systems for site-specific base conversion, deletion, and insertion, including different families of cytosine deaminases in cytosine base editors (CBEs)6'8, evolved E. coli tRNA-specific adenosine deaminases (TadA) in adenine base editors (ABEs)9'13, and various reverse transcriptases and DNA polymerases in prime editors (PEs)14, 15 and their related editors16, 17 Extension of CRISPR-based editing to RNA, which still follows the relatively straightforward protein engineering strategy, relies on the identification of catalytically competent, CRISPR-compatible RNA-editing enzymes. Efforts in programmed RNA editing so far center around a single family of effector proteins, ADARs, that act on double-stranded RNA (dsRNA). A collection of CRISPR-dependent and independent approaches have been developed to deliver ADAR-mediated site-specific A-to-I editing to RNA, including REPAIR (RNA Editing for Programmable A to I Replacement)18, 19, RESTORE (Recruiting Endogenous ADAR to Specific Transcripts for Oligonucleotide- mediated RNA Editing)20, LEAPER (Leveraging Endogenous ADAR for Programmable Editing of RNA)21, engineered ADAR guide RNAs (adRNAs)22, RNA editing with individual RNA-binding enzyme (REWIRE)23, CLUSTER guide RNAs24, and circular ADAR-recruiting RNAs25, 26. Alongside A-to-I editing, C-to-U editing can be introduced to RNA in a CRISPR- guided manner by RESCUE (RNA Editing for Specific C-to-U Exchange) using an evolved ADAR protein27 and CURE (C-to-U RNA Editor) with human cytidine deaminase AP0BEC3A (A3 A)28.
[0007] AD ARI, while essential for mouse embryogenesis29, is not highly expressed in most adult tissues in humans (FIG. 6A)30, posing challenges to RNA-editing platforms that function by recruiting endogenous ADARs. Its predominate residence in the nucleus31 also precludes efficient editing in the cytoplasm. ADAR2, another catalytically competent ADAR family protein, is expressed at low levels in most tissues (FIG. 6B)30. While AD AR2 -mediated RNA editing has been harnessed for sensing in mouse brains32, the tissue-specific expression of ADAR2 may limit its application in programmed RNA editing21. On the other hand, ADARs are estimated to impact over 100 million A sites in the human transcriptome33. RNA-editing systems that require exogeneous or overexpressed ADARs may therefore introduce excess A- to-I edits to the cellular RNA pool34. Finally, the inherent sequence preference of ADARs precludes a subset of A sites from efficient editing18. To expand the scope of RNA editing and to reduce the risk of adverse events, programmable RNA-editing platforms with alternative effector proteins are needed. SUMMARY OF THE INVENTION
[0008] Aspects herein relate to the discovery that pairing a tRNA-specific editing protein to an RNA targeting protein can lead to site specific RNA editing. In certain aspects, the tRNA- specific editing protein is a tRNA-specific deaminase. The tRNA-specific editing protein may be a TadA adenosine deaminase. In certain aspects, the tRNA-specific adenosine deaminase is an engineered TadA adenosine deaminase.
[0009] Disclosed herein are polypeptides comprising a tRNA-specific adenosine deaminase protein. Also disclosed are RNA-targeting protein. In certain aspects, the tRNA-specific adenosine deaminase protein is operably linked to the RNA-targeting protein.
[0010] The tRNA-specific adenosine deaminase protein may be an engineered adenosine deaminase (such as an engineered TadA adenosine deaminase), which may be engineered by directed evolution or mutagenesis. In some aspects, the tRNA-specific adenosine deaminase protein is not a wild-type TadA protein. The tRNA-specific adenosine deaminase may be engineered to remove wild-type editing specificity. The tRNA-specific adenosine deaminase may be any adenosine deaminase disclosed herein.
[0011] In certain aspects, the tRNA-specific adenosine deaminase protein is a TadA7.10, TadA8.20, or TadA8e adenosine deaminase. In certain aspects, the tRNA-specific adenosine deaminase protein comprises a F148A and/or V106G substitution in a tadA protein. In some aspects, the tRNA-specific adenosine deaminase protein comprises a functional portion of a TadA adenosine deaminase disclosed herein. In some aspects, the tRNA-specific adenosine deaminase protein comprises an enzymatically active portion of a TadA adenosine deaminase disclosed herein. In certain aspects, the tRNA-specific adenosine deaminase is an engineered bacterial adenosine deaminase. In certain aspects, the tRNA-specific adenosine deaminase is an engineered mammalian adenosine deaminase, such as an engineered ADAT protein. In certain aspects, the tRNA-specific adenosine deaminase is an engineered bacterial, mammalian, plant, or fungal adenosine deaminase. In certain aspects, the tRNA-specific adenosine deaminase comprises an engineered Streptococcus adenosine deaminase. In certain aspects, the tRNA-specific adenosine deaminase comprises an engineered E. coli adenosine deaminase. In certain aspects, the tRNA-specific adenosine deaminase specifically recognizes single-stranded RNA.
[0012] The RNA-targeting protein may be a targeting protein capable of binding to a specifically defined sequence on an RNA. The RNA-targeting protein may be any RNA- targeting protein disclosed herein. In certain aspects, the RNA-targeting protein is a Cas protein. In certain aspects, the RNA-targeting protein is a protospacer-adjacent motif (PAM)- targeting protein. In certain aspects, the RNA-targeting protein is a dead Cas protein. A dead Cas protein includes a Cas protein with one or more point mutations that result in a loss of nucleolytic activity of the Cas protein but, in some aspects, do not impact its binding to a target. In certain aspects, the RNA-targeting protein is a Casl3d protein. In certain aspects, the Casl3d protein is a dead Casl3d (dCasl3d) protein. In certain aspects, the dCasl3d protein is a RfxCasl3d from Ruminococcus flavefaciens XPD3002. In some aspects, the RNA-binding comprises a functional portion of a Cas protein disclosed herein.
[0013] In some aspects, the tRNA-specific adenosine deaminase protein is operably linked to the RNA-targeting protein by a flexible linker. The flexible linker may comprise any flexible linker disclosed herein. In some aspects, the flexible linker is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 in length, or any range derivable therein. In some aspects, the flexible linker is at least or at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 in length, or any range derivable therein. In certain aspects, the tRNA-specific adenosine deaminase protein is operably linked to the N-terminus of the RNA-targeting protein. In certain aspects, the tRNA-specific adenosine deaminase protein is operably linked to the C-terminus of the RNA-targeting protein. In some aspects, the polypeptide comprises at least one nuclear localization signal sequence.
[0014] Also disclosed are nucleic acids encoding any of the polypeptides disclosed herein. In some aspects, the nucleic acid encodes a short guide RNA (sgRNA). In certain aspects, the sgRNA binds to a target RNA. In certain aspects, the target RNA is a disease-causing RNA. In certain aspects, the target RNA comprises an adenosine associated with a disease. In certain aspects, the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of the adenosine associated with the disease. In certain aspects, the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of an adenosine of interest.
[0015] Also disclosed are expression vectors encoding any of the nucleic acids disclosed herein. The expression vector may also encode a promoter. The promoter may be a promoter for a target RNA, including any target RNA disclosed herein. The promoter may be a promoter for a disease-associated RNA. [0016] Also disclosed are cells comprising any of the polypeptides disclosed herein, any of the nucleic acids disclosed herein, and/or any of the expression vectors disclosed herein. In some aspects, the cell comprises an RNA having an adenosine associated with a disease. In some aspects, the cell is a diseased cell.
[0017] Also disclosed are systems comprising a polypeptide comprising a tRNA-specific adenosine deaminase protein operably linked to an RNA-targeting protein, and a short guide RNA. Also disclosed are systems comprising any polypeptide disclosed herein and any sgRNA disclosed herein.
[0018] Also disclosed are methods for modifying one or more adenosine bases, methods for editing one or more adenosine bases, and methods for altering a disease-associated mutation in a nucleic acid molecule. The method can comprise one or more steps including contacting the nucleic acid molecule with any of the polypeptides disclosed herein or contacting the nucleic acid molecule with any of the systems disclosed herein. In certain aspects, at least one of the adenosine bases are associated with a disease. In certain aspects, the method comprises contacting the nucleic acid molecule with a short guide RNA (sgRNA), including any sgRNA disclosed herein. In certain aspects, the contacting results in a conversion of at least one of the adenosine bases in the nucleic acid molecule to an inosine base. In certain aspects, the conversion is at least 20% efficient. In certain aspects, at least one additional adenosine base is not converted to an inosine.
[0019] Also disclosed are methods of treating and/or preventing a disease in a patient. In certain aspects, the method comprises one or more steps including administering to the patient an effective amount of any of the polypeptides disclosed herein, any of the nucleic acids disclosed herein, any of the expression vectors disclosed herein, any of the cells disclosed herein, or any of the systems disclosed herein. In certain aspects, the patient has at least one cell comprising a disease-associated adenosine in an RNA. In certain aspects, the polypeptide converts the adenosine to an inosine. In certain aspects, the polypeptide converts at least 20% of the disease-associated adenosines in the cell to an inosine. In certain aspects, the disease- associated adenosine is an /A7A-A-48T mutation. In certain aspects, the patient has, is suspected of having, or has been diagnosed with having a disease associated with a mutation to an adenosine in an RNA. In certain aspects, the patient has, is suspected of having, or has been diagnosed with having Van der Woude syndrome. [0020] Disclosed herein are polypeptides comprising a tRNA-specific cytosine deaminase protein. Also disclosed are RNA-targeting protein. In certain aspects, the tRNA-specific cytosine deaminase protein is operably linked to the RNA-targeting protein.
[0021] The tRNA-specific cytosine deaminase protein may be engineered by directed evolution or mutagenesis. In some aspects, the tRNA-specific cytosine deaminase protein is not a wild-type TadA protein. The tRNA-specific cytosine deaminase may be engineered to remove wild-type editing specificity. The tRNA-specific cytosine deaminase may be any cytosine deaminase disclosed herein.
[0022] In certain aspects, the tRNA-specific cytosine deaminase protein is a TadA.CBE3.1, TadA.CBE2.4, TadA.CBE.DLL, TadA.CBE.DL, TadA.CBE1.46, or TadA.CBE1.52 cytosine deaminase. In certain aspects, the tRNA-specific cytosine deaminase protein comprises at least one mutation from a wild-type TadA protein. In some aspects, the tRNA-specific cytosine deaminase protein comprises a functional portion of a TadA deaminase disclosed herein. In some aspects, the tRNA-specific cytosine deaminase protein comprises an enzymatically active portion of a TadA deaminase disclosed herein. In certain aspects, the tRNA-specific cytosine deaminase is an engineered bacterial deaminase. In certain aspects, the tRNA-specific cytosine deaminase is an engineered mammalian deaminase, such as an engineered ADAT protein. In certain aspects, the tRNA-specific cytosine deaminase is an engineered bacterial, mammalian, plant, or fungal deaminase. In certain aspects, the tRNA-specific cytosine deaminase comprises an engineered Streptococcus deaminase. In certain aspects, the tRNA-specific cytosine deaminase comprises an engineered E. coli deaminase. In certain aspects, the tRNA-specific cytosine deaminase specifically recognizes single-stranded RNA.
[0023] The RNA-targeting protein may be a targeting protein capable of binding to a specifically defined sequence on an RNA. The RNA-targeting protein may be any RNA- targeting protein disclosed herein. In certain aspects, the RNA-targeting protein is a Cas protein. In certain aspects, the RNA-targeting protein is a protospacer-adjacent motif (PAM)- targeting protein. In certain aspects, the RNA-targeting protein is a dead Cas protein. A dead Cas protein includes a Cas protein with one or more point mutations that result in a loss of nucleolytic activity of the Cas protein but, in some aspects, do not impact its binding to a target. In certain aspects, the RNA-targeting protein is a Casl3d protein. In certain aspects, the Casl3d protein is a dead Casl3d (dCasl3d) protein. In certain aspects, the dCasl3d protein is a RfxCasl3d from Ruminococcus flavefaciens XPD3002. In some aspects, the RNA-binding comprises a functional portion of a Cas protein disclosed herein. In some aspects, the tRNA- specific cytosine deaminase protein comprises an enzymatically active portion of a TadA cytosine deaminase disclosed herein.
[0024] In some aspects, the tRNA-specific cytosine deaminase protein is operably linked to the RNA-targeting protein by a flexible linker. The flexible linker may comprise any flexible linker disclosed herein. In some aspects, the flexible linker is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 in length, or any range derivable therein. In some aspects, the flexible linker is at least or at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 in length, or any range derivable therein. In certain aspects, the TadA cytosine deaminase protein is operably linked to the N-terminus of the RNA-targeting protein. In certain aspects, the TadA cytosine deaminase protein is operably linked to the C-terminus of the RNA- targeting protein. In some aspects, the polypeptide comprises at least one nuclear localization signal sequence.
[0025] Also disclosed are nucleic acids encoding any of the polypeptides disclosed herein. In some aspects, the nucleic acid encodes a short guide RNA (sgRNA). In certain aspects, the sgRNA binds to a target RNA. In certain aspects, the target RNA is a disease-causing RNA. In certain aspects, the target RNA comprises an cytosine associated with a disease. In certain aspects, the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of the cytosine associated with the disease. In certain aspects, the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of an cytosine of interest.
[0026] Also disclosed are expression vectors encoding any of the nucleic acids disclosed herein. The expression vector may also encode a promoter. The promoter may be a promoter for a target RNA, including any target RNA disclosed herein. The promoter may be a promoter for a disease-associated RNA.
[0027] Also disclosed are cells comprising any of the polypeptides disclosed herein, any of the nucleic acids disclosed herein, and/or any of the expression vectors disclosed herein. In some aspects, the cell comprises an RNA having an cytosine associated with a disease. In some aspects, the cell is a diseased cell.
[0028] Also disclosed are systems comprising a polypeptide comprising a TadA cytosine deaminase protein operably linked to an RNA-targeting protein, and a short guide RNA. Also disclosed are systems comprising any polypeptide disclosed herein and any sgRNA disclosed herein.
[0029] Also disclosed are methods for modifying one or more cytosine bases, methods for editing one or more cytosine bases, and methods for altering a disease-associated mutation in a nucleic acid molecule. The method can comprise one or more steps including contacting the nucleic acid molecule with any of the polypeptides disclosed herein or contacting the nucleic acid molecule with any of the systems disclosed herein. In certain aspects, at least one of the cytosine bases are associated with a disease. In certain aspects, the method comprises contacting the nucleic acid molecule with a short guide RNA (sgRNA), including any sgRNA disclosed herein. In certain aspects, the contacting results in a conversion of at least one of the cytosine bases in the nucleic acid molecule to an uridine base. In certain aspects, the conversion is at least 20% efficient. In certain aspects, at least one additional cytosine base is not converted to an uridine.
[0030] Also disclosed are methods of treating and/or preventing a disease in a patient. In certain aspects, the method comprises one or more steps including administering to the patient an effective amount of any of the polypeptides disclosed herein, any of the nucleic acids disclosed herein, any of the expression vectors disclosed herein, any of the cells disclosed herein, or any of the systems disclosed herein. In certain aspects, the patient has at least one cell comprising a disease-associated cytosine in an RNA. In certain aspects, the polypeptide converts the cytosine to an uridine. In certain aspects, the polypeptide converts at least 20% of the disease-associated cytosines in the cell to an uridine. In certain aspects, the patient has, is suspected of having, or has been diagnosed with having a disease associated with a mutation to an cytosine in an RNA.
[0031] Also disclosed are the following enumerated aspects:
Aspect 1. A polypeptide comprising a tRNA-specific adenosine deaminase protein operably linked to an RNA-targeting protein
Aspect 2. The polypeptide of aspect 1, wherein the tRNA-specific adenosine deaminase protein is a TadA7.10, TadA8.20, or TadA8e adenosine deaminase
Aspect 3. The polypeptide of aspect 1 or 2, wherein the tRNA-specific adenosine deaminase protein comprises an F148A and/or V106G substitution
Aspect 4. The polypeptide of any one of aspects 1 to 3, wherein the RNA-targeting protein is a Cast 3d protein Aspect 5. The polypeptide of aspect 4, wherein the Casl3d protein is a dead Casl3d (dCasl3d) protein
Aspect 6. The polypeptide of aspect 5, wherein the dCasl3d protein is an RfxCasl3d from *Ruminococcus flavefaciens* XPD3002
Aspect 7. The polypeptide of any one of aspects 1 to 6, wherein the tRNA-specific adenosine deaminase protein is operably linked to the RNA-targeting protein by a flexible linker
Aspect 8. The polypeptide of aspect 7, wherein the flexible linker is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 amino acids in length
Aspect 9. The polypeptide of any one of aspects 1 to 8, wherein the tRNA-specific adenosine deaminase protein is operably linked to the N-terminus of the RNA-targeting protein
Aspect 10. The polypeptide of any one of aspects 1 to 8, wherein the tRNA-specific adenosine deaminase protein is operably linked to the C-terminus of the RNA-targeting protein
Aspect 11. The polypeptide of any one of aspects 1 to 10, further comprising at least one nuclear localization signal sequence
Aspect 12. A nucleic acid encoding the polypeptide of any one of aspects 1 to 11
Aspect 13. The nucleic acid of aspect 12, further encoding a short guide RNA (sgRNA)
Aspect 14. The nucleic acid of aspect 13, wherein the sgRNA binds to a target RNA
Aspect 15. The nucleic acid of aspect 14, wherein the target RNA is a disease-causing RNA
Aspect 16. The nucleic acid of aspect 14 or 15, wherein the target RNA comprises an adenosine associated with a disease
Aspect 17. The nucleic acid of aspect 16, wherein the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of the adenosine associated with the disease
Aspect 18. An expression vector encoding the nucleic acid of any one of aspects 12 to 16
Aspect 19. A cell comprising the polypeptide of any one of aspects 1 to 11, the nucleic acid of any one of aspects 12 to 17, or the expression vector of aspect 18
Aspect 20. The cell of aspect 19, comprising an RNA having an adenosine associated with a disease Aspect 21. A system comprising a polypeptide comprising a tRNA-specific adenosine deaminase protein operably linked to an RNA-targeting protein; and a short guide RNA
Aspect 22. The system of aspect 21, wherein the polypeptide is the polypeptide of any one of aspects 1 to 11
Aspect 23. The system of aspect 21 or 22, wherein the short guide RNA binds to a target RNA
Aspect 24. The system of aspect 23, wherein the target RNA is a disease-causing RNA
Aspect 25. The system of aspect 23 or 24, wherein the target RNA comprises an adenosine associated with a disease
Aspect 26. A method for modifying one or more adenosine bases and/or for editing one or more adenosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the polypeptide of any one of aspects 1 to 11
Aspect 27. The method of aspect 26, wherein at least one of the adenosine bases is associated with a disease
Aspect 28. The method of aspect 26 or 27, further comprising contacting the nucleic acid molecule with a short guide RNA (sgRNA)
Aspect 29. The method of aspect 28, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one adenosine associated with the disease
Aspect 30. The method of any one of aspects 26 to 29, wherein the contacting results in a conversion of at least one of the adenosine bases in the nucleic acid molecule to an inosine base
Aspect 31. The method of aspect 30, wherein the conversion is at least 20 % efficient
Aspect 32. The method of any one of aspects 26 to 31, wherein at least one additional adenosine base is not converted to an inosine
Aspect 33. A method for modifying adenosine bases and/or for editing adenosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the system of any one of aspects 21 to 25
Aspect 34. The method of aspect 33, wherein at least one of the adenosine bases is associated with a disease Aspect 35. The method of aspect 34, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one adenosine associated with the disease
Aspect 36. The method of any one of aspects 33 to 35, wherein the contacting results in a conversion of at least one of the adenosine bases in the nucleic acid molecule to an inosine base
Aspect 37. The method of aspect 36, wherein the conversion is at least 20 % efficient
Aspect 38. The method of any one of aspects 33 to 37, wherein at least one additional adenosine base is not converted to an inosine
Aspect 39. A method of treating a disease in a patient, the method comprising administering to the patient an effective amount of the polypeptide of any one of aspects 1 to 11, the nucleic acid of any one of aspects 12 to 17, the expression vector of aspect 18, the cell of aspect 19 or 20, or the system of any one of aspects 21 to 25
Aspect 40. The method of aspect 39, wherein the patient has at least one cell comprising a disease-associated adenosine in an RNA
Aspect 41. The method of aspect 40, wherein the polypeptide converts the adenosine to an inosine
Aspect 42. The method of aspect 40, wherein the polypeptide converts at least 20 % of the disease-associated adenosines in the cell to inosine
Aspect 43. The method of any one of aspects 40 to 42, wherein the disease-associated adenosine is an IRF6-A-48T mutation
Aspect 44. The method of any one of aspects 40 to 43, wherein the patient has, is suspected of having, or has been diagnosed with a disease associated with a mutation to an adenosine in an RNA
Aspect 45. The method of any one of aspects 40 to 44, wherein the patient has, is suspected of having, or has been diagnosed with Van der Woude syndrome
Aspect 46. A polypeptide comprising a tRNA-specific cytosine deaminase protein operably linked to an RNA-targeting protein Aspect 47. The polypeptide of aspect 46, wherein the tRNA-specific cytosine deaminase protein is a TadA.CBE3.1, TadA.CBE2.4, TadA.CBE.DLL, TadA.CBE.DL, TadA.CBE1.46, or TadA.CBE1.52 cytosine deaminase
Aspect 48. The polypeptide of aspect 46 or 47, wherein the tRNA-specific cytosine deaminase protein has at least one mutation from a TadA protein
Aspect 49. The polypeptide of any one of aspects 46 to 48, wherein the RNA-targeting protein is a Casl3d protein
Aspect 50. The polypeptide of aspect 49, wherein the Casl3d protein is a dead Casl3d (dCasl3d) protein
Aspect 51. The polypeptide of aspect 50, wherein the dCasl3d protein is an RfxCasl3d from *Ruminococcus flavefaciens* XPD3002
Aspect 52. The polypeptide of any one of aspects 46 to 51, wherein the TadA cytosine deaminase protein is operably linked to the RNA-targeting protein by a flexible linker
Aspect 53. The polypeptide of aspect 52, wherein the flexible linker is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 amino acids in length
Aspect 54. The polypeptide of any one of aspects 46 to 53, wherein the TadA cytosine deaminase protein is operably linked to the N-terminus of the RNA-targeting protein
Aspect 55. The polypeptide of any one of aspects 46 to 53, wherein the TadA cytosine deaminase protein is operably linked to the C-terminus of the RNA-targeting protein
Aspect 56. The polypeptide of any one of aspects 46 to 55, further comprising at least one nuclear localization signal sequence
Aspect 57. A nucleic acid encoding the polypeptide of any one of aspects 46 to 56
Aspect 58. The nucleic acid of aspect 57, further encoding a short guide RNA (sgRNA)
Aspect 59. The nucleic acid of aspect 58, wherein the sgRNA binds to a target RNA
Aspect 60. The nucleic acid of aspect 59, wherein the target RNA is a disease-causing RNA
Aspect 61. The nucleic acid of aspect 59 or 60, wherein the target RNA comprises a cytosine associated with a disease Aspect 62. The nucleic acid of aspect 61, wherein the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of the cytosine associated with the disease
Aspect 63. An expression vector encoding the nucleic acid of any one of aspects 57 to 61
Aspect 64. A cell comprising the polypeptide of any one of aspects 46 to 56, the nucleic acid of any one of aspects 12 to 17, or the expression vector of aspect 63
Aspect 65. The cell of aspect 64, comprising an RNA having a cytosine associated with a disease
Aspect 66. A system comprising a polypeptide comprising a tRNA-specific cytosine deaminase protein operably linked to an RNA-targeting protein; and a short guide RNA
Aspect 67. The system of aspect 66, wherein the polypeptide is the polypeptide of any one of aspects 1 to 11
Aspect 68. The system of aspect 66 or 67, wherein the short guide RNA binds to a target RNA
Aspect 69. The system of aspect 68, wherein the target RNA is a disease-causing RNA
Aspect 70. The system of aspect 68 or 69, wherein the target RNA comprises a cytosine associated with a disease
Aspect 71. A method for modifying one or more cytosine bases and/or for editing one or more cytosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the polypeptide of any one of aspects 46 to 56
Aspect 72. The method of aspect 71, wherein at least one of the cytosine bases is associated with a disease
Aspect 73. The method of aspect 71 or 72, further comprising contacting the nucleic acid molecule with a short guide RNA (sgRNA)
Aspect 74. The method of aspect 73, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one cytosine associated with the disease
Aspect 75. The method of any one of aspects 71 to 74, wherein the contacting results in a conversion of at least one of the cytosine bases in the nucleic acid molecule to a uridine base Aspect 76. The method of aspect 75, wherein the conversion is at least 15 % efficient
Aspect 77. The method of any one of aspects 71 to 76, wherein at least one additional cytosine base is not converted to a uridine
Aspect 78. A method for modifying cytosine bases and/or for editing cytosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the system of any one of aspects 66 to 70
Aspect 79. The method of aspect 78, wherein at least one of the cytosine bases is associated with a disease
Aspect 80. The method of aspect 79, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one cytosine associated with the disease
Aspect 81. The method of any one of aspects 78 to 80, wherein the contacting results in a conversion of at least one of the cytosine bases in the nucleic acid molecule to a uridine base
Aspect 82. The method of aspect 81, wherein the conversion is at least 15 % efficient
Aspect 83. The method of any one of aspects 78 to 82, wherein at least one additional cytosine base is not converted to a uridine
Aspect 84. A method of treating a disease in a patient, the method comprising administering to the patient an effective amount of the polypeptide of any one of aspects 1 to 11, the nucleic acid of any one of aspects 12 to 17, the expression vector of aspect 63, the cell of aspect 19 or 20, or the system of any one of aspects 21 to 25
Aspect 85. The method of aspect 84, wherein the patient has at least one cell comprising a disease-associated cytosine in an RNA
Aspect 86. The method of aspect 85, wherein the polypeptide converts the cytosine to a uridine
Aspect 87. The method of aspect 85, wherein the polypeptide converts at least 15 % of the disease-associated cytosines in the cell to uridine
Aspect 88. The method of any one of aspects 85 to 87, wherein the patient has, is suspected of having, or has been diagnosed with a disease associated with a mutation to a cytosine in an RNA. Aspect 89. A polypeptide comprising 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any range derivable therein) identity to any one of SEQ ID NOs: l-29.
[0032] It is specifically contemplated that an aspect described herein can be combined, substituted with, or used with one or more different aspects described herein.
[0033] Throughout this application, the term “about” is used to indicate that a value includes the inherent variation of error for the measurement or quantitation method.
[0034] The use of the word “a” or “an” when used in conjunction with the term “comprising” may mean “one,” but it is also consistent with the meaning of “one or more,” “at least one,” and “one or more than one.”
[0035] The phrase “and/or” means “and” or “or”. To illustrate, A, B, and/or C includes: A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B, and C. In other words, “and/or” operates as an inclusive or.
[0036] The words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”) or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
[0037] The compositions and methods for their use can “comprise,” “consist essentially of,” or “consist of’ any of the ingredients or steps disclosed throughout the specification. Compositions and methods “consisting essentially of’ any of the ingredients or steps disclosed limits the scope of the claim to the specified materials or steps which do not materially affect the basic and novel characteristic of the claimed invention.
[0038] It is contemplated that any embodiment discussed in this specification can be implemented with respect to any method or composition of the invention, and vice versa. Furthermore, compositions of the invention can be used to achieve methods of the invention.
[0039] Other obj ects, features and advantages of the present invention will become apparent from the following detailed description. It should be understood, however, that the detailed description and the specific examples, while indicating specific embodiments of the invention, are given by way of illustration only, since various changes and modifications within the spirit and scope of the invention will become apparent to those skilled in the art from this detailed description.
BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The follow drawings provided in the Appendix form part of the present specification and are included to further demonstrate certain aspects of the present invention. The invention may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.
[0041] FIGs. 1A-1H show programmed RNA editing by dCas!3-deaminase fusions, a. DECOR-mediated RNA editing. DECOR constitutes two functional modules — an adenosine deaminase TadA8e and a dCasl3d complex. Guided by the dCasl3d gRNA, DECOR recognizes target RNA and deaminates specified adenosine to inosine. Inosine is recognized as guanosine by polymerases and ribosomes, serving as a functional equivalent of guanosine in the coding and decoding process, b. Architectures of dCas 13 -deaminase fusion proteins, c. Schematic of EZH2 mRNA showing the gRNA binding site and putative editing sites, d. C-to- U editing at EZ//2-C I I 66 and Cl 195 enforced by fusions of dCasl3d and cytosine deaminases, e. A-to-I editing at EZEE2-KW19 mediated by dCasl3d-TadA fusions, f. Sanger traces for DECOR-edited EZH 2- E W 7 . g. TadA8e-enabled A-to-I editing at EZH2 -Al 179 guided by LwaCasl3a, PspCasl3b, Casl3X. l, mini-Casl3X. l and RfxCasl3d. h. A-to-I editing across all A sites in EZH 2 -1009-1258 deposited by CRISPR-guided TadA8e. The gRNA binding site is annotated in the heatmap. Panels g and h are derived from the same set of data. All fusion proteins were evaluated in HEK293T cells with EZH2 -targeting and non-targeting (NT) gRNA (mean ± s.d., n = 3).
[0042] FIGs. 2A-2G show DECOR-mediated editing of coding and noncoding RNA. a. Design and placement of DECOR, REPAIRvl, and REPAIRv2 gRNAs. b. A-to-I editing mediated by DECOR and REPAIRvl at designated As in seven cellular RNAs. c. Heatmaps of all A edits observed in seven RNAs. The target As are aligned across seven RNAs as position 0. gRNA binding regions are boxed in grey rectangles. Editing rates are normalized to the highest editing level observed in a given RNA. For a zoomed-in view of editing distributions, see FIGs. 12A-12D. d. Box and whisker plot of editing rates delivered by different editors and gRNA designs. Design and placement of REPAIR.tl, ADAR-dCasl3d, and dCasl3d-ADAR gRNAs are provided in FIG. 15 A. REPAIR-max represents the maximum editing level observed at a given A among all REPAIR gRNAs. REPAIR-all includes all REPAIR gRNA designs. The solid line within the box represents the median, while upper and lower hinges depict the first and third quartiles. Whiskers extend from the minimum to the maximum. The plus sign (+) denotes the mean, n = 7 except REPAIRvl-all (n = 28), REPAIRv2-all (n = 28), and REPAIR.tl-all (n = 21). Statistical analysis was carried out using a two-tailed Student's t- test. n.s.: not significant, e-g. Programmed RNA editing by DECOR in HeLa (e), A549 (f) and IMR-90 human primary fibroblast (g) cells. RNA-editing experiments were carried out in HEK293T (a-d), HeLa (e), A549 (f), and IMR-90 (g) cells in the presence of targeting or nontargeting gRNA (mean ± s.d., n = 3).
[0043] FIGs. 3A-3G show transcriptome-wide specificity of DECOR, a. Transcriptomewide A-to-I editing by DECOR and REPAIRvl paired with EZH2 -targeting and non-targeting gRNA. b. Violin plots of A-to-I conversion rates observed at off-target sites. Solid lines represent the median, whereas dashed lines denote the first and third quantiles. Statistical analysis was done using a two-tailed Student's t-test. n.s. P > 0.05; ** P < 0.01; **** P < 0.0001. c. Correlation analysis of editing rates at off-target sites shared by non-targeting and /;Z7/2-targeting DECOR, d. Venn diagram showing the overlap of off-target sites between DECOR and REPAIRvl. e. Heatmap showing the overlap rates of off-target sites observed for DECOR and REPAIRvl. The percentage of overlap was calculated using the intersect number and the total number of off-target sites observed for a given RNA editing system. To simplify the plot, the inventors only present the percentage calculated based on one of the two editing systems affording fewer off-target sites. For example, the overlap of DECOR NT and REPAIRvl NT is shown as 10/888 = 1.1% instead of 10/4044 = 0.2%. f, g. Sequence logos hosting all A-to-I edits observed in the transcriptome for non-targeting DECOR and REPAIRvl. All experiments were carried out in duplicate in HEK293T cells.
[0044] FIGs. 4A-4F show efficiency and specificity of hfDECOR. a. A-to-I editing mediated by hfDECOR at designated As in eighteen cellular RNAs (mean ± s.d., n = 3). b. Transcriptome-wide A-to-I editing by hfDECOR-F148A. c. Violin plots of off-target editing rates for hfDECOR-F148A. Solid lines represent the median, whereas dashed lines denote the first and third quantiles. Statistical analysis was done using a two-tailed Student's t-test. n.s. P > 0.05. d. Venn diagram showing off-target sites shared by DECOR and hfDECOR-F148A. e. Venn diagram showing the overlap of off-target sites between REPAIRvl and hfDECOR- F148A. f. Sequence logos of A sites impacted by hfDECOR-F148A in the HEK293T transcriptome. All overlap analyses present here use data obtained with non-targeting gRNA. the inventors observe similar overlapping rates with /;Z7/2-targeting gRNA. All off-target profiling experiments were carried out in duplicate in HEK293T cells. [0045] FIGs. 5A-5I show removal of a disease-causing upstream open reading frame (uORF) and rescue of protein expression by DECOR, a. Schematic of RNA editing in the coding region of GFP carrying a premature stop codon W58X. Four gRNAs tiled 1-31 nt upstream of the UAG codon were designed and assayed, b. Fluorescence imaging of HEK293T cells expressing GFP-W58X, RFP (control), and non-targeting and targeting DECOR. gRNA b that binds 12-41 nt upstream of the target A resulted in strongest GFP recovery. Scale bar, 100 pm. c. A-to-I editing in GFP mRNA measured by high-throughput sequencing, d. Schematic of DECOR-mediated removal of a disease-causing uORF in the 5’ UTR of IRF6. The disease-causing uORF represses translation of the pORF, whereas A-to-I editing of the de novo start codon deactivates translation initiation at the uORF and reactivates expression of the pORF. Four gRNAs were tiled 1-45 nt upstream of IRF6-A- 9. A herpes simplex virus (HSV) thymidine kinase (TK) promoter or the native IRF6 promoter was used to drive the transcription of a firefly luciferase gene bearing the disease-causing uORF. Expression of the pORF was quantified by measuring luminescence. A plasmid carrying a Renilla luciferase gene was co-delivered as a transfection/cell number/cell state control, e. Normalized luciferase activity with healthy, disease-causing uORF-containing, and A-to-G mutated IRF6 5’ UTRs. f. Normalized luciferase activity. Transcription of the uORF-bearing luciferase reporter was driven by an HSV-TK promoter, g. A-to-I editing at IRF6-A- 9. h. Normalized luciferase activity. Transcription of the uORF-bearing luciferase reporter was driven by the native IRF6 promoter, i. A-to-I editing at ////’6-A-49. Cells were treated by non-targeting or ZKFd-targeting DECOR in f-i. All experiments were repeated three times (mean ± s.d., n = 3) in HEK293T cells except h (mean ± s.d., n = 6) and i (mean ± s.d., n = 4).
[0046] FIGs. 6A-6B show normalized expression levels for AD ARI (a) and ADAR2 (b) in 55 human tissues. The figure was produced with data obtained from the Human Protein Atlas (https://www.proteinatlas.org). TPM: transcript per million.
[0047] FIGs. 7A-7B show programmed editing of EZH2 mRNA by fusion proteins of dCas!3d and cytidine deaminase, a. Architectures of dCasl3d-cytidine deaminase fusion proteins. Cytidine deaminases include rAPOBECl, evoCDAl, and hAPOBEC3A. b. C-to-U editing across all C sites in EZH2-1012-1257 deposited by dCasl3d and cytidine deaminase fusions. The gRNA binding site is annotated in the heatmap. All fusion proteins were evaluated in HEK293T cells with EZH2 -targeting and non-targeting (NT) gRNA (n = 3).
[0048] FIGs. 8A-8B show programmed editing of EZH2 mRNA by adenosine deaminases guided by a Casl3d complex, a. Architectures of dCasl3d and adenosine deaminase fusion proteins. Adenosine deaminases evaluated in this assay included, coli tRNA adenosine deaminase TadA and three TadA derivatives, TadA7.10, TadA8.20, and TadA8e. b. A-to-I editing across all A sites in EZH2- 1009- 1258 deposited by dCasl3d and TadA fusions. The gRNA binding site is annotated in the heatmap. All fusion proteins were evaluated in HEK293T cells with EZH2 -targeting and non-targeting (NT) gRNA (n = 3).
[0049] FIGs. 9A-9D show programmed RNA editing by dCas!3d and TadA8e fusions with different signal peptides, a. Architectures of dCasl3d and TadA8e fusion proteins, b-d. A-to-I editing at EZH2-KW29 (b), ACTB-A1220 (c), and MALAT1-A2A 6 (d) deposited by fusions of dCasl3d and TadA8e with different signal peptides. All fusion proteins were evaluated in HEK293T cells with targeting and non-targeting (NT) gRNA (mean ± s.d., n = 3). [0050] FIGs. 10A-10B show programmed editing of ACTB mRNA by dCas!3 and TadA8e fusions, a. Architectures of dCasl3 and TadA8e fusion proteins. Casl3: LwaCasl3a, PspCasl3b, RfxCasl3d, Casl3X. l, and mini-Casl3X. l. b. A-to-I editing at ACTB-AV22Q deposited by fusions of dCasl3 and TadA8e. All fusion proteins were evaluated in HEK293T cells with targeting and non-targeting (NT) gRNA (mean ± s.d., n = 3).
[0051] FIG. 11 shows A-to-I editing of B4GALNT1-A95 by DECOR paired with nontargeting (NT) gRNA and gRNAs targeting B4GALNT1, EZH2 and KRAS. All experiments were carried out in HEK293T cells (mean ± s.d., n = 3).
[0052] FIGs. 12A-12D show heatmaps of DECOR- and REPAIRvl-mediated A-to-I editing in seven cellular RNAs. A positions in RNA refer to the coordinates in NCBI accession (Table 1). Black lines indicate the gRNA binding regions. Editing rates are presented as the mean of three independent replicates, a. ACTB (left) and B4GALNT1 (right), b. EZH2 (left) and KRAS (right), c. MAL ATI (left) and METTL3 (right), d. STAT3.
[0053] FIGs. 13A-13C show programmed RNA editing by DECOR with varying linker lengths between TadA8e and dCasl3. a. Architectures of DECOR, b. A-to-I editing mediated by DECOR of varying linker lengths at designated As in three cellular RNAs. c. Heatmaps of all A edits observed in target RNAs. All fusion proteins were evaluated in HEK293T cells with targeting and non-targeting (NT) gRNA (mean ± s.d., n = 3).
[0054] FIGs. 14A-14C show context preferences of DECOR, a. Schematic diagram of the identification process for flanking sequences preferred by DECOR. DECOR was designed with gRNA binding 11 nt upstream of the target NAN sequences in GFP mRNA to enable editing. High-throughput sequencing reads of reverse-transcribed GFP mRNA were demultiplexed and analyzed for A-to-I editing, b. Editing rates across all possible NAN motifs, c. Heatmap of editing rates observed at the target A with different 5’ and 3’ flanking nucleotides. The context preferences of DECOR were characterized in HEK293T cells (mean ± s.d., n = 3).
[0055] FIGs. 15A-15B show Editing of coding and noncoding RNA by DECOR and ADAR-based RNA-editing systems, a. Design and placement of REPAIR.tl, ADAR- dCasl3d, and dCasl3d-ADAR gRNAs. Design and placement of DECOR, REPAIRvl, and REPAIRv2 gRNAs are provided in FIG. 2A. b. A-to-I editing mediated by DECOR and ADAR-based RNA-editing systems at designated As across seven cellular RNAs. All experiments were carried out in HEK293T cells (mean ± s.d., n = 3). The same set of data was used to plot FIG. 2D.
[0056] FIGs. 16A-16G show transcriptome-wide specificity of DECOR and ADAR- based RNA-editing systems, a. A-to-I editing detected at EZH2-KW19 by whole- transcriptome sequencing, b. A-to-I editing detected at AXES'- A 504 and MALAT1- 2A 6 by whole-transcriptome sequencing, c. Transcriptome-wide off-target sites of different editors detected in individual replicates and those persisting across two replicates, d. Venn diagram showing the overlap of off-target sites between EZH2 -targeting and NT REPAIRvl. e-f. Pie charts depicting off-target sites of REPAIRvl paired with EZH2 -targeting and non-targeting gRNAs that showed non-zero editing in untreated cells, g. Venn diagram showing off-target sites shared by EZH2 -targeting and NT DECOR. All experiments were carried out in duplicate in HEK293T cells.
[0057] FIGs. 17A-17J show gRNA-independent off-target DNA editing by DECOR, a. Overview of the R-loop assay adapted to evaluate the off-target DNA-editing activity of DNA base editors and DECOR. DECOR or SpABE8e was delivered into HEK293T cells alongside a catalytically inactive SaCas9 (dSaCas9) and a Sa sgRNA that targets a specified genomic locus. A-to-G conversion within the R-loop created by the dSaCas9 complex serves a surrogate for gRNA-independent off-target effects in DNA. b. On-target A-to-G editing ? HEK4-M by SpABE8e when codelivered with dSaCas9 and a Sa sgRNA. Six experiments were carried out in parallel to generate R-loops at 6 separate genomic loci. c. On-target A-to-I RNA editing at EZH2-A1179 by DECOR when codelivered with dSaCas9 and a Sa sgRNA targeting genomic site 1. d-i. Off-target A-to-G editing by DECOR and SpABE8e at R-loops created by dSaCas9. A sites edited to greater than 1% are plotted, j. Box and whisker plot of editing rates delivered by SpABE8e and DECOR at 6 R-loops as shown in d-i. The solid line within the box represents the median, while upper and lower hinges depict the first and third quartiles. Whiskers extend from the minimum to the maximum. The plus sign (+) denotes the mean. Statistical analysis was done using a two-tailed Student's t-test. ** P < 0.01 (n = 12). All experiments were carried out in HEK293T cells (mean ± s.d., n = 3).
[0058] FIGs. 18A-18G show efficiency and specificity of hfDECOR. a. A-to-I editing mediated by hfDECOR detected at EZH2 -Al 179 by whole-transcriptome sequencing, b. Transcriptome-wide off-target sites of hfDECOR detected by individual replicates and those persisting across replicates, c. Transcriptome-wide off-target A-to-I editing by hfDECOR- V106G. d. Violin plots of editing rates detected at off-target sites hit by hfDECOR-V106G. Solid lines represent the median, whereas dashed lines denote the first and third quantile. Statistical analysis was done using a two-tailed Student's t-test. n.s. P>0.05. e. Venn diagram showing off-target sites shared by DECOR and hfDECOR-V106G. f. Venn diagram showing the overlap of off-target sites between REPAIRvl and hfDECOR-V106G. g. Sequence logos of A sites impacted by hfDECOR-V106G in the HEK293T transcriptome. All overlap analysis presented here uses data obtained with non-targeting gRNA. The inventors observe similar overlapping rates with EZH2 -targeting gRNA. All experiments were carried out in duplicate in HEK293T cells.
[0059] FIG. 19 shows programmed RNA editing with an evolved bacterial adenosine deaminase.
[0060] FIGs. 20A-20D shows RNA editing in HEK293T cells, a. EZH2 RNA editing in HEK293T cells, b. RPS5 RNA editing in HEK293T cells, c. UBE3A RNA editing in HEK293T cells, d. XIST RNA editing in HEK293T cells.
DETAILED DESCRIPTION OF THE INVENTION
[0061] Programmed RNA editing is an attractive strategy to treat genetic disease. It has so far been limited to a single family of effector proteins, adenosine deaminases acting on RNA (ADARs). Aspects herein include bacterial deaminase-enabled recoding of RNA (DECOR), a CRISPR-based, ADAR-independent RNA-editing platform that delivers robust adenosine-to- inosine editing to user-specified sites in the human transcriptome. DECOR can exploit an evolved bacterial tRNA adenosine deaminase, TadA8e, and shows transcriptome-wide off- target effects significantly lower than ADAR-overexpressing RNA-editing platforms. Installation of high-fidelity mutations into the effector protein TadA8e can further reduce the off-target effects of DECOR to almost basal level. By expanding the current portfolio of effector proteins, DECOR can unlock new territory in RNA-editing technologies. [0062] Advantages of certain aspects herein include 1) it is the first ADAR-independent A- to-I editing platform, DECOR expands the current portfolio of programmed RNA-editing technologies. DECOR has features distinct from existing technologies, e.g., sequence preference, cellular localization, and target scope (single-stranded RNA versus doublestranded RNA). 2) DECOR is broadly compatible with RNA-targeting CRISPR systems, delivering editing activity comparable or higher than existing RNA-editing technologies. 3) DECOR has lower off-target effects than REPAIRvl, the a previous RNA base editing system. 4) DECOR-enabled removal of a disease causing upstream open reading frame by RNA editing, demonstrating the potential in therapeutic applications. 5) programmed C-to-U editing in RNA is an unsolved problem. With deaminases described herein, programmed C-to-U editing in RNA can be achieved.
I. Proteinaceous Compositions
[0063] As used herein, a “protein” “peptide” or “polypeptide” refers to a molecule comprising at least five amino acid residues. As used herein, the term “wild-type” refers to the endogenous version of a molecule that occurs naturally in an organism. In some aspects, wildtype versions of a protein or polypeptide are employed, however, in many aspects of the disclosure, a modified protein or polypeptide is employed to generate an immune response. The terms described above may be used interchangeably. A “modified protein” or “modified polypeptide” or a “variant” refers to a protein or polypeptide whose chemical structure, particularly its amino acid sequence, is altered with respect to the wild-type protein or polypeptide. In some aspects, a modified/variant protein or polypeptide has at least one modified activity or function (recognizing that proteins or polypeptides may have multiple activities or functions). It is specifically contemplated that a modified/variant protein or polypeptide may be altered with respect to one activity or function yet retain a wild-type activity or function in other respects, such as immunogenicity.
[0064] Where a protein is specifically mentioned herein, it is in general a reference to a native (wild-type) or recombinant (modified) protein or, optionally, a protein in which any signal sequence has been removed. The protein may be isolated directly from the organism of which it is native, produced by recombinant DNA/exogenous expression methods, or produced by solid-phase peptide synthesis (SPPS) or other in vitro methods. In particular aspects, there are isolated nucleic acid segments and recombinant vectors incorporating nucleic acid sequences that encode a polypeptide (e.g., a tRNA-specific deaminase or fragment thereof and/or an RNA-binding protein or fragment thereof). The term “recombinant” may be used in conjunction with a polypeptide or the name of a specific polypeptide, and this generally refers to a polypeptide produced from a nucleic acid molecule that has been manipulated in vitro or that is a replication product of such a molecule.
[0065] In certain aspects the size of a protein or polypeptide (wild-type or modified) may comprise, but is not limited to, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22,
23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47,
48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72,
73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,
98, 99, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 525, 550, 575, 600, 625, 650, 675, 700, 725, 750, 775, 800, 825, 850, 875, 900, 925, 950, 1000, 1200, 1400, 1600, 1800, or 2000 amino acid residues or nucleic acid residues or greater, and any range derivable therein, or derivative of a corresponding amino sequence described or referenced herein. It is contemplated that polypeptides may be mutated by truncation, rendering them shorter than their corresponding wild-type form, also, they might be altered by fusing or conjugating a heterologous protein or polypeptide sequence with a particular function (e.g., for targeting or localization, for enhanced immunogenicity, for purification purposes, etc.).
[0066] The polypeptides, proteins, or polynucleotides encoding such polypeptides or proteins of the disclosure may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, or 50 (or any derivable range therein) or more variant amino acids or nucleic acid substitutions or be at least 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any derivable range therein) similar, identical, or homologous to at least, or at most 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123,
124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,
143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161,
162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200 or more contiguous amino acids, or any range derivable therein, of a sequence disclosed herein. In specific aspects, the peptide or polypeptide is or is based on a human sequence. In certain aspects, the peptide or polypeptide is not naturally occurring and/or is in a combination of peptides or polypeptides.
[0067] The polypeptides of the disclosure may include at least, at most, or exactly 1, 2, 3,
4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30,
31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55,
56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80,
81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123,
124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,
143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161,
162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180,
181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, or 200 substitutions (or any range derivable therein).
[0068] In some aspects, the polypeptide comprises one or more substitutions at one or more amino acid positions selected from amino acid 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16,
17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41,
42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66,
67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91,
92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131,
132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150,
151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169,
170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188,
189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 300, 400, 500, 600, 700, 800, 900,
1000, 1100, and/or 1200 of any of the sequences disclosed herein, wherein each substitution is independently chosen from an amino acid selected from alanine, arginine, asparagine, aspartic acid, cysteine, glutamine, glutamic acid, glycine, histidine, isoleucine, leucine, lysine, methionine, phenylalanine, proline, serine, threonine, tryptophan, tyrosine, or valine; and wherein the polypeptide is or is at least 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any derivable range therein) sequence identity to one sequence disclosed herein.
[0069] In some aspects, the protein or polypeptide may comprise amino acids 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123,
124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,
143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161,
162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180,
181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199,
200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200 (or any derivable range therein) of any sequence disclosed herein.
[0070] In some aspects, the protein or polypeptide may comprise amino acids 1 to 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123,
124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142,
143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161,
162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180,
181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199,
200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200 (or any derivable range therein) of a sequence disclosed herein and have or have at least 60%, 61%, 62%, 63%, 64%, 65%, 66%,
67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%,
83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%,
99%, or 100% (or any derivable range therein) sequence identity to one sequence disclosed herein.
[0071] In some aspects, the protein may comprise, comprise at least, or comprise at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122,
123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141,
142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160,
161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179,
180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198,
199, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200 (or any derivable range therein) contiguous amino acids or nucleic acids of a sequence disclosed herein..
[0072] In some aspects, the polypeptide, protein, or nucleic acid may comprise at least, at most, or exactly 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24,
25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49,
50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74,
75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99,
100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118,
119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137,
138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156,
157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175,
176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194,
195, 196, 197, 198, 199, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200 (or any derivable range therein) contiguous amino acids of a sequence disclosed herein that are at least, at most, or exactly 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% (or any derivable range therein) similar, identical, or homologous to one of a sequence disclosed herein.
[0073] In some aspects there is a nucleic acid molecule or polypeptide starting at position 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28,
29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53,
54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78,
79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102,
103, 104, 105, 106, 107, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121,
122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140,
141, 142, 143, 144, 145, 146, 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168, 169, 170, 171, 172, 173, 174, 175, 176, 177, 178,
179, 180, 181, 182, 183, 184, 185, 186, 187, 188, 189, 190, 191, 192, 193, 194, 195, 196, 197,
198, 199, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1100, 1200 of any sequence disclosed herein and comprising at least, at most, or exactly 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15,
16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40,
41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65,
66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90,
91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 110, 111,
112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130,
131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146, 147, 148, 149,
150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160, 161, 162, 163, 164, 165, 166, 167, 168,
169, 170, 171, 172, 173, 174, 175, 176, 177, 178, 179, 180, 181, 182, 183, 184, 185, 186, 187,
188, 189, 190, 191, 192, 193, 194, 195, 196, 197, 198, 199, 200, 300, 400, 500, 600, 700, 800,
900, 1000, 1100, 1200 (or any derivable range therein) contiguous amino acids or nucleotides of any sequence disclosed herein.
[0074] The nucleotide as well as the protein, polypeptide, and peptide sequences for various genes have been previously disclosed, and may be found in the recognized computerized databases. Two commonly used databases are the National Center for Biotechnology Information’s Genbank and GenPept databases (on the World Wide Web at ncbi.nlm.nih.gov/) and The Universal Protein Resource (UniProt; on the World Wide Web at uniprot.org). The coding regions for these genes may be amplified and/or expressed using the techniques disclosed herein or as would be known to those of ordinary skill in the art.
[0075] It is contemplated that in compositions of the disclosure, there is between about 0.001 mg and about 10 mg of total polypeptide, peptide, and/or protein per ml. The concentration of protein in a composition can be about, at least about or at most about 0.001, 0.010, 0.050, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0, 1.5, 2.0, 2.5, 3.0, 3.5, 4.0, 4.5, 5.0, 5.5, 6.0, 6.5, 7.0, 7.5, 8.0, 8.5, 9.0, 9.5, 10.0 mg/ml or more (or any range derivable therein).
[0076] The following is a discussion of changing the amino acid subunits of a protein to create an equivalent, or even improved, second-generation variant polypeptide or peptide. For example, certain amino acids may be substituted for other amino acids in a protein or polypeptide sequence with or without appreciable loss of binding capacity or enzymatic activity. Since it is the interactive capacity and nature of a protein that defines that protein’s functional activity, certain amino acid substitutions can be made in a protein sequence and in its corresponding DNA coding sequence, and nevertheless produce a protein with similar or desirable properties. It is thus contemplated by the inventors that various changes may be made in the DNA sequences of genes which encode proteins without appreciable loss of their biological utility or activity.
[0077] The term “functionally equivalent codon” is used herein to refer to codons that encode the same amino acid, such as the six different codons for arginine. Also considered are “neutral substitutions” or “neutral mutations” which refers to a change in the codon or codons that encode biologically equivalent amino acids.
[0078] Amino acid sequence variants of the disclosure can be substitutional, insertional, or deletion variants. A variation in a polypeptide of the disclosure may affect 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, or more non-contiguous or contiguous amino acids of the protein or polypeptide, as compared to wild-type (or any range derivable therein). A variant can comprise an amino acid sequence that is at least 50%, 60%, 70%, 80%, or 90%, including all values and ranges there between, identical to any sequence provided or referenced herein. A variant can include 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more substitute amino acids.
[0079] It also will be understood that amino acid and nucleic acid sequences may include additional residues, such as additional N- or C-terminal amino acids, or 5' or 3' sequences, respectively, and yet still be essentially identical as set forth in one of the sequences disclosed herein, so long as the sequence meets the criteria set forth above, including the maintenance of biological protein activity where protein expression is concerned. The addition of terminal sequences particularly applies to nucleic acid sequences that may, for example, include various non-coding sequences flanking either of the 5' or 3' portions of the coding region.
[0080] Deletion variants typically lack one or more residues of the native or wild type protein. Individual residues can be deleted or a number of contiguous amino acids can be deleted. A stop codon may be introduced (by substitution or insertion) into an encoding nucleic acid sequence to generate a truncated protein.
[0081] Insertional mutants typically involve the addition of amino acid residues at a nonterminal point in the polypeptide. This may include the insertion of one or more amino acid residues. Terminal additions may also be generated and can include fusion proteins which are multimers or concatemers of one or more peptides or polypeptides described or referenced herein.
[0082] Substitutional variants typically contain the exchange of one amino acid for another at one or more sites within the protein or polypeptide, and may be designed to modulate one or more properties of the polypeptide, with or without the loss of other functions or properties. Substitutions may be conservative, that is, one amino acid is replaced with one of similar chemical properties. “Conservative amino acid substitutions” may involve exchange of a member of one amino acid class with another member of the same class. Conservative substitutions are well known in the art and include, for example, the changes of: alanine to serine; arginine to lysine; asparagine to glutamine or histidine; aspartate to glutamate; cysteine to serine; glutamine to asparagine; glutamate to aspartate; glycine to proline; histidine to asparagine or glutamine; isoleucine to leucine or valine; leucine to valine or isoleucine; lysine to arginine; methionine to leucine or isoleucine; phenylalanine to tyrosine, leucine or methionine; serine to threonine; threonine to serine; tryptophan to tyrosine; tyrosine to tryptophan or phenylalanine; and valine to isoleucine or leucine. Conservative amino acid substitutions may encompass non-naturally occurring amino acid residues, which are typically incorporated by chemical peptide synthesis rather than by synthesis in biological systems. These include peptidomimetics or other reversed or inverted forms of amino acid moieties.
[0083] Alternatively, substitutions may be “non-conservative”, such that a function or activity of the polypeptide is affected. Non-conservative changes typically involve substituting an amino acid residue with one that is chemically dissimilar, such as a polar or charged amino acid for a nonpolar or uncharged amino acid, and vice versa. Non-conservative substitutions may involve the exchange of a member of one of the amino acid classes for a member from another class.
[0084] One skilled in the art can determine suitable variants of polypeptides as set forth herein using well-known techniques. One skilled in the art may identify suitable areas of the molecule that may be changed without destroying activity by targeting regions not believed to be important for activity. The skilled artisan will also be able to identify amino acid residues and portions of the molecules that are conserved among similar proteins or polypeptides. In further aspects, areas that may be important for biological activity or for structure may be subject to conservative amino acid substitutions without significantly altering the biological activity or without adversely affecting the protein or polypeptide structure. [0085] In making such changes, the hydropathy index of amino acids may be considered. The hydropathy profile of a protein is calculated by assigning each amino acid a numerical value (“hydropathy index”) and then repetitively averaging these values along the peptide chain. Each amino acid has been assigned a value based on its hydrophobicity and charge characteristics. They are: isoleucine (+4.5); valine (+4.2); leucine (+3.8); phenylalanine (+2.8); cysteine/cysteine (+2.5); methionine (+1.9); alanine (+1.8); glycine (-0.4); threonine (-0.7); serine (-0.8); tryptophan (-0.9); tyrosine (-1.3); proline (1.6); histidine (-3.2); glutamate (-3.5); glutamine (-3.5); aspartate (-3.5); asparagine (-3.5); lysine (-3.9); and arginine (-4.5). The importance of the hydropathy amino acid index in conferring interactive biologic function on a protein is generally understood in the art (Kyte et al., J. Mol. Biol. 157: 105-131 (1982)). It is accepted that the relative hydropathic character of the amino acid contributes to the secondary structure of the resultant protein or polypeptide, which in turn defines the interaction of the protein or polypeptide with other molecules, for example, enzymes, substrates, receptors, DNA, antibodies, antigens, and others. It is also known that certain amino acids may be substituted for other amino acids having a similar hydropathy index or score, and still retain a similar biological activity. In making changes based upon the hydropathy index, in certain aspects, the substitution of amino acids whose hydropathy indices are within ±2 is included. In some aspects of the invention, those that are within ±1 are included, and in other aspects of the invention, those within ±0.5 are included.
[0086] It also is understood in the art that the substitution of like amino acids can be effectively made based on hydrophilicity. U.S. Patent 4,554,101, incorporated herein by reference, states that the greatest local average hydrophilicity of a protein, as governed by the hydrophilicity of its adjacent amino acids, correlates with a biological property of the protein. In certain aspects, the greatest local average hydrophilicity of a protein, as governed by the hydrophilicity of its adjacent amino acids, correlates with its immunogenicity and antigen binding, that is, as a biological property of the protein. The following hydrophilicity values have been assigned to these amino acid residues: arginine (+3.0); lysine (+3.0); aspartate (+3.0+1); glutamate (+3.0+1); serine (+0.3); asparagine (+0.2); glutamine (+0.2); glycine (0); threonine (_0.4); proline (-0.5+1); alanine (_0.5); histidine (_0.5); cysteine (—1.0); methionine (-1.3); valine (-1.5); leucine (-1.8); isoleucine (-1.8); tyrosine (-2.3); phenylalanine (-2.5); and tryptophan (-3.4). In making changes based upon similar hydrophilicity values, in certain aspects, the substitution of amino acids whose hydrophilicity values are within ±2 are included, in other aspects, those which are within ±1 are included, and in still other aspects, those within ±0.5 are included. In some instances, one may also identify epitopes from primary amino acid sequences based on hydrophilicity. These regions are also referred to as “epitopic core regions.” It is understood that an amino acid can be substituted for another having a similar hydrophilicity value and still produce a biologically equivalent and immunologically equivalent protein.
[0087] Additionally, one skilled in the art can review structure-function studies identifying residues in similar polypeptides or proteins that are important for activity or structure. In view of such a comparison, one can predict the importance of amino acid residues in a protein that correspond to amino acid residues important for activity or structure in similar proteins. One skilled in the art may opt for chemically similar amino acid substitutions for such predicted important amino acid residues.
[0088] One skilled in the art can also analyze the three-dimensional structure and amino acid sequence in relation to that structure in similar proteins or polypeptides. In view of such information, one skilled in the art may predict the alignment of amino acid residues of an protein with respect to its three-dimensional structure. One skilled in the art may choose not to make changes to amino acid residues predicted to be on the surface of the protein, since such residues may be involved in important interactions with other molecules. Moreover, one skilled in the art may generate test variants containing a single amino acid substitution at each desired amino acid residue. These variants can then be screened using standard assays for binding and/or activity, thus yielding information gathered from such routine experiments, which may allow one skilled in the art to determine the amino acid positions where further substitutions should be avoided either alone or in combination with other mutations. Various tools available to determine secondary structure can be found on the world wide web at expasy.org/proteomics/protein structure.
[0089] In some aspects of the invention, amino acid substitutions are made that: (1) reduce susceptibility to proteolysis, (2) reduce susceptibility to oxidation, (3) alter binding affinity for forming protein complexes, (4) alter ligand or antigen binding affinities, and/or (5) confer or modify other physicochemical or functional properties on such polypeptides. For example, single or multiple amino acid substitutions (in certain aspects, conservative amino acid substitutions) may be made in the naturally occurring sequence. In such aspects, conservative amino acid substitutions can be used that do not substantially change the structural characteristics of the protein or polypeptide (e.g., one or more replacement amino acids that do not disrupt the secondary structure that characterizes the native protein). II. Nucleic Acids
[0090] In certain aspects, nucleic acid sequences can exist in a variety of instances such as: isolated segments and recombinant vectors of incorporated sequences or recombinant polynucleotides encoding a protein (including any polypeptide disclosed herein), or a fragment, derivative, mutein, or variant thereof, polynucleotides sufficient for use as hybridization probes, PCR primers or sequencing primers for identifying, analyzing, mutating or amplifying a polynucleotide encoding a polypeptide, anti-sense nucleic acids for inhibiting expression of a polynucleotide, and complementary sequences of the foregoing described herein. Nucleic acids that encode one or multiple domains of the polypeptides disclosed herein are provided herein are also provided. Nucleic acids encoding fusion proteins that include these peptides are also provided. The nucleic acids can be single-stranded or double-stranded and can comprise RNA and/or DNA nucleotides and artificial variants thereof (e.g., peptide nucleic acids).
[0091] The term “polynucleotide” refers to a nucleic acid molecule that either is recombinant or has been isolated from total genomic nucleic acid. Included within the term “polynucleotide” are oligonucleotides (nucleic acids 100 residues or less in length), recombinant vectors, including, for example, plasmids, cosmids, phage, viruses, and the like. Polynucleotides include, in certain aspects, regulatory sequences, isolated substantially away from their naturally occurring genes or protein encoding sequences. Polynucleotides may be single- stranded (coding or antisense) or double- stranded, and may be RNA, DNA (genomic, cDNA or synthetic), analogs thereof, or a combination thereof. Additional coding or noncoding sequences may, but need not, be present within a polynucleotide.
[0092] In this respect, the term “gene,” “polynucleotide,” or “nucleic acid” is used to refer to a nucleic acid that encodes a protein, polypeptide, or peptide (including any sequences required for proper transcription, post-translational modification, or localization). As will be understood by those in the art, this term encompasses genomic sequences, expression cassettes, cDNA sequences, and smaller engineered nucleic acid segments that express, or may be adapted to express, proteins, polypeptides, domains, peptides, fusion proteins, and mutants. A nucleic acid encoding all or part of a polypeptide may contain a contiguous nucleic acid sequence encoding all or a portion of such a polypeptide. It also is contemplated that a particular polypeptide may be encoded by nucleic acids containing variations having slightly different nucleic acid sequences but, nonetheless, encode the same or substantially similar protein. [0093] In certain aspects, there are polynucleotide variants having substantial identity to the sequences disclosed herein; those comprising at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% or higher sequence identity, including all values and ranges there between, compared to a polynucleotide sequence provided herein using the methods described herein (e.g., BLAST analysis using standard parameters). In certain aspects, the isolated polynucleotide will comprise a nucleotide sequence encoding a polypeptide that has at least 90%, preferably 95% and above, identity to an amino acid sequence described herein, over the entire length of the sequence; or a nucleotide sequence complementary to said isolated polynucleotide.
[0094] The nucleic acid segments, regardless of the length of the coding sequence itself, may be combined with other nucleic acid sequences, such as promoters, polyadenylation signals, additional restriction enzyme sites, multiple cloning sites, other coding segments, and the like, such that their overall length may vary considerably. The nucleic acids can be any length. They can be, for example, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 75, 100, 125, 175, 200, 250, 300, 350, 400, 450, 500, 750, 1000, 1500, 3000, 5000 or more nucleotides in length, and/or can comprise one or more additional sequences, for example, regulatory sequences, and/or be a part of a larger nucleic acid, for example, a vector. It is therefore contemplated that a nucleic acid fragment of almost any length may be employed, with the total length preferably being limited by the ease of preparation and use in the intended recombinant nucleic acid protocol. In some cases, a nucleic acid sequence may encode a polypeptide sequence with additional heterologous coding sequences, for example to allow for purification of the polypeptide, transport, secretion, post-translational modification, or for therapeutic benefits such as targeting or efficacy. As discussed above, a tag or other heterologous polypeptide may be added to the modified polypeptide-encoding sequence, wherein “heterologous” refers to a polypeptide that is not the same as the modified polypeptide.
A. Hybridization
[0095] The nucleic acids that hybridize to other nucleic acids under particular hybridization conditions. Methods for hybridizing nucleic acids are well known in the art. See, e.g., Current Protocols in Molecular Biology, John Wiley and Sons, N.Y. (1989), 6.3.1-6.3.6. As defined herein, a moderately stringent hybridization condition uses a prewashing solution containing 5* sodium chloride/sodium citrate (SSC), 0.5% SDS, 1.0 mM EDTA (pH 8.0), hybridization buffer of about 50% formamide, 6* SSC, and a hybridization temperature of 55° C. (or other similar hybridization solutions, such as one containing about 50% formamide, with a hybridization temperature of 42° C), and washing conditions of 60° C. in 0.5*SSC, 0.1% SDS. A stringent hybridization condition hybridizes in 6*SSC at 45° C., followed by one or more washes in 0.1 * SSC, 0.2% SDS at 68° C. Furthermore, one of skill in the art can manipulate the hybridization and/or washing conditions to increase or decrease the stringency of hybridization such that nucleic acids comprising nucleotide sequence that are at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99% identical to each other typically remain hybridized to each other.
[0096] The parameters affecting the choice of hybridization conditions and guidance for devising suitable conditions are set forth by, for example, Sambrook, Fritsch, and Maniatis (Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., chapters 9 and 11 (1989); Current Protocols in Molecular Biology, Ausubel et al., eds., John Wiley and Sons, Inc., sections 2.10 and 6.3-6.4 (1995), both of which are herein incorporated by reference in their entirety for all purposes) and can be readily determined by those having ordinary skill in the art based on, for example, the length and/or base composition of the DNA.
B. Mutation
[0097] Changes can be introduced by mutation into a nucleic acid, thereby leading to changes in the amino acid sequence of a polypeptide that it encodes. Mutations can be introduced using any technique known in the art. In one aspect, one or more particular amino acid residues are changed using, for example, a site-directed mutagenesis protocol. In another aspect, one or more randomly selected residues are changed using, for example, a random mutagenesis protocol. However it is made, a mutant polypeptide can be expressed and screened for a desired property.
[0098] Mutations can be introduced into a nucleic acid without significantly altering the biological activity of a polypeptide that it encodes. For example, one can make nucleotide substitutions leading to amino acid substitutions at non-essential amino acid residues. Alternatively, one or more mutations can be introduced into a nucleic acid that selectively changes the biological activity of a polypeptide that it encodes. See, eg., Romain Studer et al., Biochem. J. 449:581-594 (2013). For example, the mutation can quantitatively or qualitatively change the biological activity. Examples of quantitative changes include increasing, reducing or eliminating the activity. C. Probes
[0099] In another aspect, nucleic acid molecules are suitable for use as primers or hybridization probes for the detection of nucleic acid sequences. A nucleic acid molecule can comprise only a portion of a nucleic acid sequence encoding a full-length polypeptide, for example, a fragment that can be used as a probe or primer or a fragment encoding an active portion of a given polypeptide.
[00100] In another aspect, the nucleic acid molecules may be used as probes or PCR primers for specific protein sequences. For instance, a nucleic acid molecule probe may be used in diagnostic methods or a nucleic acid molecule PCR primer may be used to amplify regions of DNA that could be used, inter alia, to isolate nucleic acid sequences for use in producing proteins. See, eg., Gaily Kivi et al., BMC Biotechnol. 16:2 (2016). In a preferred aspect, the nucleic acid molecules are oligonucleotides.
[00101] Probes based on the desired sequence of a nucleic acid can be used to detect the nucleic acid or similar nucleic acids, for example, transcripts encoding a polypeptide of interest. The probe can comprise a label group, e.g., a radioisotope, a fluorescent compound, an enzyme, or an enzyme co-factor. Such probes can be used to identify a cell that expresses the polypeptide.
III. Polypeptide Expression
[00102] In some aspects, there are nucleic acid molecule encoding polypeptides or peptides of the disclosure. These may be generated by methods known in the art, e.g., expressed in any suitable recombinant expression system and allowed to assemble to form protein compositions.
A. Expression
[00103] The nucleic acid molecules may be used to express large quantities of polypeptides.
B. Vectors
[00104] In some aspects, contemplated are expression vectors comprising a nucleic acid molecule encoding a polypeptide of the desired sequence or a portion thereof (e.g., a fragment containing one or more polypeptide or protein domains). In addition to control sequences that govern transcription and translation, vectors and expression vectors may contain nucleic acid sequences that serve other functions as well.
[00105] To express the polypeptides or peptides of the disclosure, DNAs encoding the polypeptides or peptides are inserted into expression vectors such that the gene area is operatively linked to transcriptional and translational control sequences. Typically, expression vectors used in any of the host cells contain sequences for plasmid or virus maintenance and for cloning and expression of exogenous nucleotide sequences. Such sequences, collectively referred to as “flanking sequences” typically include one or more of the following operatively linked nucleotide sequences: a promoter, one or more enhancer sequences, an origin of replication, a transcriptional termination sequence, a complete intron sequence containing a donor and acceptor splice site, a sequence encoding a leader sequence for polypeptide secretion, a ribosome binding site, a polyadenylation sequence, a polylinker region for inserting the nucleic acid encoding the polypeptide to be expressed, and a selectable marker element. Such sequences and methods of using the same are well known in the art.
C. Expression Systems
[00106] Numerous expression systems exist that comprise at least a part or all of the expression vectors discussed above. Prokaryote- and/or eukaryote-based systems can be employed for use with an aspect to produce nucleic acid sequences, or their cognate polypeptides, proteins and peptides. Commercially and widely available systems include in but are not limited to bacterial, mammalian, yeast, and insect cell systems. Different host cells have characteristic and specific mechanisms for the post-translational processing and modification of proteins. Appropriate cell lines or host systems can be chosen to ensure the correct modification and processing of the foreign protein expressed. Those skilled in the art are able to express a vector to produce a nucleic acid sequence or its cognate polypeptide, protein, or peptide using an appropriate expression system.
IV. Methods of Gene Transfer
[00107] Suitable methods for nucleic acid delivery to effect expression of compositions are anticipated to include virtually any method by which a nucleic acid (e.g., DNA, including viral and nonviral vectors) can be introduced into a cell, a tissue or an organism, as described herein or as would be known to one of ordinary skill in the art. Such methods include, but are not limited to, direct delivery of DNA such as by injection (U.S. Patents 5,994,624,5,981,274, 5,945,100, 5,780,448, 5,736,524, 5,702,932, 5,656,610, 5,589,466 and 5,580,859, each incorporated herein by reference), including microinjection (Harland and Weintraub, 1985; U.S. Patent 5,789,215, incorporated herein by reference); by electroporation (U.S. Patent No. 5,384,253, incorporated herein by reference); by calcium phosphate precipitation (Graham and Van Der Eb, 1973; Chen and Okayama, 1987; Rippe et al., 1990); by using DEAE dextran followed by polyethylene glycol (Gopal, 1985); by direct sonic loading (Fechheimer et al., 1987); by liposome mediated transfection (Nicolau and Sene, 1982; Fraley et al., 1979; Nicolau et al., 1987; Wong et al., 1980; Kaneda et al., 1989; Kato et al., 1991); by microprojectile bombardment (PCT Application Nos. WO 94/09699 and 95/06128; U.S. Patents 5,610,042; 5,322,783, 5,563,055, 5,550,318, 5,538,877 and 5,538,880, and each incorporated herein by reference); by agitation with silicon carbide fibers (Kaeppler et al., 1990; U.S. Patents 5,302,523 and 5,464,765, each incorporated herein by reference); by Agrobacterium mediated transformation (U.S. Patents 5,591,616 and 5,563,055, each incorporated herein by reference); or by PEG mediated transformation of protoplasts (Omirulleh et al., 1993; U.S. Patents 4,684,611 and 4,952,500, each incorporated herein by reference); by desiccation/inhibition mediated DNA uptake (Potrykus et al., 1985). Other methods include viral transduction, such as gene transfer by lentiviral or retroviral transduction.
A. Host Cells
[00108] In another aspect, contemplated are the use of host cells into which a recombinant expression vector has been introduced. Proteins can be expressed in a variety of cell types. An expression construct encoding an protein can be transfected into cells according to a variety of methods known in the art. Vector DNA can be introduced into prokaryotic or eukaryotic cells via conventional transformation or transfection techniques. Some vectors may employ control sequences that allow it to be replicated and/or expressed in both prokaryotic and eukaryotic cells. In certain aspects, the protein expression construct can be placed under control of a promoter that is linked to a disease-associated RNA. One of skill in the art would understand the conditions under which to incubate host cells to maintain them and to permit replication of a vector. Also understood and known are techniques and conditions that would allow large- scale production of vectors, as well as production of the nucleic acids encoded by vectors and their cognate polypeptides, proteins, or peptides.
[00109] For stable transfection of mammalian cells, it is known, depending upon the expression vector and transfection technique used, only a small fraction of cells may integrate the foreign DNA into their genome. In order to identify and select these integrants, a selectable marker (e.g., for resistance to antibiotics) is generally introduced into the host cells along with the gene of interest. Cells stably transfected with the introduced nucleic acid can be identified by drug selection (e.g., cells that have incorporated the selectable marker gene will survive, while the other cells die), among other methods known in the arts. B. Isolation
[00110] The nucleic acid molecule encoding the polypeptides may be obtained from any source that produces proteins. Methods of isolating mRNA encoding an protein are well known in the art. See e.g., Sambrook et al., supra. The sequences of human heavy and light chain constant region genes are also known in the art. See, e.g., Kabat et al., 1991, supra.
V. Kits
[00111] The present disclosure additionally provides kits for modifying and/or detecting modified adenosines in a target DNA. Each kit may also include additional components that are useful for amplifying the nucleic acid, or sequencing the nucleic acid, or other applications of the present disclosure as described herein. The kit may optionally provide additional components that are useful in the procedure. These optional components include buffers, capture reagents, developing reagents, labels, reacting surfaces, means for detection, control samples, instructions, and interpretive information. The kit may also include reagents for DNA isolation and/or purification.
VI. Sequences
[00112] The present disclosure provides polypeptides capable of modifying nucleotides in certain RNA molecules, including tRNA-specific deaminases linked to an RNA-targeting molecule. The polypeptide sequences of such polypeptides can include any of the following:
TadA-linker-dCasl3d
[00113] MSEVEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPI GRHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFG ARDAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKK AQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFAKG MGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGYK IGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVIH NILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNND KLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAHW VVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNYIA ETLGINPAEF AEQ YFRF SIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFD SIRT KVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKDAL YYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSKLMY ALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADEL RLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK
HGMRNFIINNVISNKRFHYLIRYGDPAHLHEIAKNEAVVKFVLGRIADIQKKQGQNG
KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE
KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS
SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE
FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA
YINDIAEVNSYFQLYHYIMQRIIMNERYEI<SSGI<VSEYFDAVNDEI<I<YNDRLLI<LLC
VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK
DDDDK (SEQ ID NO:1)
TadA7.10-linker-dCasl3d
[00114] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAI
GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG
VRNAKTGAAGSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQK
KAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFAK
GMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGY
KIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVI
HNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNN
DKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAH
WVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNY
IAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFDSI
RTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKD
ALYYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSKLM
YALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADE
LRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK
HGMRNFIINNVISNKRFHYLIRYGDPAHLHEIAKNEAVVKFVLGRIADIQKKQGQNG
KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE
KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS
SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE
FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA
YINDIAEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKLLC
VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK
DDDDK (SEQ ID NO:2)
TadA8.20-linker-dCasl3d [00115] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAI
GLHDPTAHAEIMALRQGGLVMQNYRLYDATLYSTFEPCVMCAGAMIHSRIGRVVFG
VRNAKTGAAGSLMDVLHHPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQK
KAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFAK
GMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGY
KIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVI
HNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNN
DKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAH
WVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNY
IAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFDSI
RTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKD
ALYYDEANRIWRI<LENIMHNII<EFRGNI<TREYI< I<I<DAPRLPRILPAGRDVSAFSI<LM
YALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADE
LRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK
HGMRNFIINNVISNI<RFHYLIRYGDPAHLHEIAI<NEAVVI<FVLGRIADIQI<I<QGQNG
KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE
KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS
SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE
FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA
YINDIAEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKLLC
VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK
DDDDK (SEQ ID NO:3)
TadA8e-linker-dCasl3d
[00116] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAI
GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG
VRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQK
KAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFAK
GMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGY
KIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVI
HNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNN
DKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAH
WVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNY
IAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFDSI
RTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKD ALYYDEANRIWRI<LENIMHNII<EFRGNI<TREYI< I<I<DAPRLPRILPAGRDVSAFSI<LM
YALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADE
LRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK
HGMRNFIINNVISNI<RFHYLIRYGDPAHLHEIAI<NEAVVI<FVLGRIADIQI<I<QGQNG
KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE
KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS
SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE
FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA
YINDIAEVNSYFQLYHYIMQRIIMNERYEI<SSGI<VSEYFDAVNDEI<I<YNDRLLI<LLC
VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK
DDDDK (SEQ ID N0:4)
TadA8e-F148A-linker-dCasl3d
[00117] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAI
GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG
VRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDAYRMPRQVFNAQK
KAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFAK
GMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGY
KIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVI
HNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNN
DKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAH
WVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNY
IAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFDSI
RTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKD
ALYYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSKLM
YALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADE
LRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK
HGMRNFIINNVISNKRFHYLIRYGDPAHLHEIAKNEAVVKFVLGRIADIQKKQGQNG
KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE
KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS
SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE
FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA
YINDIAEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKLLC
VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK
DDDDK (SEQ ID NO: 5) TadA8e-V106G-linker-dCasl3d
[00118] MSEVEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAI GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG GRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQK KAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFAK GMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGY KIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVI HNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNN DKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAH WVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNY lAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFDSI RTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKD ALYYDEANRIWRI<LENIMHNII<EFRGNI<TREYI< I<I<DAPRLPRILPAGRDVSAFSI<LM YALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADE LRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK HGMRNFIINNVISNI<RFHYLIRYGDPAHLHEIAI<NEAVVI<FVLGRIADIQI<I<QGQNG KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA YINDIAEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKLLC VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK DDDDK (SEQ ID NO: 6)
TadA.CBE1.46-linker-dCasl3d
[00119] MSE VEYSHE YWMRH ALTLAKRARDERH VP VGA VLVLNNR VIGEGWNRA KGLHDPTAHAEIMALRQGGLVMQNYRLWGATLYTTFEPCVMCAGAMIHSRIGRVV FGVCNAKTHACMSLMDVLGHPGMNHRVEITEGILADECEALLCRFFRMPRRVFNAQ KKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFA KGMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAG YKIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQ VIHNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFN NNDKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGL AHWVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANV NYIAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFD SIRTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQ I<DALYYDEANRIWRI<LENIMHNH<EFRGNI<TREYI< I<I<DAPRLPRILPAGRDVSAFSI<
LMYALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIA DELRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKK GI<HGMRNFIINNVISNI<RFHYLIRYGDPAHLHEIAI<NEAVVI<FVLGRIADIQI<I<QGQ NGKNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAE REKFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKG FSSVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKA EEFTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYV HAYINDIAEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKL LCVPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADY KDDDDK (SEQ ID NO: 7)
TadA.CBE1.52-linker-dCasl3d
[00120] MSEVEYSHEYWMRHALTLAKRARDERHVPVGAVLVLNNRVIGEGWNRA KGLHDPTAHAEIMALRQGGLVMQNYRLWGATLYTTFEPCVMCAGAMIHSRIGRVV FGVCNAKTHACMSLMDVLGHPGMPHRVEITEGILADECEALLCRFFRMPRRVFNAQ KKAQSSTDSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFA KGMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAG YKIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQ VIHNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFN NNDKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGL
AHWVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANV NYIAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFD SIRTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQ KDALYYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSK LMYALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIA DELRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKK GKHGMRNFIINNVISNKRFHYLIRYGDPAHLHEIAKNEAVVKFVLGRIADIQKKQGQ NGKNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAE REKFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKG FSSVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKA EEFTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYV HAYINDIAEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKL LCVPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADY KDDDDK (SEQ ID NO: 8)
TadA.CBE2.4-linker-dCasl3d
[00121] MSEVEFSHEYWMRHALTLAKRAWDERDVPVGAVLVHNNRVIGEGWNRA IGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYSTLEPCVMCAGAMIHSRIGRVVFG ARYLRMGAAGSLMNVLHYPGIKHRVEITEGILADECAALLSDFFRMPRRVFKAQKK AQ S SINSGGS SGGS SGSETPGT SES ATPES SGGS SGGS SPKKKRKVE ASIEKKKSF AKG MGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGYK IGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVIH NILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNND KLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAHW VVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNYIA ETLGINPAEF AEQ YFRF SIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFD SIRT KVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKDAL YYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSKLMY ALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADEL RLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK HGMRNFIINNVISNI<RFHYLIRYGDPAHLHEIAI<NEAVVI<FVLGRIADIQI<I<QGQNG KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA YINDIAEVNSYFQLYHYIMQRIIMNERYEI<SSGI<VSEYFDAVNDEI<I<YNDRLLKLLC
VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK DDDDK (SEQ ID NO: 9)
TadA.CBE3.1-linker-dCasl3d
[00122] MSEVEFSHEYWMRHALTLAKRAWDEAHAPVGAVLVHNNRVIGEGWNRA IGRHDPTAHAEIMALRQGGLVMQNYRLIDATLYSTLEPCVMCAGAMIHSRIGRVVFG ARYLRMGAAGSLMNVLHYPGIKHRVEITEGILADECAALLSDFFRMPRRVFKAQKK AQ S SINSGGS SGGS SGSETPGT SES ATPES SGGS SGGS SPKKKRKVE ASIEKKKSF AKG MGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGYK IGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVIH NILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNND KLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAHW VVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNYIA
ETLGINPAEF AEQ YFRF SIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFD SIRT
KVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKDAL
YYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSKLMY
ALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADEL
RLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK
HGMRNFIINNVISNKRFHYLIRYGDPAHLHEIAKNEAVVKFVLGRIADIQKKQGQNG
KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE
KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS
SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE
FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA
YINDIAEVNSYFQLYHYIMQRIIMNERYEI<SSGI<VSEYFDAVNDEI<I<YNDRLLI<LLC
VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK
DDDDK (SEQ ID NO: 10)
TadA.CBE.DLL-linker-dCasl3d
[00123] MSEVEFSHEYWMRHALTLAKRARDERRVPVGAVLVLNNRVIGEGWLRAI
GLHDPTAHAEIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFG
VRNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQK
KAQSSINSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFAK
GMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGY
KIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVI
HNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNN
DKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAH
WVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNY
IAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFDSI
RTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKD
ALYYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSKLM
YALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADE
LRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK
HGMRNFIINNVISNKRFHYLIRYGDPAHLHEIAKNEAVVKFVLGRIADIQKKQGQNG
KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE
KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS
SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE
FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA YINDIAEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKLLC VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK DDDDK (SEQ ID NO: 11)
TadA.CBE.DL-linker-dCas 13d
[00124] MSEVEFSHEYWMRHALTLAKRARDERKAPVGAVLVLNNRVIGEGWNRAI GLHDPTAHAEIIALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMINSRIGRVVFGV RNSKRGAAGSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKK AQ S SINSGGS SGGS SGSETPGT SES ATPES SGGS SGGS SPKKKRKVE ASIEKKKSF AKG MGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGYK IGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVIH NILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNND KLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAHW VVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNYIA ETLGINPAEF AEQ YFRF SIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFD SIRT KVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKDAL YYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSKLMY ALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADEL RLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK HGMRNFIINNVISNI<RFHYLIRYGDPAHLHEIAI<NEAVVI<FVLGRIADIQI<I<QGQNG KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA YINDIAEVNSYFQLYHYIMQRIIMNERYEI<SSGI<VSEYFDAVNDEI<I<YNDRLLI<LLC VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK DDDDK (SEQ ID NO: 12) dCas!3d-linker-TadA
[00125] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE
VEFSHEYWMRHALTLAKRAWDEREVPVGAVLVHNNRVIGEGWNRPIGRHDPTAHA EIMALRQGGLVMQNYRLIDATLYVTLEPCVMCAGAMIHSRIGRVVFGARDAKTGAA GSLMDVLHHPGMNHRVEITEGILADECAALLSDFFRMRRQEIKAQKKAQSSTD (SEQ ID NO: 13) dCasl3d-linker-TadA7.10
[00126] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA
HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHA EIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGAA GSLMDVLHYPGMNHRVEITEGILADECAALLCYFFRMPRQVFNAQKKAQSSTD (SEQ ID N0:14) dCasl3d-linker-TadA8.20
[00127] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA
HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHA EIMALRQGGLVMQNYRLYDATLYSTFEPCVMCAGAMIHSRIGRVVFGVRNAKTGA AGSLMDVLHHPGMNHRVEITEGILADECAALLCRFFRMPRRVFNAQKKAQSSTD (SEQ ID NO:15) dCas!3d-linker-TadA8e
[00128] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHA EIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAA GSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 16) dCas!3d-linker-TadA8e-F148A
[00129] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA
HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHA EIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAA GSLMNVLNYPGMNHRVEITEGILADECAALLCDAYRMPRQVFNAQKKAQSSIN (SEQ ID NO:17) dCasl3d-linker-TadA8e-V106G
[00130] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA
HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRARDEREVPVGAVLVLNNRVIGEGWNRAIGLHDPTAHA EIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGGRNSKRGAA GSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO: 18) dCasl3d-linker-TadA.CBE1.46
[00131] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE
VEYSHEYWMRHALTLAKRARDERHVPVGAVLVLNNRVIGEGWNRAKGLHDPTAH AEIMALRQGGLVMQNYRLWGATLYTTFEPCVMCAGAMIHSRIGRVVFGVCNAKTH ACMSLMDVLGHPGMNHRVEITEGILADECEALLCRFFRMPRRVFNAQKKAQ S STD (SEQ ID N0:19) dCasl3d-linker-TadA.CBE1.52
[00132] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA
HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEYSHEYWMRHALTLAKRARDERHVPVGAVLVLNNRVIGEGWNRAKGLHDPTAH AEIMALRQGGLVMQNYRLWGATLYTTFEPCVMCAGAMIHSRIGRVVFGVCNAKTH ACMSLMDVLGHPGMPHRVEITEGILADECEALLCRFFRMPRRVFNAQKKAQSSTD (SEQ ID NO:20) dCas!3d-linker-TadA.CBE2.4
[00133] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA
HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRAWDERDVPVGAVLVHNNRVIGEGWNRAIGRHDPTAH AEIMALRQGGLVMQNYRLIDATLYSTLEPCVMCAGAMIHSRIGRVVFGARYLRMGA AGSLMNVLHYPGIKHRVEITEGILADECAALLSDFFRMPRRVFKAQKKAQSSIN (SEQ ID NO:21) dCasl3d-linker-TadA.CBE3.1
[00134] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER
YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRAWDEAHAPVGAVLVHNNRVIGEGWNRAIGRHDPTAH AEIMALRQGGLVMQNYRLIDATLYSTLEPCVMCAGAMIHSRIGRVVFGARYLRMGA AGSLMNVLHYPGIKHRVEITEGILADECAALLSDFFRMPRRVFKAQKKAQSSIN (SEQ ID NO:22) dCasl3d-linker-TadA.CBE.DLL
[00135] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA
HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRARDERRVPVGAVLVLNNRVIGEGWLRAIGLHDPTAHA EIMALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMIHSRIGRVVFGVRNSKRGAA GSLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO:23) dCasl3d-linker-TadA.CBE.DL
[00136] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA
HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSE VEFSHEYWMRHALTLAKRARDERKAPVGAVLVLNNRVIGEGWNRAIGLHDPTAHA EIIALRQGGLVMQNYRLIDATLYVTFEPCVMCAGAMINSRIGRVVFGVRNSKRGAAG SLMNVLNYPGMNHRVEITEGILADECAALLCDFYRMPRQVFNAQKKAQSSIN (SEQ ID NO:24) rAPOBEC 1-linker-dCas 13d
[00137] MSSETGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWR HTSQNTNKHVEVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVT LFIYIARLYHHADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPR YPHLWVRLYVLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGL KSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSPKKKRKVEASIEKKKSFAKGMG VKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGYKIGN AKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVIHNIL DIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLI NAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAHWVV ANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNYIAET LGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFDSIRTKV YTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYY DEANRIWRI<LENIMHNH<EFRGNI<TREYI< I<I<DAPRLPRILPAGRDVSAFSI<LMYALT MFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIK SFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGKHGM RNFIINNVISNKRFHYLIRYGDPAHLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQI DRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKK IISL YLT VIYHILKNIVNINARYVIGFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTK LCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQI NREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHAYINDI AEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFG YCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYKDDDD K (SEQ ID NO:25) dCasl3d-linker-rAPOBECl
[00138] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSMSSE TGPVAVDPTLRRRIEPHEFEVFFDPRELRKETCLLYEINWGGRHSIWRHTSQNTNKHV EVNFIEKFTTERYFCPNTRCSITWFLSWSPCGECSRAITEFLSRYPHVTLFIYIARLYHH ADPRNRQGLRDLISSGVTIQIMTEQESGYCWRNFVNYSPSNEAHWPRYPHLWVRLY VLELYCIILGLPPCLNILRRKQPQLTFFTIALQSCHYQRLPPHILWATGLKGSAAADYK DDDDK (SEQ ID NO:26)
APOBEC-evoCDAl-linker- dCas!3d
[00139] MKRTADGSEFESPKKKRKVSTDAEYVRIHEKLDIYTFKKQFSNNKKSVSH RCYVLFELKRRGERRACFWGYAVNKPQSGTERGIHAEIFSIRKVEEYLRDNPGQFTIN WYSSWSPCADCAEKILEWYNQELRGNGHTLKIWVCKLYYEKNARNQIGLWNLRDN GVGLNVMVSEHYQCCRKIFIQSSHNQLNENRWLEKTLKRAEKRRSELSIMFQVKILH TTKSPAVSGGSSGGSSGSETPGTSESATPESSGGSSGGSSPKKKRKVEASIEKKKSFAK GMGVKSTLVSGSKVYMTTFAEGSDARLEKIVEGDSIRSVNEGEAFSAEMADKNAGY KIGNAKFSHPKGYAVVANNPLYTGPVQQDMLGLKETLEKRYFGESADGNDNICIQVI HNILDIEKILAEYITNAAYAVNNISGLDKDIIGFGKFSTVYTYDEFKDPEHHRAAFNNN DKLINAIKAQYDEFDNFLDNPRLGYFGQAFFSKEGRNYIINYGNECYDILALLSGLAH WVVANNEEESRISRTWLYNLDKNLDNEYISTLNYLYDRITNELTNSFSKNSAANVNY
IAETLGINPAEFAEQYFRFSIMKEQKNLGFNITKLREVMLDRKDMSEIRKNHKVFDSI RTKVYTMMDFVIYRYYIEEDAKVAAANKSLPDNEKSLSEKDIFVINLRGSFNDDQKD ALYYDEANRIWRKLENIMHNIKEFRGNKTREYKKKDAPRLPRILPAGRDVSAFSKLM YALTMFLDGKEINDLLTTLINKFDNIQSFLKVMPLIGVNAKFVEEYAFFKDSAKIADE LRLIKSFARMGEPIADARRAMYIDAIRILGTNLSYDELKALADTFSLDENGNKLKKGK HGMRNFIINNVISNKRFHYLIRYGDPAHLHEIAKNEAVVKFVLGRIADIQKKQGQNG KNQIDRYYETCIGKDKGKSVSEKVDALTKIITGMNYDQFDKKRSVIEDTGRENAERE KFKKIISLYLTVIYHILKNIVNINARYVIGFHCVERDAQLYKEKGYDINLKKLEEKGFS SVTKLCAGIDETAPDKRKDVEKEMAERAKESIDSLESANPKLYANYIKYSDEKKAEE FTRQINREKAKTALNAYLRNTKWNVIIREDLLRIDNKTCTLFANKAVALEVARYVHA
YINDIAEVNSYFQLYHYIMQRIIMNERYEKSSGKVSEYFDAVNDEKKYNDRLLKLLC VPFGYCIPRFKNLSIEALFDRNEAAKFDKEKKKVSGNSGSGPKKKRKVGSAAADYK DDDDK (SEQ ID NO:27) dCasl3d-linker-APOBEC-evoCDAl
[00140] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII
REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSTDAE YVRIHEKLDIYTFKKQFSNNKKSVSHRCYVLFELKRRGERRACFWGYAVNKPQSGT ERGIHAEIF SIRK VEE YLRDNPGQFTIN W YS SWSPC ADCAEKILEWYNQELRGNGHTL KIWVCKLYYEKNARNQIGLWNLRDNGVGLNVMVSEHYQCCRKIFIQSSHNQLNENR WLEKTLKRAEKRRSELSIMFQVKILHTTKSPAV (SEQ ID NO:28) dCas!3d-hAPOBEC3A
[00141] MSPKKKRKVEASIEKKKSFAKGMGVKSTLVSGSKVYMTTFAEGSDARLE KIVEGDSIRSVNEGEAFSAEMADKNAGYKIGNAKFSHPKGYAVVANNPLYTGPVQQ DMLGLKETLEKRYFGESADGNDNICIQVIHNILDIEKILAEYITNAAYAVNNISGLDKD IIGFGKFSTVYTYDEFKDPEHHRAAFNNNDKLINAIKAQYDEFDNFLDNPRLGYFGQ AFFSKEGRNYIINYGNECYDILALLSGLAHWVVANNEEESRISRTWLYNLDKNLDNE YISTLNYLYDRITNELTNSFSKNSAANVNYIAETLGINPAEFAEQYFRFSIMKEQKNLG FNITKLREVMLDRKDMSEIRKNHKVFDSIRTKVYTMMDFVIYRYYIEEDAKVAAAN KSLPDNEKSLSEKDIFVINLRGSFNDDQKDALYYDEANRIWRKLENIMHNIKEFRGN KTREYKKKDAPRLPRILPAGRDVSAFSKLMYALTMFLDGKEINDLLTTLINKFDNIQS FLKVMPLIGVNAKFVEEYAFFKDSAKIADELRLIKSFARMGEPIADARRAMYIDAIRI LGTNLSYDELKALADTFSLDENGNKLKKGKHGMRNFIINNVISNKRFHYLIRYGDPA HLHEIAKNEAVVKFVLGRIADIQKKQGQNGKNQIDRYYETCIGKDKGKSVSEKVDA LTKIITGMNYDQFDKKRSVIEDTGRENAEREKFKKIISLYLTVIYHILKNIVNINARYVI GFHC VERD AQL YKEKGYDINLKKLEEKGF S S VTKLC AGIDETAPDKRKD VEKEMAE RAKESIDSLESANPKLYANYIKYSDEKKAEEFTRQINREKAKTALNAYLRNTKWNVII REDLLRIDNKTCTLFANKAVALEVARYVHAYINDIAEVNSYFQLYHYIMQRIIMNER YEKSSGKVSEYFDAVNDEKKYNDRLLKLLCVPFGYCIPRFKNLSIEALFDRNEAAKF DKEKKKVSGNSGSGPKKKRKVSGGSSGGSSGSETPGTSESATPESSGGSSGGSEASPA SGPRHLMDPHIFTSNFNNGIGRHKTYLCYEVERLDNGTSVKMDQHRGFLHNQAKNL LCGFYGRHAELRFLDLVPSLQLDPAQIYRVTWFISWSPCFSWGCAGEVRAFLQENTH
VRLRIFAARIYDYDPLYKEALQMLRDAGAQVSIMTYDEFKHCWDTFVDHQGCPFQP WDGLDEHSQALSGRLRAILQNQGN (SEQ ID NO:29)
VII. Examples
[00142] The examples provided in the Appendix are included to demonstrate preferred embodiments of the invention. It should be appreciated by those of skill in the art that the techniques disclosed in the examples which follow represent techniques discovered by the inventor to function well in the practice of the invention, and thus can be considered to constitute preferred modes for its practice. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the spirit and scope of the invention.
Example 1 - Programmed RNA Editing with an Evolved Bacterial Adenosine Deaminase
[00143] In this work, the inventors present DECOR (bacterial deaminase-enabled recoding of RNA) as a new programmable RNA-editing platform (FIG. 1A). The inventors focus on deaminases that act on single-stranded nucleic acids, a substrate profile orthogonal to ADARs and ADAR derivatives. DECOR exploits TadASe11, a variant of the E.coli TadA, to deposit A- to-I editing to sites programmed by CRISPR. DECOR has a relaxed editing window, with the most strongly edited A positioned ~13 nucleotides (nt) downstream of the guide RNA-binding site. The inventors demonstrate that DECOR functions robustly in a variety of human cells, delivering up to 53.1% A-to-I conversion to endogenous transcripts. Meanwhile, DECOR impacts 88% fewer off-target sites in the human transcriptome than REPAIRvl, the state-of- the-art CRISPR-based RNA-editing platform. With a prokaryotic deaminase catalyzing the A- to-I transition, DECOR shows an off-target profile orthogonal to REPAIR. Installation of high- fidelity mutations into TadA8e further reduces the off-target effects of DECOR to almost basal level while sustaining on-target activity. The inventors applied DECOR to remove a diseasecausing upstream open reading frame (uORF) in the 5’ untranslated region (UTR) of interferon regulatory factor 6 (IRF6) and rescued translation of the primary ORF (pORF) from 12.3% to 36.5%. As the first ADAR-independent A-to-I editing platform, DECOR complements and expands the current portfolio of programmed RNA-editing technologies.
[00144] To identify effector proteins compatible with RNA editing, the inventors screened rAPOBECl, a cytosine deaminase that accepts both DNA and RNA35; A3 A, an APOB EC family protein highly active in deaminating cytosine in single-stranded DNA (ssDNA)8; evoCDAl, a deaminase evolved for robust deamination of deoxy cytidine36; E. coli tRNA adenosine deaminase TadA37; and TadA derivatives evolved for nucleic acid editing including TadA7.109, TadA8.2O10, and TadASe11. To enable programmability, the inventors chose RfxCasl3d, an RNA-targeting CRISPR system identified from Ruminococcus flavefaciens XPD300238, 39 , hereby referred to as Cast 3d. Effector proteins were fused to dCasl3d, a nuclease-deficient mutant of Casl3d, through a flexible linker (FIG. IB). The inventors appended two nuclear localization signal (NLS) peptides to dCasl3d to maximize its RNA- targeting efficiency38. Both termini of Casl3d are exposed and not intimately involved in RNA engagement based on the crystal structure40. The inventors therefore evaluated RNA-editing efficiency for constructs with effector proteins situated at either terminus of dCasl3d.
[00145] Plasmids expressing dCasl3d fusion proteins, together with plasmids encoding processed guide RNAs (gRNAs), were delivered into human embryonic kidney (HEK) 293 T cells through transient transfection. For on-target activity screening, the inventors designed gRNA for EZEI2, a histone lysine A-methyltransferase (Table 1 and 2). gRNA targeting a sequence absent in the human transcriptome was applied as a negative control (non-targeting gRNA or NT gRNA). The inventors harvested cells 28 h post-transfection and reverse transcribed the EZH2 mRNA into its complementary DNA (cDNA), which was subsequently amplified by polymerase chain reaction (PCR) and analyzed by next-generation sequencing (NGS). As reverse transcriptases pair U with A and I with C, C-to-U and A-to-I edits are detected as C-to-T and A-to-G mutations.
[00146] No editing above background was detected for A3A, consistent with previous findings28. Low levels of C-to-U editing (0.4-1.7%) were observed for evoCDAl at Cl 174, 10 nt downstream of the gRNA binding site (EZH2- \ 135-1164, FIG. 1C and FIGs. 7A-7B). rAPOBECl was more robust than evoCDAl, delivering up to 4.3% C-to-U editing at Cl 166 and Cl 195 (FIG. ID). The evoCDAl -edited C is proceeded by G, while both Cs targeted by rAPOBECl are in AC, highlighting the different contexts preferred by individual deaminases41. The effector orientation in the fusion protein also impacted the editing outcome, as the most strongly edited Cs for rAPOBECl-dCasl3d and dCasl3d-rAPOBECl are 29 nt apart. Little editing (<0.4%) was observed in the absence of the AZ//? -targeting gRNA, confirming that RNA editing was programmed. The basal to modest editing level is consistent with the tight regulation of cytosine editing in RNA42 and the known preference for DNA over RNA by some cytosine deaminases43, 44.
[00147] The tRNA adenosine deaminase TadA acts on the anti-codon loop of Arg tRNA (ACG) in E. coli 5. With its substrate RNA strictly defined, wildtype TadA is unlikely to crosstalk with the cellular mRNA pool, even when guided by CRISPR. Laboratory evolution has largely liberated the substrate restraint of TadA, enabling nucleic acid editing with minimal context dependence9. The inventors therefore hypothesized that evolved TadAs are better suited than wildtype TadA as effector proteins for programmed RNA editing. Indeed, editing was absent in cells expressing dCasl3d fused to wildtype TadA, but was evident with TadA7.10, TadA8.20, and TadA8e fusions, regardless of the effector orientation (FIGs. le, If, and FIGs. 8A-8B). Strongest activity was detected at AZ//2-A I I 79, 15 nt downstream of the gRNA binding site, which was edited to 6.3-13.5%, 16.6-23.3%, and 40.9-52.9% by TadA7.10, TadA8.20, and TadA8e, respectively. The activity increment from TadA7.10 to TadA8.20 and TadA8e was anticipated, as the catalytic function of the latter two variants was improved via further directed evolution10'12. Similar boosted editing activity has been reported for TadA8.20 and TadA8e when targeting DNA. Among different effector proteins and fusion designs, TadA8e-dCasl3d delivered the most robust editing. NLS peptides are essential for the observed activity, as their removal and replacement by nuclear export signal (NES) peptides both abolished editing (FIGs. 9A-9D). The inventors moved forward with this architecture and named the platform bacterial deaminase-enabled recoding of RNA (DECOR, FIG. 1A).
[00148] DECOR consists of two functional modules — TadA8e that catalyzes the A-to-I conversion and the dCasl3d/gRNA complex responsible for target search. The latter function is broadly shared by RNA-targeting CRISPR systems. The inventors therefore hypothesized that DECOR may be compatible with CRISPR systems beyond RfxCasl3d. The inventors tethered TadA8e to the N- or C-terminus of nuclease-deficient LwaCasl3a46, PspCasl3b18, Casl3X. l (also known as Casl3bt3)47, 48, and minimized Casl3X.l (mini-Casl3X.l). In the presence of the cognate EZH2 -targeting gRNA (Table 2), TadA8e-dPspCasl3b, dPspCasl3b- TadA8e, TadA8e-dCasl3X.l, dCasl3X. l-TadA8e, TadA8e-mini-dCasl3X.l, and mini- dCasl3X. l-TadA8e all showed noticeable A-to-I editing (FIGs. lb, 1g, and Ih). Specifically, TadA8e-dPspCasl3b and mini-dCasl3X. l-TadA8e edited EZH2- Al 179 to 7.8% and 10.0%, respectively. Minimal editing was observed when TadA8e was tethered to dLwaCasl3a, indicating that DECOR is less compatible with dLwaCasl3a. These observations are in general agreement with activity levels reported for ADAR-based platforms, where PspCasl3b was demonstrated more potent thanLwaCasl3a18, 49. Nevertheless, none of these CRISPR systems outperform RfxCasl3d for TadA8e-mediated RNA editing.
[00149] To test if the activity levels observed for different CRISPR systems can be extended beyond EZH2, the inventors designed a set of new gRNAs for ACTB (FIGs. 10A-10B, Table 1 and 2). 4(77>-targeting gRNAs yielded results similar to EZH2 -targeting gRNAs, wherein RfxCasl3d delivered the strongest RNA-editing signals, followed by mini-Casl3X. l, PspCasl3b, and Casl3X. l (FIGs. 10A-10B). Interestingly, with similar gRNA placement, strongest editing was observed at Al 179 in EZH2 and A1220 in ACTB regardless of the CRISPR system and effector orientation, indicating that mRNA carries intrinsic determinants to refine site selectivity within the editing window, including but not limited to the sequence context and the secondary structure covering a candidate A site. Collectively, the inventors demonstrate that DECOR facilitates RNA editing when guided by a variety of CRISPR systems. RfxCas 13 d-based DECOR was used in all further characterizations.
[00150] We next tested whether DECOR could deliver programmed RNA editing across the human transcriptome and selected five more loci (Table 1). Both ADAR and TadA support CRISPR-programmed A-to-I editing. The inventors therefore benchmarked side by side REPAIRvl, the state-of-the-art CRISPR-based RNA-editing platform that employs the adenosine deaminase domain of ADAR2 bearing a hyperactivating mutation E488Q (ADARDD-E488Q)18; REPAIRV2, a high-fidelity variant of REPAIRvl that harnesses ADAR2DD(E488Q/T375G)18; REPAIR.tl, constructed using Casl3btl and ADARDD- E488Q48; Casl3d-derived REPAIR systems49; and DECOR across seven target sites.
[00151] To enable activity comparison, the inventors designed DECOR and REPAIR systems to edit the same As — EZH2 -Al 179, ACTB-A122Q, B 4G ALN 7-A958, A7MS-A504, MALAT1-A2A 6, METTL3-A95Q, and 57XZ3-A1402 (Table 2). REPAIRvl functions with spacer lengths ranging from 30 to 70 nt18. However, hybridization length beyond 30 nt increases CRISPR-independent editing and local off-target editing without benefiting on-target activity49, 50. Hence, the inventors chose 30 nt spacers. ADARs deaminate As in dsRNA with a preference for As in A:C mismatches2, 51. REPAIRvl shows high-level editing when the mismatch is close to the center of the spacer18, but screening is recommended to identify the optimal placement of the A:C mismatch. The inventors designed four REPAIRvl gRNAs for each locus with the target A situated at different positions in the protospacer (13, 18, 22, and 26, FIG. 2A and Table 2). Indeed, the inventors did not find a single mismatch position that was constantly favored, and variable editing levels were observed for the four gRNAs at individual target sites (FIG. 2B).
[00152] Both DECOR and REPAIRvl showed minimal A-to-I conversion with the nontargeting gRNA, 0.8-1.8% for DECOR and 0.1-0.7% for REPAIRvl, with the exception of B 4G ALNT 1-^95 for DECOR (17.4%) and A7US-A504 for REPAIRvl (8.4%, FIGs. 2b and 2c). Background editing of B4GALNT1-A95 by DECOR can be partially attributed to nonspecific binding of the control gRNA, as gRNAs targeting EZH2 and KRAS both led to lower levels of editing at B4GALNT1- 95'& (7.5-7.9%, FIG. 11). Nevertheless, the A-to-I editing levels were significantly boosted in the presence of the targeting gRNA, reaching 33.1% for B4GALNT1 with DECOR and 59.7% for KRAS with REPAIRvl . Among the seven target sites, REPAIRvl delivered 0.6-63.5% on-target editing (26.4 ± 19.1%, mean ± s.d.; 37.5 ± 21.0% if only considering the best gRNA design at each target site), whereas 19.7-52.9% editing was attained by DECOR (37.4 ± 10.2%, FIG. 2D). The long non-coding RNA MALAT1 was a particularly challenging substrate for REPAIRvl as all four gRNAs performed poorly at MALAT1- 2A 6 (0.6-4.1%). In contrast, DECOR edited this site robustly, delivering 39.8% A-to-I editing. The difference in editing efficiency may be explained by the intrinsic difficulty in targeting nuclear-localized MALAT1 with cytoplasm-localized REPAIRvl . This observation also highlights the functional complementarity of REPAIRvl and DECOR.
[00153] In addition to the best edited sites, both REPAIRvl and DECOR show moderate editing in regions surrounding the target A sites (FIG. 2C and FIGs. 12A-12D). Shortening the linker between TadA8e and dCasl3d had a modest impact on editing potency and did not result in a narrower editing window (FIGs. 13A-13C). Notably, editing activity persisted even when the 33 amino acid (aa) linker was truncated to two amino acids, indicating considerable flexibility in the protein termini. The inventors dissected DECOR’s context preferences using a high-throughput assay and found that DECOR was more potent towards YA (Y = U and C) (Supplementary Note 1 and FIGs. 14A-14C), a context preference likely inherited from the parent Tad A protein52.
[00154] Both REPAIRv2 and REPAIR.tl conferred lower editing across seven target sites with 3-4 gRNAs designed to vary spacer placement: 10.1 ± 12.5% and 1.2 ± 2.4% considering all target site and gRNA combinations; 17.6 ± 18.0% and 2.3 ± 3.9% considering the best gRNA design at each target site (FIGs. 2a, 2d, and FIGs. 15A-15B). dCasl3d-ADARoD-E488Q was generally more active than ADARoD-E488Q-dCasl3d (16.5 ± 12.4% vs. 5.8 ± 9.1% editing across seven target sites; FIG. 2D); however, these editing levels were substantially lower than those achieved by REPAIRvl and DECOR. Among the five editing platforms evaluated, REPAIRvl and DECOR consistently showed the highest levels of on-target editing. [00155] To assess the function of DECOR in cell models beyond HEK293T, the inventors delivered DECOR targeting EZH2-S 179, A7 AS'-A504, MALAT1-R2A 6, and METTL 3 -A950 to HeLa and A549 cells and quantified RNA-editing activity. While non-targeting DECOR afforded 0.3-0.8% editing (0.5 ± 0.2%), on-target DECOR delivered robust editing — 18.1- 26.9% in HeLa cells (a cervical cancer cell line, 22.7 ± 4.3%; FIG. 2E) and 6.9-19.4% in A549 cells (a lung cancer cell line, 14.4 ± 5.5%; FIG. 2F). The A-to-I conversion rate was boosted 23- to 67-fold in the presence of the targeting gRNA compared to the non-targeting gRNA, again confirming that the observed RNA editing is programmed. In human lung fibroblasts, DECOR edited EZH2-KW 79 to 3.0 ± 0.1%, a level significantly higher than background editing (0.4 ± 0.1%) generated using a non-targeting gRNA (FIG. 2G). Collectively, the inventors demonstrate that DECOR functions broadly in different cell types, including primary cells.
[00156] DECOR showed non-negligible editing with non-targeting gRNA in some cases, which prompted us to interrogate the off-target effects of DECOR on the whole transcriptome and compared side by side with REPAIRvl, REPAIRv2, REPAIR.tl, and dCasl3d-ADARoD- E488Q. All five RNA-editing platforms were assayed with EZH2 -targeting and non-targeting gRNAs in HEK293T cells in duplicate and yielded editing levels consistent with the previous observations (FIGs. 16A). The inventors included two additional gRNAs targeting KRAS and MALAT1 for DECOR to fully benchmark its off-target behavior (FIG. 16B). Untransfected HEK293T cells were employed as a control. The inventors called out A-to-I edits using REDItools, a pipeline developed to profile RNA editing in RNA-sequencing data53. To mask endogenous RNA-editing events and isolate perturbation-related mutations, the inventors discarded sites that are edited to similar levels in transfected and untransfected cells. Specifically, A-to-I mutations that fit the following criteria are identified as off-target edits: 1) sequencing depth > 10 in samples and untransfected controls; 2) A-to-G frequencies in transfected cells significantly higher than those in untransfected controls (Fisher’s exact test P- value < 0.05); 3) persisting in two independent replicates.
[00157] In the presence of the non-targeting and EZH2 -targeting gRNA, the inventors detected 4,044 and 4,004 off-target sites for REPAIRvl with median editing levels of 30.6% (first and third quantile: 23.5% and 39.2%) and 29.8% (first and third quantile: 22.8% and 38.2%), respectively (FIGs. 3a, 3b, and FIG. 16C). Many off-target editing sites appeared to be gRNA-independent, as 2,426 off-target sites are shared between non-targeting and EZH2- targeting REPAIRvl (FIG. 16D). The vast majority of off-target edits are de novo — only 0.07- 0.12% of the off-target sites showed non-zero editing in untreated cells (FIGs. 16E and 1 If).
[00158] REPAIRv2, REPAIR.tl, and dCasl3d-ADARDD-E488Q reduced the number of off- target sites to 11-29, 1,992-2,139, and 2,881-2,929 (FIG. 16C), respectively, consistent with their modest on-target activity. Note that the absolute number of off-target sites can vary with different sequencing depths and analysis pipelines. The numbers the inventors present here may not be directly comparable with numbers reported in other studies. Nevertheless, the results are in general agreement with previous reports that with an embedded hyperactive ADAR variant, REPAIRvl shows considerable transcriptome-wide off-target effects in human cells18212250.
[00159] We detected in total 888, 840, 896 and 2,067 off-target sites for non-targeting, EZH2 -targeting, AVMS'-targeting, and Afd/.d 77-targeting DECOR, respectively (FIG. 3 A and FIG. 16C). Given that DECOR performed consistently across four different gRNAs, the inventors focused the subsequent analyses on off-target editing generated using non-targeting and EZH2 -targeting gRNAs. The number of off-target sites hit by DECOR is -88% lower than that of REPAIRvl. The median editing level detected at these off-target sites is 20.2-20.6% (first and third quantile: 16.2-16.6% and 26.1-26.4%, FIG. 3B), also markedly lower than that generated by REPAIRvl . Similar to what was observed for REPAIRvl, a large portion of these off-target sites (345) overlap between non-targeting and EZH2 -targeting DECOR and showed consistent A-to-I editing levels (FIG. 3C and FIG. 16G). DECOR generated DNA off-target editing 56.4% lower than that of ABE8e (Supplementary Note 2 and FIGs. 17A-17J)1154'57, a DNA base editor derived from TadA8e, suggesting that DECOR, while capable of acting on DNA, maintains a tightly controlled DNA off-target profile. Taken together, these results suggest that DECOR is among the cleanest CRISPR-based RNA-editing systems.
[00160] We overlapped the off-target sites hit by DECOR and REPAIRvl with non-targeting gRNAs and found 10 common sites (FIGs. 3D and 3E). The almost mutually exclusive off- target profiles for DECOR and REPAIRvl indicate that the deaminase, rather than the CRISPR complex, dictates how these RNA-editing systems interplay with the human transcriptome. TadA8e, a variant of a bacterial tRNA adenosine deaminase, and ADAR2, a mammalian adenosine deaminase acting on dsRNA, may naturally accept sites of distinct features. Indeed, off-target editing sites for DECOR are enriched in YA sequences, UA being most favored, while REPAIRvl preferred G and C following the target A (FIGs. 3F and 3G).
[00161] A handful of mutations were reported to reduce the off-target editing capacity of evolved TadAs58'60. The inventors chose two of these mutations, F148A and V106G, to mitigate the off-target effects of DECOR and named these variants high-fidelity DECOR (hfDECOR). hfDECOR-F148A and hfDECOR-V106G generated 23.1 ± 10.5% and 24.0 ± 9.2% A-to-I editing across 18 endogenous RNAs, respectively, slightly lower than what was observed for DECOR (35.4 ± 12.1%; FIG. 4A and FIG. 18A).
[00162] When complexed with the non-targeting gRNA, hfDECOR showed drastically reduced off-target effects — 47 and 126 off-target sites for hfDECOR-F148A and hfDECOR- V106G, respectively (FIG. 4B, FIG. 18B and 18C), 94.7% and 85.8% lower than DECOR. The median editing levels observed at these off-target sites were 17.0% for hfDECOR-F148A and 19.5% for hfDECOR-V106G (FIG. 4C and FIG. 18D). The majority of the off-target sites impacted by hfDECOR overlap with those of DECOR, which account for 33 of the 47 off- target sites observed with hfDECOR-F148A, and 73 of the 126 off-target sites hit by hfDECOR-V106G (FIG. 4D and FIG. 18E). Off-target profiles of hfDECOR-F148A and REPAIRvl are orthogonal, while 3 off-target sites are shared between hfDECOR-V106G and REPAIRvl (FIG. 4E and FIG. 18F). The sequence contexts hosting off-target edits for hfDECOR are again skewed towards YA (FIG. 4F and FIG. 18G). These results, taken together, show that hfDECOR delivers robust and clean RNA editing.
[00163] With the RNA-editing properties of DECOR characterized, the inventors next attempted to leverage DECOR to recode GFP mRNA that harbors a premature stop codon (FIG. 5A). GFP-W58X was overexpressed in HEK293T cells with RFP serving as a transfection control. A-to-I editing mediated by DECOR changes the UAG codon to UIG, which is read as UGG by the ribosome for tryptophan (W) installation, restoring GFP fluorescence. The inventors designed four gRNAs that place UAG 2-32 nt downstream of the protospacer (FIG. 5A). Among them, gRNA b that binds 11 to 40 nt upstream of the UAG codon led to strong recovery of GFP fluorescence (FIG. 5B). Minimal GFP signals were detected in cells expressing non-targeting DECOR. The editing level at the UAG codon was quantified by high- throughput sequencing to be 7.0% (FIG. 5C), confirming that the observed GFP restoration was indeed a consequence of programmed RNA editing.
[00164] Lastly, The inventors explored potential clinical applications of DECOR by targeting a disease-causing uORF in the 5’ UTR of IRFt Mutations in IRF6 are the primary cause of Van der Woude syndrome (VWS), an orofacial clefting disorder that affects 1-3 newborns in every 100,00062. In addition to coding mutations that may directly impact protein function, a subpopulation of VWS patients carry an A-48T mutation that creates a de novo start codon in the 5’ UTR of IRF6. The new start codon cannot produce IRF6 protein owing to a shifted coding frame. Instead, translation of the uORF suppresses initiation at the native start site, leading to insufficient IRF6 expression (FIG. 5D)63. No existing RNA-editing technology, including DECOR, can reverse an A-to-U transversion. Alternatively, the inventors propose to edit A-49 to mute the uORF, with which the disease-causing start codon AUG will be converted into IUG, a much less potent translation initiation signal. To quantify the impact of the 5’ UTR on expression of the pORF, the inventors adapted a customized dual-luciferase reporter assay in which the IRF65’ UTR was employed to drive expression of firefly luciferase (FIG. 5D)63.
[00165] We confirmed that the disease-causing mutation A-48T indeed reduced expression of the pORF by 87.7% (FIG. 5E). Interestingly, mutation of A-49 into G on the A-48T background boosted expression drastically to 3.3-fold of the level observed for the healthy UTR (FIG. 5E). These results support the editing strategy as A-49I will likely rescue IRF6 expression by removing the pathogenic uORF. The inventors screened four gRNAs tiled 1-45 nt upstream of the uORF (FIG. 5D) and found that gRNA d, targeting IRF6 -94 to -65, enabled 3.0-fold gain of expression for firefly luciferase (FIG. 5F). The inventors quantified the editing level at A-49 again by high-throughput sequencing and detected 10.2% A-to-I editing (FIG. 5g). While the diseased UTR confers a pORF expression strength 12.3% of the healthy UTR, RNA editing boosts the expression level to 36.5%, an impact that can likely translate in the clinic. To better mimic editing of the IRF6 transcript in diseased cells, the inventors replaced the herpes simplex virus (HSV) thymidine kinase (TK) promoter with the endogenous IRF6 promoter (FIG. 5D). With the native promoter, the inventors detected 31.2% editing at A-49 and a 2.1-fold increase in protein expression (FIGs. 5H and 51). Collectively, DECOR offers an attractive strategy to treat VWS in patients carrying the 7KF6-A-48T mutation.
[00166] We present in this work the development of DECOR, a new CRISPR-based RNA- editing platform. As the first ADAR-independent A-to-I editing system, DECOR employs an evolved bacterial tRNA adenosine deaminase to edit As flagged by CRISPR. The inventors demonstrate that DECOR functions broadly in human cells, delivering A-to-I editing at levels comparable to ADAR-mediated RNA editing. Importantly, DECOR shows significantly lower transcriptome-wide off-target effects than RNA-editing systems with overexpressed ADARs, and an orthogonal off-target profile. High-fidelity DECOR further reduces the off-target effects close to basal level. [00167] DECOR showcases that CRISPR-guided RNA editing can be achieved using deaminases acting on single-stranded nucleic acids, which opens up exciting opportunities for the development of programmed RNA-editing machinery. Although the several cytosine deaminases the inventors evaluated in this work delivered low to moderate editing, the inventors envision robust C-to-U editing may be attained by other families of cytosine deaminases64'66, such as the DYW-deaminase domains responsible for hundreds of C-to-U editing sites in plant mitochondria and chloroplasts67, or evolved deaminase variants. TadA has recently been engineered to deaminate C instead of A68'70, with which DECOR can readily expand its territory into C-to-U editing. Although deaminase evolution has primarily focused on DNA-editing applications9'11, 13, 36, similar strategies can be employed to isolate deaminases tailored for RNA editing71. Additionally, RNA editing is particularly diverse and abundant in symbiotic dinoflagellates72. For example, RNA editing occurs in 1.6% of all nuclear-encoded genes of Symbiodinium microadriaticum, with all 12 types of base substitutions observed73. The involved nucleobase conversion enzymes, although poorly characterized at the moment, may offer an ample pool of effector proteins to further expand the scope of programmed RNA editing.
[00168] DECOR has a relaxed editing window with the most strongly edited A jointly defined by the position of the gRNA, the context preference of TadA8e, and possibly intrinsic properties of the target mRNA. With all tested sites considered, the optimal distance between the target A and the gRNA binding site is 13 ± 10 nt. As a start point for design, the inventors recommend to place the target A 1-30 nt downstream of the gRNA-binding site and prioritize YA targets. gRNA screening can be carried out to secure an optimal editing profile, as demonstrated by correcting a premature stop codon in GFP.
[00169] We frequently observe minor editing at A sites surrounding the target A, which can be undesired if these bystander edits lead to deleterious consequences. Current DECOR designs are therefore less suited for applications where a clean, single A-to-I edit needs to be installed. Nonetheless, the inventors envision DECOR will be particularly useful for editing coding sequences, perturbing protein-binding sites, altering splicing patterns, and eliminating undesired features in UTRs when multiple edits in a short RNA region are needed or when low-level bystander editing is not a concern, as demonstrated by removal of a disease-causing uORF in IRF6. Moreover, many deaminases show natural selectivity for sequence contexts65, 66; DECOR may leverage these context-specific deaminases to achieve clean and precise editing patterns. [00170] DECOR may demonstrate its ultimate clinical utility for editing transcripts that function transiently in development. For example, temporary upregulation of IRF6, as enabled by DECOR-mediated editing of IRF6-A- 9 at the early stage of embryonic development, may be sufficient to prevent VWS without changing the patient’s genome. In summary, DECOR is a new programmable RNA-editing platform with features distinct from and complementary to existing RNA-editing technologies.
Example 2 - Materials and Methods for Aspects Herein
Plasmid construction
[00171] Oligonucleotides were ordered from Integrated DNA Technologies (IDT). Sources of plasmid and synthetic DNA sequences are listed in Table 3. DNA sequences of signaling peptides, linker peptides, and //U6 5’ UTR are listed in Table 4. PCR fragments were generated using Phusion U DNA polymerase (Thermo Scientific, F555L) and assembled by USER assembly (New England Biolabs, M5505L) following manufactures’ protocols. DH5-alpha competent cells (New England Biolabs, C2992I) were used for cloning. The sequences of all plasmids were confirmed by Sanger sequencing (Azenta Life Science). Cas 13 -deaminase fusions and GFP/RFP reporters are expressed under a cytomegalovirus (CMV) promotor. Firefly luciferase reporters are expressed with a HSV-TK promoter or the native IRF6 promoter. Guide RNAs are expressed with a U6 promoter. Key plasmids generated in this work will be deposited to Addgene.
Cell culture
[00172] HEK293T, HeLa, A549 cell lines and IMR-90 human primary fibroblasts were cultured in Dulbecco’s Modified Eagle Medium with high glucose, sodium pyruvate, and L- Glutamine (Coming) supplemented with 10% fetal bovine serum (Coming). Cells were maintained in 5% CO2 and passaged regularly at confluence of 80%.
Editing of endogenous RNA
[00173] For endogenous RNA editing in HEK293T, HeLa and A549 cells, transfection was performed with jetOPTIMUS DNA transfection reagent (Polyplus, 101000006) following the manufacturer’s instructions. Briefly, cells were seeded 16 h prior to transfection to ensure -70% confluence at the time of transfection. To each well of a 96-well plate, 40 ng of editor plasmid and 80 ng of gRNA plasmid were delivered with 0.12 pL of jetOPTIMUS DNA transfection reagent in 12.5 pL of jetOPTIMUS buffer. Cells were incubated for 28 h followed by RNA extraction using either a QuickExtract RNA extraction kit (Lucigen) or a Direct-zol RNA Microprep Kit (Zymo Research) per the manufactures’ instructions. RNA was reverse transcribed using GoScript reverse transcriptase (Promega). The resulting cDNA was subjected to two rounds of PCR to add Illumina adaptors and barcodes. The pooled library was gel purified and sequenced on an Illumina MiSeq instrument. The sequences of primers used for cDNA amplification are listed in Table 5.
[00174] For RNA editing in IMR-90 human primary fibroblasts, DNA was delivered by nucleofection (Lonza, V4XC-1032) with a 4D-Nucleofector. Briefly, 2 x 105 cells were transfected with program CM- 120 in a 20 pL Nucleocuvette containing 1000 ng of plasmid DNA expressing both the editor and the gRNA. Cells were resuspended in 500 pL of DMEM and incubated in a 24-well plate for an additional 28 h. Total RNA was extracted using a Direct- zol RNA Microprep Kit (Zymo Research) followed by reverse transcription and deep sequencing as described above.
Profiling the context preferences of DECOR
[00175] HEK293T cells were transfected with jetOPTIMUS DNA transfection reagent. To each well of a 96-well plate, 30 ng of editor plasmid, 60 ng of gRNA plasmid, and 10 ng of GFP plasmid containing NAN motifs were co-delivered with 0.1 pL of jetOPTIMUS DNA transfection reagent in 12.5 pL of jetOPTIMUS buffer. Cells were incubated for an additional 28 h post-transfection. Total RNA was extracted with TRIzol (Invitrogen, 15596026) followed by DNase I (Lucigen, QER09015) digestion. Reverse transcription, cDNA amplification, and high-throughput sequencing followed the same protocol as described above.
Correction of a premature stop codon in GFP
[00176] HEK293T cells were transfected with 10 ng of GFP-W58X plasmid, 10 ng of RFP plasmid, 30 ng of editor plasmid, and 60 ng of gRNA plasmid using 0.12 pL of jetOPTIMUS DNA transfection reagent in 12.5 pL of jetOPTIMUS buffer. After 28 h of incubation, cells were imaged using an Olympus BX53 microscope. Total RNA was extracted from the same cells with TRIzol followed by DNase I digestion. RNA editing was quantified by high- throughput sequencing as described above. Editing 0URF65’ UTR
[00177] To each well of a 96-well plate pre-seeded with HEK293T cells, 10 ng of firefly luciferase plasmid and 10 ng of Renilla luciferase plasmid (used as control to normalize firefly luciferase activity) were co-transfected with 30 ng of editor and 60 ng of gRNA plasmids using 0.12 pL of jetOPTIMUS DNA transfection reagent in 12.5 pL of jetOPTIMUS buffer. After incubation for 28 h, luciferase activity was measured with the Dual-Glo Luciferase Assay System (Promega) using a BioTek Synergy H4 Hybrid plate reader. To quantify RNA editing, total RNA was extracted by TRIzol and treated with DNase I. High-throughput sequencing was carried out as described above.
RNA sequencing
[00178] Transcriptome wide RNA off target editing experiments were carried out in HEK293T cells. To each well in a 24-well plate, 200 ng of editor plasmid and 400 ng of gRNA plasmid were delivered with 0.6 pL of jetOPTIMUS DNA transfection reagent in 62.5 pL of jetOPTIMUS buffer. Cells were incubated for 28 h followed by total RNA extraction with TRIzol (Invitrogen). mRNA was enriched by a NEBNext Poly(A) mRNA Magnetic Isolation Module (New England Biolabs, E7490L) and prepared for sequencing using an NEBNext Ultra II Directional RNA Library Prep Kit for Illumina (New England Biolabs, E7760L). The libraries were sequenced on an Illumina NovaSeq instrument (single-end R1 121) or a NextSeq 2000 instrument (paired-end R1 25 R2 97 or paired-end R1 62 R2 160).
Sequencing data analysis
[00179] For targeted sequencing, MiSeq reads were aligned to reference sequences with BWA-MEM74 and sorted using Samtools 1.1375. Editing rates at given positions were calculated as G-reads/total-reads. For whole-transcriptome RNA-seq, quality control was conducted using Fastqc 0.11.5 (http://www.bioinformatics.babraham.ac.uk/projects/fastqc/). To account for variables across different batches of sequencing experiments, only R2 reads of paired-end sequencing data were used for analysis. All reads were trimmed to the same length (97 nt). Reads were aligned to the hgl9 (GRCh37) reference genome using STAR 2.6.1b76. Reads that could be aligned to more than one genomic location were discarded. PCR duplicates were removed using Samtools 1.13. The deduplicated bam files were randomly downsampled to 4.5 million reads to ensure the same sequencing depth across samples. A-to-G mutations were extracted from aligned BAMs by REDItools2.0. Common single nucleotide polymorphisms (SNPs) were excluded from downstream analysis. The inventors only consider A sites of sequencing depths > 10 in both treated and non-treated samples. Fisher’s exact tests were performed using G reads and A reads detected in treated and non-treated cells. Sites with a P-value greater than 0.05 were removed as these are A sites edited to similar levels regardless of the treatment, i.e., sites hosting naturally occurring A-to-I editing. Off-target sites were considered significant if detected in both replicates. Mean editing rates of two replicates were used for all plots involving editing frequencies. Overlap of off-target sites among different samples was analyzed by BioVenn77. Sequence motif analysis was done using WebLogo 378.
Statistical analysis
[00180] Unpaired Student's t-tests (two tailed) were used to compare samples, n.s. P > 0.05; * P < 0.05; ** p < 0.01; *** P < 0.001; **** p < 0.0001. Investigators were not blinded to experimental conditions or outcome assessments.
Data availability
[00181] All RNA-seq data have been deposited to the NCBI Gene Expression Omnibus (GEO) and can be accessed through GEO series accession number GSE251875.
Example 3 - DECOR favors A proceeded by a pyrimidine
[00182] The physiological function of E. coli TadA is to deaminate the A at the wobble position of Arg tRNA(ACG)79. Proceeded by a U-turn, this A is situated in a UAC motif. Both -1 U and +1 C interact with TadA during substrate engagement80. The tertiary structure of the anti-codon loop was proposed both essential and sufficient for TadA-catalyzed adenosine deamination81. TadA8e is a hyperactive deaminase evolved from E. coli TadA. Although the context restriction has been largely eased through multiple rounds of directed evolution, the preference for YA (Y = U and C) persists in TadA8e82. The inventors suspected that DECOR might have inherited such context preferences for RNA editing.
[00183] To dissect DECOR’s context preferences, the inventors installed an NAN motif (N = A, C, G, and U) in GFP mRNA (position 172-174) and delivered 16 plasmids encoding all possible combinations into HEK293T cells (FIG. 14A). DECOR paired with gRNA targeting GFP-132-161 was co-delivered. A-to-I editing in NAN was quantified by RT-PCR followed by NGS. Among different sequence contexts, UAN displayed the highest level of editing (6.7- 10.6%, FIGs. 14B and 14C). The second most favored sequence category appeared to be CAN (2.2-3.6%). Editing at RAN (0.9- 1.4%, R = A or G), albeit higher than background defined by the non-targeting gRNA, is less potent (FIGs. 14B and 14C). Such a trend suggests that DECOR has inherited the context preference from TadA and favors A following a pyrimidine instead of a purine. The overall lower editing level observed in GFP mRNA compared to endogenous mRNA may be a result of GFP overexpression and/or variations in CRISPR- targeting efficiency.
Example 4 - DECOR has a tightly controlled off-target profile in the human genome
[00184] While Casl3d specifically targets RNA, TadA8e can act on both RNA and DNA, leading to potential off-target A:T-to-G:C editing in the genome. To assess CRISPR- independent off-target effects of DECOR in DNA, the inventors adapted an orthogonal R-loop assay83'86 originally developed to evaluate the fidelity of DNA base editors. To carry out the assay, DECOR targeting EZH2-KW19 or ABE8e targeting the HEK4 site was co-delivered with catalytically deficient Staphylococcus aureus Cas9 (dSaCas9) and a cognate Sa gRNA into HEK293T cells. The dSaCas9 complex forms a constant R loop susceptible to TadA8e- mediated deamination; editing at the R loop therefore serves as a surrogate for potential genome-wide off-target editing (FIGs. 17A-17J). Both DECOR and ABE8e showed anticipated on-target RNA and DNA editing, respectively, in the presence of these orthogonal R loops.
[00185] While both editors generated similar editing patterns at six R loops presented by the dSaCas9 complex, editing by DECOR was consistently lower than that of ABE8e at individual A sites (11.5 ± 6.6% for ABE8e and 4.8 ± 3.0% for DECOR across 12 A sites edited to greater than 1%; mean ± s.d.). The inventors further summed all editing events observed at individual R loops and found that editing generated by DECOR was 36.4-73.3% lower than that of ABE8e (56.4 ± 12.8% lower across six R loops). Collectively, DECOR maintains a tightly controlled off-target profile in the human genome. The lower DNA off-target editing observed with DECOR may stem from laboratory evolution of TadA8e as part of the ABE architecture87, which prompts specialization of TadA8e for robust deamination within the ABE context. This optimal DNA deamination is not expected to extend to DECOR, where TadA8e is fused to an RNA-targeting CRISPR protein.
Tables
[00186] Table 1. NCBI accession numbers for target cellular RNAs.
[00187] Table 2. gRNAs used in this work.
[00188] Table 3 Sources of plasmids and synthetic DNA.
[00189] Table 4 DNA sequences of signal peptides, linker peptides, and IRF65’ UTR.
[00190] Table 5 Primers for preparing high-throughput sequencing libraries.
[00191] Table 6
* * *
[00192] All of the methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the compositions and methods of this invention have been described in terms of preferred embodiments, it will be apparent to those of skill in the art that variations may be applied to the methods and in the steps or in the sequence of steps of the method described herein without departing from the concept, spirit and scope of the invention. More specifically, it will be apparent that certain agents which are both chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the invention as defined by the appended claims. REFERENCES
The references cited throughout the disclosure, to the extent that they provide exemplary procedural or other details supplementary to those set forth herein, are specifically incorporated herein by reference.
1. Gott, J.M. & Emeson, R.B. Functions and mechanisms of RNA editing. Annu. Rev. Genet. 34, 499-531 (2000).
2. Bass, B.L. RNA editing by adenosine deaminases that act on RNA. Annu. Rev. Biochem. 71, 817-846 (2002).
3. Porto, E.M., Komor, A.C., Slaymaker, I.M. & Yeo, G.W. Base editing: advances and therapeutic opportunities. Nat. Rev. Drug. Discov. 19, 839-859 (2020).
4. Khosravi, H.M. & Jantsch, M.F. Site-directed RNA editing: recent advances and open challenges. RNA Biol. 18, 1-10 (2021).
5. Pickar-Oliver, A. & Gersbach, C.A. The next generation of CRISPR-Cas technologies and applications. Nat. Rev. Mol. Cell Biol. 20, 490-507 (2019).
6. Komor, A.C., Kim, Y.B., Packer, M.S., Zuris, J. A. & Liu, D.R. Programmable editing of a target base in genomic DNA without double-stranded DNA cleavage. Nature 533, 420- 424 (2016).
7. Nishida, K. et al. Targeted nucleotide editing using hybrid prokaryotic and vertebrate adaptive immune systems. Science 353, DOI: 10.1126/science.aaf8729 (2016).
8. Gehrke, J.M. et al. An APOBEC3A-Cas9 base editor with minimized bystander and off-target activities. Nat. BiotechnoL 36, 977-982 (2018).
9. Gaudelli, N.M. et al. Programmable base editing of A*T to G*C in genomic DNA without DNA cleavage. Nature 551, 464-471 (2017).
10. Gaudelli, N.M. et al. Directed evolution of adenine base editors with increased activity and therapeutic application. Nat. BiotechnoL 38, 892-900 (2020).
11. Richter, M.F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. BiotechnoL 38, 883-891 (2020).
12. Lapinaite, A. et al. DNA capture by a CRISPR-Cas9-guided adenine base editor. Science 369, 566-571 (2020). 13. Xiao, Y.L., Wu, Y. & Tang, W. Directed evolution of an adenine base editor with increased activity and context compatibility. Nat. BiotechnoL, in press.
14. Anzalone, A. V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149-157 (2019).
15. Grunewald, J. et al. Engineered CRISPR prime editors with compact, untethered reverse transcriptases. Nat. BiotechnoL (2022).
16. Liu, B. et al. Targeted genome editing with a DNA-dependent DNA polymerase and exogenous DNA-containing templates. Nat. BiotechnoL (2023).
17. da Silva, J.F. et al. Click editing enables programmable genome writing using DNA polymerases and HUH endonucleases. bioRxiv (2023).
18. Cox, D.B.T. et al. RNA editing with CRISPR-Casl3. Science 358, 1019-1027 (2017).
19. Liu, Y. et al. REPAIRx, a specific yet highly efficient programmable A > I RNA base editor. EMBO J. 39, el04748 (2020).
20. Merkle, T. et al. Precise RNA editing by recruiting endogenous ADARs with antisense oligonucleotides. Nat. BiotechnoL 37, 133-138 (2019).
21. Qu, L. et al. Programmable RNA editing by recruiting endogenous ADAR using engineered RNAs. Nat. BiotechnoL 37, 1059-1069 (2019).
22. Katrekar, D. et al. In vivo RNA editing of point mutations via RNA-guided adenosine deaminases. Nat. Methods 16, 239-242 (2019).
23. Han, W. et al. Programmable RNA base editing with a single gRNA-free enzyme. Nucleic Acids Res. 50, 9580-9595 (2022).
24. Reautschnig, P. et al. CLUSTER guide RNAs enable precise and efficient RNA editing with endogenous ADAR enzymes in vivo. Nat. BiotechnoL 40, 759-768 (2022).
25. Yi, Z. et al. Engineered circular ADAR-recruiting RNAs increase the efficiency and fidelity of RNA editing in vitro and in vivo. Nat. BiotechnoL 40, 946-955 (2022).
26. Katrekar, D. et al. Efficient in vitro and in vivo RNA editing via recruitment of endogenous ADARs using circular guide RNAs. Nat. BiotechnoL 40, 938-945 (2022).
27. Abudayyeh, O.O. et al. A cytosine deaminase for programmable single-base RNA editing. Science 365, 382-386 (2019). 28. Huang, X. et al. Programmable C-to-U RNA editing using the human APOBEC3A deaminase. EMBO J. 39, el04741 (2020).
29. Hartner, J.C. et al. Liver disintegration in the mouse embryo caused by deficiency in the RNA-editing enzyme AD ARI. J. Biol. Chem. 279, 4894-4902 (2004).
30. Uhlen, M. et al. Proteomics. Tissue-based map of the human proteome. Science 347, 1260419 (2015).
31. Savva, Y.A., Rieder, L.E. & Reenan, R.A. The ADAR protein family. Genome Biol. 13, 252 (2012).
32. Qian, Y. et al. Programmable RNA sensing for cell monitoring and manipulation. Nature 610, 713-721 (2022).
33. Walkley, C.R. & Li, J.B. Rewriting the transcriptome: adenosine-to-inosine RNA editing by ADARs. Genome Biol. 18, 205 (2017).
34. Vallecillo- Viejo, I.C., Liscovitch-Brauer, N., Montiel-Gonzalez, M.F., Eisenberg, E. & Rosenthal, J. J.C. Abundant off-target edits from site-directed RNA editing can be reduced by nuclear localization of the editing enzyme. RNA Biol. 15, 104-114 (2018).
35. Harris, R.S., Petersen-Mahrt, S.K. & Neuberger, M.S. RNA editing enzyme APOBEC1 and some of its homologs can act as DNA mutators. Mol. Cell 10, 1247-1253 (2002).
36. Thuronyi, B.W. et al. Continuous evolution of base editors with expanded target compatibility and improved activity. Nat. BiotechnoL 37, 1070-1079 (2019).
37. Gerber, A.P. & Keller, W. An adenosine deaminase that generates inosine at the wobble position of tRNAs. Science 286, 1146-1149 (1999).
38. Konermann, S. et al. Transcriptome engineering with RNA-targeting type VLD CRISPR effectors. Cell 173, 665-676 (2018).
39. Wei, J. et al. Deep learning and CRISPR-Casl3d ortholog discovery for optimized RNA targeting. Cell Syst (2023).
40. Zhang, C. et al. Structural basis for the RNA-guided ribonuclease activity of CRISPR- Casl3d. Ce/Z 175, 212-223 (2018).
41. Conticello, S.G. The AID/APOBEC family of nucleic acid mutators. Genome Biol. 9, 229 (2008). 42. Wolfe, A.D., Li, S., Goedderz, C. & Chen, X.S. The structure of APOBEC1 and insights into its RNA and DNA substrate selectivity. NAR Cancer 2, zcaa027 (2020).
43. Nabel, C.S., Lee, J.W., Wang, L.C. & Kohli, R.M. Nucleic acid determinants for selective deamination of DNA over RNA by activation-induced deaminase. Proc. Natl. Acad. Sci. U. S. A. 110, 14225-14230 (2013).
44. Barka, A. et al. The base-editing enzyme APOBEC3 A catalyzes cytosine deamination in RNA with low proficiency and high selectivity. ACS Chem. Biol. 17, 629-636 (2022).
45. Wolf, J., Gerber, A.P. & Keller, W. TadA, an essential tRNA-specific adenosine deaminase from Escherichia coli. EMBO J. 21, 3841-3851 (2002).
46. Abudayyeh, O.O. et al. C2c2 is a single-component programmable RNA-guided RNA- targeting CRISPR effector. Science 353, DOI: 10.1126/science.aaf5573 (2016).
47. Xu, C. et al. Programmable RNA editing with compact CRISPR-Casl3 systems from uncultivated microbes. Nat. Methods 18, 499-506 (2021).
48. Kannan, S. et al. Compact RNA editors with small Cast 3 proteins. Nat. BiotechnoL 40, 194-197 (2022).
49. Marina, R.J., Brannan, K.W., Dong, K.D., Yee, B.A. & Yeo, G.W. Evaluation of engineered CRISPR-Cas-mediated systems for site-specific RNA editing. Cell Rep. 33, 108350 (2020).
50. Vogel, P. et al. Efficient and precise editing of endogenous transcripts with SNAP- tagged ADARs. Nat. Methods 15, 535-538 (2018).
51. Wong, S.K., Sato, S. & Lazinski, D.W. Substrate recognition by AD ARI and ADAR2. RNA 7, 846-858 (2001).
52. Song, M. et al. Sequence-specific prediction of the efficiencies of adenine and cytosine base editors. Nat. BiotechnoL 38, 1037-1043 (2020).
53. Lo Giudice, C., Tangaro, M.A., Pesole, G. & Picardi, E. Investigating RNA editing in deep transcriptome datasets with REDItools and REDIportal. Nat. Protoc. 15, 1098-1131 (2020).
54. Doman, J.L., Raguram, A., Newby, G.A. & Liu, D.R. Evaluation and minimization of Cas9-independent off-target DNA editing by cytosine base editors. Nat. BiotechnoL 38, 620- 628 (2020). 55. Yu, Y. et al. Cytosine base editors with minimized unguided DNA and RNA off-target events and high on-target activity. Nat. Commun. 11, 2052 (2020).
56. Wang, L. et al. Eliminating base-editor-induced genome-wide and transcriptome-wide off-target mutations. Nat. Cell Biol. 23, 552-563 (2021).
57. Jin, S. et al. Rationally Designed APOBEC3B Cytosine Base Editors with Improved Specificity. Mol. Cell 79, 728-740 e726 (2020).
58. Zhou, C. et al. Off-target RNA mutation induced by DNA base editing and its elimination by mutagenesis. Nature 571, 275-278 (2019).
59. Rees, H.A., Wilson, C., Doman, J.L. & Liu, D.R. Analysis and minimization of cellular RNA editing by DNA adenine base editors. Sci. Adv. 5, eaax5717 (2019).
60. Grunewald, J. et al. CRISPR DNA base editors with reduced RNA off-target and selfediting activities. Nat. Biotechnol. 37, 1041-1048 (2019).
61. Kondo, S. et al. Mutations in IRF6 cause Van der Woude and popliteal pterygium syndromes. Nat. Genet. 32, 285-289 (2002).
62. de Lima, R.L. et al. Prevalence and nonrandom distribution of exonic mutations in interferon regulatory factor 6 in 307 families with Van der Woude syndrome and 37 families with popliteal pterygium syndrome. Genet. Med. 11, 241-247 (2009).
63. Calvo, S.E., Pagliarini, D.J. & Mootha, V.K. Upstream open reading frames cause widespread reduction of protein expression and are polymorphic among humans. Proc. Natl. Acad. Sci. U. S. A. 106, 7507-7512 (2009).
64. Krishnan, A., Iyer, L.M., Holland, S.J., Boehm, T. & Aravind, L. Diversification of AID/APOBEC-like deaminases in metazoa: multiplicity of clades and widespread roles in immunity. Proc. Natl. Acad. Sci. U. S. A. 115, E3201-E3210 (2018).
65. Vaisvila, R. et al. Mol. Cell, accepted.
66. Huang, J. et al. Discovery of deaminase functions by structure-based protein clustering. Ce/Z 186, 3182-3195 e3114 (2023).
67. Wagoner, J. A., Sun, T., Lin, L. & Hanson, M.R. Cytidine deaminase motifs within the DYW domain of two pentatricopeptide repeat-containing proteins are required for site-specific chloroplast RNA editing. J. Biol. Chem. 290, 2957-2968 (2015). 68. Chen, L. et al. Re-engineering the adenine deaminase TadA-8e for efficient and specific CRISPR-based cytosine base editing. Nat. BiotechnoL 41, 663-672 (2023).
69. Neugebauer, M.E. et al. Evolution of an adenine base editor into a small, efficient cytosine base editor with low off-target activity. Nat. BiotechnoL 41, 673-685 (2023).
70. Lam, D.K. et al. Improved cytosine base editors generated from TadA variants. Nat. BiotechnoL 41, 686-697 (2023).
71. Lin, Y. et al. RNA molecular recording with an engineered RNA deaminase. Nat. Methods 20, 1887-1899 (2023).
72. Mungpakdee, S. et al. Massive gene transfer and extensive RNA editing of a symbiotic dinoflagellate plastid genome. Genome Biol. Evol. 6, 1408-1422 (2014).
73. Liew, Y.J., Li, Y., Baumgarten, S., Voolstra, C.R. & Aranda, M. Condition-specific RNA editing in the coral symbiont Symbiodinium microadriaticum. PLoS Genet. 13, el006619 (2017).
74. Li, H. & Durbin, R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25, 1754-1760 (2009).
75. Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10 (2021).
76. Dobin, A. et al. STAR: ultrafast universal RNA-seq aligner. Bioinformatics 29, 15-21 (2013).
77. Hulsen, T., de Vlieg, J. & Alkema, W. BioVenn - a web application for the comparison and visualization of biological lists using area-proportional Venn diagrams. BMC Genomics 9, 488 (2008).
78. Crooks, G.E., Hon, G., Chandonia, J.M. & Brenner, S.E. WebLogo: a sequence logo generator. Genome Res. 14, 1188-1190 (2004).
79. Wolf, J., Gerber, A.P. & Keller, W. TadA, an essential tRNA-specific adenosine deaminase from Escherichia coli. EMBO J. 21, 3841-3851 (2002).
80. Losey, H.C., Ruthenburg, A.J. & Verdine, G.L. Crystal structure of Staphylococcus aureus tRNA adenosine deaminase TadA in complex with RNA. Nat. Struct. Mol. BioL 13, 153-159 (2006). 81. Elias, Y. & Huang, R.H. Biochemical and structural studies of A-to-I editing by tRNA:A34 deaminases at the wobble position of transfer RNA. Biochemistry 44, 12057-12065 (2005).
82. Song, M. et al. Sequence-specific prediction of the efficiencies of adenine and cytosine base editors. Nat. Biotechnol. 38, 1037-1043 (2020).
83. Doman, J.L., Raguram, A., Newby, G.A. & Liu, D.R. Evaluation and minimization of Cas9-independent off-target DNA editing by cytosine base editors. Nat. Biotechnol. 38, 620- 628 (2020).
84. Yu, Y. et al. Cytosine base editors with minimized unguided DNA and RNA off-target events and high on-target activity. Nat. Commun. 11, 2052 (2020).
85. Wang, L. et al. Eliminating base-editor-induced genome-wide and transcriptome-wide off-target mutations. Nat. Cell Biol. 23, 552-563 (2021).
86. Jin, S. et al. Rationally Designed APOBEC3B Cytosine Base Editors with Improved Specificity. Mol. Cell 79, 728-740 e726 (2020).
87. Richter, M.F. et al. Phage-assisted evolution of an adenine base editor with improved Cas domain compatibility and activity. Nat. Biotechnol. 38, 883-891 (2020).

Claims

1. A polypeptide comprising a tRNA-specific adenosine deaminase protein operably linked to an RNA-targeting protein.
2. The polypeptide of claim 1, wherein the tRNA-specific adenosine deaminase protein is a TadA7.10, TadA8.20, or TadA8e adenosine deaminase.
3. The polypeptide of claim 1 or 2, wherein the tRNA-specific adenosine deaminase protein comprises a F148A and/or V106G substitution.
4. The polypeptide of any one of claims 1 to 3, wherein the RNA-targeting protein is a Casl3d protein.
5. The polypeptide of claim 4, wherein the Casl3d protein is a dead Casl3d (dCasl3d) protein.
6. The polypeptide of claim 5, wherein the dCasl3d protein is a RfxCasl3d from Ruminococcus flavefaciens XPD3002.
7. The polypeptide of any one of claims 1 to 6, wherein the tRNA-specific adenosine deaminase protein is operably linked to the RNA-targeting protein by a flexible linker.
8. The polypeptide of claim 7, wherein the flexible linker is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 in length.
9. The polypeptide of any one of claims 1 to 8, wherein the tRNA-specific adenosine deaminase protein is operably linked to the N-terminus of the RNA-targeting protein.
10. The polypeptide of any one of claims 1 to 8, wherein the tRNA-specific adenosine deaminase protein is operably linked to the C-terminus of the RNA-targeting protein.
11. The polypeptide of any one of claims 1 to 10, further comprising at least one nuclear localization signal sequence.
12. A nucleic acid encoding the polypeptide of any one of claims 1 to 11.
13. The nucleic acid of claim 12, further encoding a short guide RNA (sgRNA).
14. The nucleic acid of claim 13, wherein the sgRNA binds to a target RNA.
15. The nucleic acid of claim 14, wherein the target RNA is a disease-causing RNA.
16. The nucleic acid of claim 14 or 15, wherein the target RNA comprises an adenosine associated with a disease.
17. The nucleic acid of claim 16, wherein the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of the adenosine associated with the disease.
18. An expression vector encoding the nucleic acid of any one of claims 12 to 16.
19. A cell comprising the polypeptide of any one of claims 1 to 11, the nucleic acid of any one of claims 12 to 17, or the expression vector of claim 18.
20. The cell of claim 19, comprising an RNA having an adenosine associated with a disease.
21. A system comprising: a polypeptide comprising a TRNA-specific adenosine deaminase protein operably linked to an RNA-targeting protein; and a short guide RNA.
22. The system of claim 21, wherein the polypeptide is the polypeptide of any one of claims 1 to 11.
23. The system of claim 21 or 22, wherein the short guide RNA binds to a target RNA.
24. The system of claim 23, wherein the target RNA is a disease-causing RNA.
25. The system of claim 23 or 24, wherein the target RNA comprises an adenosine associated with a disease.
26. A method for modifying one or more adenosine bases and/or for editing one or more adenosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the polypeptide of any one of claims 1 to 11.
27. The method of claim 26, wherein at least one of the adenosine bases are associated with a disease.
28. The method of claim 26 or 27, further comprising contacting the nucleic acid molecule with a short guide RNA (sgRNA).
29. The method of claim 28, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one adenosine associated with the disease.
30. The method of any one of claims 26 to 29, wherein the contacting results in a conversion of at least one of the adenosine bases in the nucleic acid molecule to an inosine base.
31. The method of claim 30, wherein the conversion is at least 20% efficient.
32. The method of any one of claims 26 to 31, wherein at least one additional adenosine base is not converted to an inosine.
33. A method for modifying adenosine bases and/or for editing adenosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the system of any one of claims 21 to 25.
34. The method of claim 33, wherein at least one of the adenosine bases are associated with a disease.
35. The method of claim 34, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one adenosine associated with the disease.
36. The method of any one of claims 33 to 35, wherein the contacting results in a conversion of at least one of the adenosine bases in the nucleic acid molecule to an inosine base.
37. The method of claim 36, wherein the conversion is at least 20% efficient.
38. The method of any one of claims 33 to 37, wherein at least one additional adenosine base is not converted to an inosine.
39. A method of treating a disease in a patient, the method comprising administering to the patient an effective amount of the polypeptide of any one of claims 1 to 11, the nucleic acid of any one of claims 12 to 17, the expression vector of claim 18, the cell of claim 19 or 20, or the system of any one of claims 21 to 25.
40. The method of claim 39, wherein the patient has at least one cell comprising a disease- associated adenosine in an RNA.
41. The method of claim 40, wherein the polypeptide converts the adenosine to an inosine.
42. The method of claim 40, wherein the polypeptide converts at least 20% of the disease- associated adenosines in the cell to an inosine.
43. The method of any one of claims 40 to 42, wherein the disease-associated adenosine is an /A7A-A-48T mutation.
44. The method of any one of claims 40 to 43, wherein the patient has, is suspected of having, or has been diagnosed with having a disease associated with a mutation to an adenosine in an RNA.
45. The method of any one of claims 40 to 44, wherein the patient has, is suspected of having, or has been diagnosed with having Van der Woude syndrome.
46. A polypeptide comprising an tRNA-specific cytosine deaminase protein operably linked to an RNA-targeting protein.
47. The polypeptide of claim 46, wherein the tRNA-specific cytosine deaminase protein is a TadA.CBE3.1, TadA.CBE2.4, TadA.CBE.DLL, TadA.CBE.DL, TadA.CBE1.46, or TadA.CBE1.52 cytosine deaminase.
48. The polypeptide of claim 46 or 47, wherein the tRNA-specific cytosine deaminase protein has at least one mutation from a TadA protein.
49. The polypeptide of any one of claims 46 to 48, wherein the RNA-targeting protein is a Cast 3d protein.
50. The polypeptide of claim 49, wherein the Casl3d protein is a dead Casl3d (dCasl3d) protein.
51. The polypeptide of claim 50, wherein the dCasl3d protein is a RfxCasl3d from Ruminococcus flavefaciens XPD3002.
52. The polypeptide of any one of claims 46 to 51, wherein the TadA cytosine deaminase protein is operably linked to the RNA-targeting protein by a flexible linker.
53. The polypeptide of claim 52, wherein the flexible linker is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35 in length.
54. The polypeptide of any one of claims 46 to 53, wherein the TadA cytosine deaminase protein is operably linked to the N-terminus of the RNA-targeting protein.
55. The polypeptide of any one of claims 46 to 53, wherein the TadA cytosine deaminase protein is operably linked to the C-terminus of the RNA-targeting protein.
56. The polypeptide of any one of claims 46 to 55, further comprising at least one nuclear localization signal sequence.
57. A nucleic acid encoding the polypeptide of any one of claims 46 to 56.
58. The nucleic acid of claim 57, further encoding a short guide RNA (sgRNA).
59. The nucleic acid of claim 58, wherein the sgRNA binds to a target RNA.
60. The nucleic acid of claim 59, wherein the target RNA is a disease-causing RNA.
61. The nucleic acid of claim 59 or 60, wherein the target RNA comprises a cytosine associated with a disease.
62. The nucleic acid of claim 61, wherein the sgRNA binds to the target RNA 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of the cytosine associated with the disease.
63. An expression vector encoding the nucleic acid of any one of claims 57 to 61.
64. A cell comprising the polypeptide of any one of claims 46 to 56, the nucleic acid of any one of claims 12 to 17, or the expression vector of claim 63.
65. The cell of claim 64, comprising an RNA having a cytosine associated with a disease.
66. A system comprising: a polypeptide comprising a tRNA-specific cytosine deaminase protein operably linked to an RNA-targeting protein; and a short guide RNA.
67. The system of claim 66, wherein the polypeptide is the polypeptide of any one of claims 1 to 11.
68. The system of claim 66 or 67, wherein the short guide RNA binds to a target RNA.
69. The system of claim 68, wherein the target RNA is a disease-causing RNA.
70. The system of claim 68 or 69, wherein the target RNA comprises a cytosine associated with a disease.
71. A method for modifying one or more cytosine bases and/or for editing one or more cytosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the polypeptide of any one of claims 46 to 56.
72. The method of claim 71, wherein at least one of the cytosine bases are associated with a disease.
73. The method of claim 71 or 72, further comprising contacting the nucleic acid molecule with a short guide RNA (sgRNA).
74. The method of claim 73, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one cytosine associated with the disease.
75. The method of any one of claims 71 to 74, wherein the contacting results in a conversion of at least one of the cytosine bases in the nucleic acid molecule to a uridine base.
76. The method of claim 75, wherein the conversion is at least 15% efficient.
77. The method of any one of claims 71 to 76, wherein at least one additional cytosine base is not converted to a uridine.
78. A method for modifying cytosine bases and/or for editing cytosine bases in a nucleic acid molecule comprising contacting the nucleic acid molecule with the system of any one of claims 66 to 70.
79. The method of claim 78, wherein at least one of the cytosine bases are associated with a disease.
80. The method of claim 79, wherein the sgRNA binds to the nucleic acid molecule 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 nucleotides upstream of at least one cytosine associated with the disease.
81. The method of any one of claims 78 to 80, wherein the contacting results in a conversion of at least one of the cytosine bases in the nucleic acid molecule to a uridine base.
82. The method of claim 81, wherein the conversion is at least 15% efficient.
83. The method of any one of claims 78 to 82, wherein at least one additional cytosine base is not converted to a uridine.
84. A method of treating a disease in a patient, the method comprising administering to the patient an effective amount of the polypeptide of any one of claims 1 to 11, the nucleic acid of any one of claims 12 to 17, the expression vector of claim 63, the cell of claim 19 or 20, or the system of any one of claims 21 to 25.
85. The method of claim 84, wherein the patient has at least one cell comprising a disease- associated cytosine in an RNA.
86. The method of claim 85, wherein the polypeptide converts the cytosine to a uridine.
87. The method of claim 85, wherein the polypeptide converts at least 15% of the disease- associated cytosines in the cell to a uridine.
88. The method of any one of claims 85 to 87, wherein the patient has, is suspected of having, or has been diagnosed with having a disease associated with a mutation to an cytosine in an RNA.
PCT/US2025/033195 2024-06-12 2025-06-11 Polypeptides and methods for modifying nucleic acids Pending WO2025259780A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202463659189P 2024-06-12 2024-06-12
US63/659,189 2024-06-12

Publications (1)

Publication Number Publication Date
WO2025259780A1 true WO2025259780A1 (en) 2025-12-18

Family

ID=98051589

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2025/033195 Pending WO2025259780A1 (en) 2024-06-12 2025-06-11 Polypeptides and methods for modifying nucleic acids

Country Status (1)

Country Link
WO (1) WO2025259780A1 (en)

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200308571A1 (en) * 2019-02-04 2020-10-01 The General Hospital Corporation Adenine dna base editor variants with reduced off-target rna editing
WO2023034959A2 (en) * 2021-09-03 2023-03-09 The University Of Chicago Polypeptides and methods for modifying nucleic acids

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200308571A1 (en) * 2019-02-04 2020-10-01 The General Hospital Corporation Adenine dna base editor variants with reduced off-target rna editing
WO2023034959A2 (en) * 2021-09-03 2023-03-09 The University Of Chicago Polypeptides and methods for modifying nucleic acids

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
ZHOU CHANGYANG; SUN YIDI; YAN RUI; LIU YAJING; ZUO ERWEI; GU CHAN; HAN LINXIAO; WEI YU; HU XINDE; ZENG RONG; LI YIXUE; ZHOU HAIBO;: "Off-target RNA mutation induced by DNA base editing and its elimination by mutagenesis", NATURE, vol. 571, no. 7764, 10 June 2019 (2019-06-10), pages 275 - 278, XP036831896, DOI: 10.1038/s41586-019-1314-0 *

Similar Documents

Publication Publication Date Title
KR102851101B1 (en) Methods and compositions for editing RNA
JP7605852B2 (en) Class II V-type CRISPR system
AU2019357450B2 (en) Methods and compositions for editing RNAs
KR102852347B1 (en) Method for substituting pathogenic amino acids using a programmable base editor system
JP2023075118A (en) RNA TARGETING OF MUTATIONS VIA SUPPRESSOR tRNAs AND DEAMINASES
US20240352439A1 (en) Polypeptides and methods for modifying nucleic acids
KR20180069898A (en) Nucleobase editing agents and uses thereof
KR20250107288A (en) Uses of adenosine base editors
JP2020534795A (en) Methods and Compositions for Evolving Base Editing Factors Using Phage-Supported Continuous Evolution (PACE)
CN115698278A (en) Compositions comprising Cas12i2 variant polypeptides and uses thereof
JP2025532113A (en) Novel adenine deaminase mutant and base editing method using the same
CA2683960A1 (en) Methods of genetically encoding unnatural amino acids in eukaryotic cells using orthogonal trna/synthetase pairs
CA3163741A1 (en) Compositions comprising a nuclease and uses thereof
CA3173526A1 (en) Rna-guided genome recombineering at kilobase scale
US20230383288A1 (en) Systems, methods, and compositions for rna-guided rna-targeting crispr effectors
JP2018011525A (en) Genome editing method
US20230295589A1 (en) Compositions comprising a variant cas12i4 polypeptide and uses thereof
WO2025007044A1 (en) Methods and compositions comprising iscb variants
WO2024044329A1 (en) Crispr base editor
CN117210435A (en) Editing system for regulating and controlling RNA methylation modification and application thereof
US20250369958A1 (en) Methods for prediction and treatment of limb-girdle muscular dystrophy
US20230323335A1 (en) Miniaturized cytidine deaminase-containing complex for modifying double-stranded dna
HK40081918A (en) Methods and compositions for editing rna
HK40081918B (en) Methods and compositions for editing rna
HK40056042A (en) Methods and compositions for editing rnas

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25822007

Country of ref document: EP

Kind code of ref document: A1