WO2024259294A2 - Proteomimetic scaffolds and uses thereof - Google Patents
Proteomimetic scaffolds and uses thereof Download PDFInfo
- Publication number
- WO2024259294A2 WO2024259294A2 PCT/US2024/034092 US2024034092W WO2024259294A2 WO 2024259294 A2 WO2024259294 A2 WO 2024259294A2 US 2024034092 W US2024034092 W US 2024034092W WO 2024259294 A2 WO2024259294 A2 WO 2024259294A2
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- side chain
- scaffold
- compound
- amino acid
- rna
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C07—ORGANIC CHEMISTRY
- C07K—PEPTIDES
- C07K7/00—Peptides having 5 to 20 amino acids in a fully defined sequence; Derivatives thereof
- C07K7/04—Linear peptides containing only normal peptide links
- C07K7/06—Linear peptides containing only normal peptide links having 5 to 11 amino acids
-
- C—CHEMISTRY; METALLURGY
- C07—ORGANIC CHEMISTRY
- C07K—PEPTIDES
- C07K2319/00—Fusion polypeptide
Definitions
- the present invention relates to proteomimetic scaffolds. More specifically, the present invention is directed to proteomimetic scaffolds and their uses for sequence-specific recognition and binding of duplex RNA.
- RNA Base-specific recognition of RNA is possible by antisense strategies (Crooke et al., “Antisense Technology: An Overview and Prospectus,” Nature Reviews Drug Discovery 20 ATI - 453 (2021); Rinaldi et al., “Antisense Oligonucleotides: the Next Frontier for Treatment of Neurological Disorders,” Nature Reviews Neurology 14:9-21 (2018)), including synthetic oligomers such as morpholinos (Corey et al., “Morpholino Antisense Oligonucleotides: Tools for Investigating Vertebrate Development,” Genome Biology llSy.xQvxosNS 1015.1-1015.3 (2001)) and peptide nucleic acids (PNAs) (Nielsen, P.
- a-helix or P-hairpin motif in principle, provides a similar opportunity for dsRNA recognition, but these protein secondary structure scaffolds are limited in the number of side chains they can project into the RNA grooves, which in turn lowers their potential as a general platform for sequence-specific nucleic acid binding.
- the present disclosure provides an encodable scaffold that engages the major groove of double-stranded RNA.
- the present disclosure provides a compound having a structure according to Formula (I), Formula (II), Formula (III), Formula (IV), Formula (V), Formula (VI), or Formula (VII):
- Q is independently at each occurrence selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain;
- the present disclosure provides a method of binding doublestranded RNA. This method comprises contacting the double-stranded RNA with the compound according to the present disclosure, the composition according to the present disclosure, or the dosage form according to the present disclosure. [00014] In another aspect, the present disclosure provides a method of treating a repeat expansion disorder in a subject in need of such treatment. This method comprises administering to the subject a therapeutically effective amount of the compound according to the present disclosure, the composition according to the present disclosure, or the dosage form according to the present disclosure.
- the present disclosure provides a method for treating a Huntington’s disease in a subject in need of such treatment.
- This method comprises administering to the subject a therapeutically effective amount of the compound according to the present disclosure, the composition according to the present disclosure, or the dosage form according to the present disclosure.
- Figures 1A-1C depict nucleic acid recognition by small molecule and proteins.
- Fig. 1A shows DNA recognition by hairpin polyamide, pdb 3omj; zinc-finger protein, pdb Imey; and basic-leucine zipper, pdb Ifos.
- Fig. IB shows RNA recognition by Rev a-helix peptide, pdb letf, and cyclic Tat b-hairpin mimetic, pdb 2kdq.
- Fig. 1C shows Hoogsteen hydrogen bonding between arginine and guanine, and asparagine or glutamine with adenine.
- FIGS 2A-2B depict the design of the crosslinked helical fork (CHF) scaffold.
- Fig. 2A shows that from wild-type TAV2b protein, al helix was retained and truncated to an 8- residue sequence containing R26-R33 (ribbon). The individual al-derived peptides were covalently crosslinked to obtain the CHF scaffold.
- Fig. 2B shows side-chain interactions of R26- R33 with duplex RNA from the native complex, pdb 2zi0. Black dashes represent side-chain interaction with the RNA phosphate backbone. Wedges represent Hoogsteen hydrogen bonds between amino acid side chain to nucleobases, dotted box.
- Figure 3A shows binding affinities of TAV2b-ul Flu and CHF-l Flu for designed oligonucleotide hairpins according to certain embodiments of the present invention, RNA 1 and DNA 1.
- TAV2b-al Flu and CHF-l Flu sequences are listed in brackets. Highlights on helices and amino acid sequence represent residues involved in Hoogsteen recognition.
- Nle Norleucine.
- Ri Fluorescein-b-Alanine.
- R2 Fluorescein-b-Alanyl-glycine.
- R3 Acetyl -tyrosine. All peptide C- termini are carboxamides.
- Figure discloses SEQ ID NOS 36 and 40, respectively, in order of appearance.
- Figure 3B depicts designed hairpin oligonucleotides. Oligonucleotide sequences are listed in Table 3. Figure discloses SEQ ID NOS 26 and 27, respectively, in order of appearance. [00020] Figures 4A-4B outline a comparison of monomeric and dimeric peptide recognition of hairpin oligonucleotides. Figure 4A shows the sequence and secondary structure of oligonucleotides. Schematics were generated with NUPACK. Square box indicates guanine nucleobases engaged in Hoogsteen recognition. Figure discloses SEQ ID NOS 26-29, respectively, in order of appearance.
- Figure 4B shows binding affinities of fluorescein-labeled peptides for different hairpin oligonucleotides as measured in a fluorescence polarization assay.
- Figure discloses SEQ ID NOS 4, 37, and 5, respectively, in order of appearance.
- Ac Acetyl.
- Flu Fluorescein.
- X 4-pentenoic acid.
- G* A’-allylglycine.
- O Melamine. All peptide C-termini are carboxamides.
- Figures 5A-5B show analysis of specific and non-specific contacts in CHF’RNA recognition.
- Figure 5A shows that CHF-1 contains two solvent exposed lysine residues (K27 and K30) along with residues that potentially engage the RNA Hoogsteen surface; substitution of these cationic residues with alanine provided CHF-2.
- CHF-3 represents a negative control in which the nucleotide-engaging residues (R26 and H29) have been substituted with alanine.
- Models were generated with UCSF ChimeraX using pdb 2zi0. Binding affinities of fluorescein-labeled CHF compounds as determined in a fluorescence polarization assay. Bar graph of affinities represented in log(Ka).
- Figure discloses SEQ ID NOS 3, 38, and 39, respectively, in order if appearance.
- Figure 6 depicts hydroxyl radical footprinting of 5 ’-32 P-radiolabeled RNA 1 with CHF-2.
- Footprint signatures are labeled in brackets (i) to (iv). Footprints corresponding to RNA 1 sequence are denoted with Roman numerals and dashed boxes.
- Overlay of CHF-2 corresponds to binding site for RNA 1.
- Figures 7A-7C show an assessment of conformational changes in peptide (CHF- 2) and r(CUG)io upon complex formation.
- Figure 7A shows circular dichroism spectra of 30 pM CHF-2 (black) and signal of 7 pM r(CUG)io subtracted from CD spectrum of 10 pM CHF-2 and 7 nM r(CUG)io complex (purple).
- Figure 7B shows 7 pM r(CUG)io, 10 pM CHF-2 and 7 pM r(CUG)io complex and 4 pM of d(CTG)io. All measurements were conducted in 10 mM Tris and 10 mM KF, pH 7.4.
- Figure 7C shows a comparison of major grooves between TAV2b:siRNA complex (pdb 2zi0) and a canonical Bform DNA (pdb 3bse). Two overlaid views of the bound siRNA and DNA are shown. Dashed lines indicate major grooves.
- Figure 8 shows secondary structures of ssRNA 1, RNA 1, DNA 1, r(CUG)io, d(CTG)io, and r(CUG)Mut as predicted and generated from NUPACK (Zadeh et al., “NUPACK: Analysis and Design of Nucleic Acid Systems,” J. Comput. Chem. 32(1): 170-173 (2011), which is hereby incorporated by reference in its entirety) and RNAfold (Lorenz et al., “ViennaRNA Package 2.0,” Algorithms Mol. Biol. 6:26 (2011), which is hereby incorporated by reference in its entirety). Black boxes indicate Hoogsteen recognition sites. Grey boxes represent guanine to adenine mutation. Figure discloses SEQ ID NOS 26-30, 32 and 33, respectively, in order of appearance.
- Figure 9A shows analytical HPLC traces of peptides according to certain embodiments of the invention.
- Figure 9B shows analytical HPLC traces of Synthetic Analogues of CHF-2 according to certain embodiments of the invention.
- Figure 10 shows fluorescence polarization assay binding measurements between peptide compounds according to embodiments of the present disclosure and oligonucleotides.
- Figure 11 shows microscale thermophoresis binding measurement between CHF- 2 and 5 nM Cy5-r(CUG)io.
- Figure 12 depicts an electrophoretic mobility gel shift assay for the direct binding of 32 P-ssRNA 1 and CHF-2.
- Figure 13A describes direct binding of CHF-2 Flu . Direct binding was performed using Fluorescence Polarization Assay between CHF-2 Flu and listed oligonucleotides.
- Figure discloses SEQ ID NOS 28, 32 and 33, respectively, in order of appearance.
- Figure 13B describes binding of “Stapled” Analogue of CHF-2.
- Figure 13C describes binding of “D- isomer/RetroInverso” Analogue of CHF-2.
- Figure 13D describes binding of the “OHM” Analogue of CHF-2.
- Figures 14A-14C show analysis of dsDNA and dsRNA recognition by Max homodimer, a basic helix loop motif.
- Figure 14B shows comparison of dsDNA and dsRNA major groove interactions with a-helices: (i) positioning of base pairs within major grooves; (ii-iv) structural analysis of DNA in complex with Max homodimer (PDB code Ihlo) and RNA in complex with Rev peptide (PDB code letf) and TAV2b (PDB code 2zi0) reveals that the angle of entry of dimeric oc -helical domains into dsDNA and dsRNA diverge.
- Figure 14C shows that replacement of the leucine zipper region in Max homodimer with a flexible linker leads to CHF-Max ( Figures 15 and 16). CHF-Max binding to Ebox DNA and RNA was evaluated in a fluorescence polarization assay.
- Ka(CHF-Max/Ebox DNA) 0.69 ⁇ 0.21 pM
- Ka(CHF-Max/Ebox RNA) 0.48 ⁇ 0.22 pM.
- Figures 15A-15B show the distance between truncated helical dimers in DNA and RNA major grooves.
- Figure 15A shows crosslinking distance of 16 A for CHF-Max.
- Figure 15A was generated from PDB code Ihlo.
- Figure 15B shows crosslinking distance of 8 A for TAV2b- al Flu and 19 A for the minimal CHF constructs.
- Figure 15B was generated using PDB code 2zi0.
- Figure 16 shows the chemical structure of CHF-Max.
- Figure discloses SEQ ID NOS 23 and 24, respectively, in order of appearance.
- the terms “treat” or “treatment” of a state, disorder or condition include: (1) preventing, delaying, or reducing the incidence and/or likelihood of the appearance of at least one clinical or sub-clinical symptom of the state, disorder or condition developing in a subject that may be afflicted with or predisposed to the state, disorder or condition but does not yet experience or display clinical or subclinical symptoms of the state, disorder or condition; or (2) inhibiting the state, disorder or condition, i.e., arresting, reducing or delaying the development of the disease or a relapse thereof or at least one clinical or sub-clinical symptom thereof; or (3) relieving the disease, i.e., causing regression of the state, disorder or condition or at least one of its clinical or sub-clinical symptoms.
- the benefit to a subject to be treated is either statistically significant or at least perceptible to the patient or to the physician.
- the subject is a human.
- the term “effective” applied to dose or amount refers to that quantity of a compound or pharmaceutical composition that is sufficient to result in a desired activity upon administration to a subject in need thereof.
- the effective amount of the combination may or may not include amounts of each ingredient that would have been effective if administered individually. The exact amount required will vary from subject to subject, depending on the species, age, and general condition of the subject, the severity of the condition being treated, the particular drug or drugs employed, the mode of administration, and the like.
- compositions of the invention refers to molecular entities and other ingredients of such compositions that are physiologically tolerable and do not typically produce untoward reactions when administered to a mammal (e.g., a human).
- pharmaceutically acceptable means approved by a regulatory agency of the Federal or a state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeia for use in mammals, and more particularly in humans.
- Ranges can be expressed herein as from “about” or “approximately” one particular value and/or to “about” or “approximately” another particular value. When such a range is expressed, another embodiment includes from the one particular value and/or to the other particular value.
- aliphatic or “aliphatic group”, as used herein, means a straight-chain (i.e., unbranched) or branched, substituted or unsubstituted hydrocarbon chain that is completely saturated or that contains one or more units of unsaturation, or a monocyclic hydrocarbon, bicyclic hydrocarbon, or tricyclic hydrocarbon that is completely saturated or that contains one or more units of unsaturation, but which is not aromatic (also referred to herein as “carbocycle,” “cycloaliphatic” or “cycloalkyl”), that has a single point of attachment to the rest of the molecule. Unless otherwise specified, aliphatic groups contain 1—30 aliphatic carbon atoms.
- aliphatic groups contain 1-20 aliphatic carbon atoms. In other embodiments, aliphatic groups contain 1-10 aliphatic carbon atoms. In still other embodiments, aliphatic groups contain 1-6 aliphatic carbon atoms, and in yet other embodiments, aliphatic groups contain 1, 2, 3, or 4 aliphatic carbon atoms.
- Suitable aliphatic groups include, but are not limited to, linear or branched, substituted or unsubstituted alkyl, alkenyl, alkynyl groups and hybrids thereof such as (cycloalkyl)alkyl, (cycloalkenyl)alkyl or (cycloalkyl)alkenyl.
- cycloaliphatic refers to saturated or partially unsaturated cyclic aliphatic monocyclic, bicyclic, or polycyclic ring systems, as described herein, having from 3 to 14 members, wherein the aliphatic ring system is optionally substituted as defined above and described herein.
- Cycloaliphatic groups include, without limitation, cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, cycloheptenyl, cyclooctyl, cyclooctenyl, norbomyl, adamantyl, and cyclooctadienyl.
- the cycloalkyl has 3-6 carbons.
- cycloaliphatic may also include aliphatic rings that are fused to one or more aromatic or nonaromatic rings, such as decahydronaphthyl or tetrahydronaphthyl, where the radical or point of attachment is on the aliphatic ring.
- a carbocyclic group is bicyclic.
- a 'carbocyclic group is tricyclic.
- a carbocyclic group is polycyclic.
- cycloaliphatic refers to a monocyclic C3-C6 hydrocarbon, or a Cs-Cio bicyclic hydrocarbon that is completely saturated or that contains one or more units of unsaturation, but which is not aromatic, that has a single point of attachment to the rest of the molecule, or a C9-C16 tricyclic hydrocarbon that is completely saturated or that contains one or more units of unsaturation, but which is not aromatic, that has a single point of attachment to the rest of the molecule.
- alkyl is given its ordinary meaning in the art and may include saturated aliphatic groups, including straight-chain alkyl groups, branched-chain alkyl groups, cycloalkyl (alicyclic) groups, alkyl substituted cycloalkyl groups, and cycloalkyl substituted alkyl groups.
- a straight chain or branched chain alkyl has about 1-20 carbon atoms in its backbone (e.g., C1-C20 for straight chain, C2-C20 for branched chain), and alternatively, about 1-10 carbon atoms, or about 1 to 6 carbon atoms.
- a cycloalkyl ring has from about 3-20 carbon atoms in their ring structure where such rings are monocyclic or bicyclic, and alternatively about 5, 6 or 7 carbons in the ring structure.
- an alkyl group may be a lower alkyl group, wherein a lower alkyl group comprises 1-4 carbon atoms (e.g., C1-C4 for straight chain lower alkyls).
- alkenyl refers to an alkyl group, as defined herein, having one or more double bonds.
- alkynyl refers to an alkyl group, as defined herein, having one or more triple bonds.
- heteroalkyl is given its ordinary meaning in the art and refers to alkyl groups as described herein in which one or more carbon atoms is replaced with a heteroatom (e.g., oxygen, nitrogen, sulfur, and the like). Examples of heteroalkyl groups include, but are not limited to, alkoxy, poly(ethylene glycol)-, alkyl -substituted amino, tetrahydrofuranyl, piperidinyl, morpholinyl, etc.
- heterocycloalkyl is given its ordinary meaning in the art and refers to cycloalkyl groups as described herein in which one or more carbon atoms is replaced with a heteroatom (e.g., oxygen, nitrogen, sulfur, and the like).
- heterocycloalkyl groups include, but are not limited to, tetrahydrofuranyl, piperidinyl, morpholinyl, etc.
- aryl used alone or as part of a larger moiety as in “aralkyl,” “aralkoxy,” or “aryloxyalkyl,” refers to monocyclic or bicyclic ring systems having a total of five to fourteen ring members, wherein at least one ring in the system is aromatic and wherein each ring in the system contains 3 to 7 ring members.
- aryl may be used interchangeably with the term “aryl ring.”
- aryl refers to an aromatic ring system which includes, but not limited to, phenyl, biphenyl, naphthyl, binaphthyl, anthracyi and the like, which may bear one or more substituents.
- aryl is a group in which an aromatic ring is fused to one or more nonaromatic rings, such as indanyl, phthalimidyl, naphthimidyl, phenanthridinyl, or tetrahydronaphthyl, and the like.
- arylene means a group obtained by removal of a hydrogen atom from an aryl group.
- Non-limiting examples of arylene include phenylene and naphthylene.
- heteroaryl and “heteroar-,” used alone of as part of a larger moiety, e.g., “heteroaralkyl,” or “heteroaralkoxy,” refer to groups having 5 to 10 ring atoms (i.e., monocyclic or bicyclic), in some embodiments 5, 6, 9, or 10 ring atoms. In some embodiments, such rings have 6, 10, or 147t electrons shared in a cyclic array; and having, in addition to carbon atoms, from one to five heteroatoms.
- heteroatom refers to nitrogen, oxygen, or sulfur, and includes any oxidized form of nitrogen or sulfur, and any quatemized form of a basic nitrogen.
- Heteroaryl groups include, without limitation, thienyl, furanyl, pyrrolyl, imidazolyl, pyrazolyl, triazolyl, tetrazolyl, oxazolyl, isoxazolyl, oxadiazolyl, thiazolyl, isothiazolyl, thiadiazolyl, pyridyl, pyridazinyl, pyrimidinyl, pyrazinyl, indolizinyl, purinyl, naphthyridinyl, and pteridinyl.
- a heteroaryl is a heterobiaryl group, such as bipyridyl and the like.
- heteroaryl and “heteroar-”, as used herein, also include groups in which a heteroaromatic ring is fused to one or more aryl, cycloaliphatic, or heterocyclyl rings, where the radical or point of attachment is on the heteroaromatic ring.
- Nonlimiting examples include indolyl, isoindolyl, benzothienyl, benzofuranyl, dibenzofuranyl, indazolyl, benzimidazolyl, benzthiazolyl, quinolyl, isoquinolyl, cinnolinyl, phthalazinyl, quinazolinyl, quinoxalinyl, 4H — quinolizinyl, carbazolyl, acridinyl, phenazinyl, phenothiazinyl, phenoxazinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, and pyrido[2,3-b]-l,4-oxazin-3(4H)-one.
- a heteroaryl group may be monocyclic, bicyclic, tricyclic, tetracyclic, and/or otherwise polycyclic.
- heteroaryl may be used interchangeably with the terms “heteroaryl ring,” “heteroaryl group,” or “heteroaromatic,” any of which terms include rings that are optionally substituted.
- heteroarylkyl refers to an alkyl group substituted by a heteroaryl, wherein the alkyl and heteroaryl portions independently are optionally substituted.
- heterocycle As used herein, the terms “heterocycle,” “heterocyclyl,” “heterocyclic radical,” and “heterocyclic ring” are used interchangeably and refer to a stable 5- to 7-membered monocyclic or 7-10-membered bicyclic heterocyclic moiety that is either saturated or partially unsaturated, and having, in addition to carbon atoms, one or more, preferably one to four, heteroatoms, as defined above.
- nitrogen includes a substituted nitrogen.
- a heterocyclic ring can be attached to its pendant group at any heteroatom or carbon atom that results in a stable structure and any of the ring atoms can be optionally substituted.
- saturated or partially unsaturated heterocyclic radicals include, without limitation, tetrahydrofuranyl, tetrahydrothiophenyl pyrrolidinyl, piperidinyl, pyrrolinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, decahydroquinolinyl, oxazolidinyl, piperazinyl, dioxanyl, dioxolanyl, diazepinyl, oxazepinyl, thiazepinyl, morpholinyl, and quinuclidinyl.
- heterocycle used interchangeably herein, and also include groups in which a heterocyclyl ring is fused to one or more aryl, heteroaryl, or cycloaliphatic rings, such as indolinyl, 3H-indolyl, chromanyl, phenanthridinyl, or tetrahydroquinolinyl.
- a heterocyclyl group may be monocyclic, bicyclic, tricyclic, tetracyclic, and/or otherwise polycyclic.
- heterocyclylalkyl refers to an alkyl group substituted by a heterocyclyl, wherein the alkyl and heterocyclyl portions independently are optionally substituted.
- partially unsaturated refers to a ring moiety that includes at least one double or triple bond.
- partially unsaturated is intended to encompass rings having multiple sites of unsaturation, but is not intended to include aryl or heteroaryl moieties, as herein defined.
- heteroatom means one or more of oxygen, sulfur, nitrogen, phosphorus, or silicon (including, any oxidized form of nitrogen, sulfur, phosphorus, or silicon; the quaternized form of any basic nitrogen or; a substitutable nitrogen of a heterocyclic ring.
- halogen means F, Cl, Br, or I; the term “halide” refers to a halogen radical or substituent, namely -F, -Cl, -Br, or -I.
- compounds of the invention may contain “optionally substituted” moieties.
- substituted whether preceded by the term “optionally” or not, means that one or more hydrogens of the designated moiety are replaced with a suitable substituent.
- an “optionally substituted” group may have a suitable substituent at each substitutable position of the group, and when more than one position in any given structure may be substituted with more than one substituent selected from a specified group, the substituent may be either the same or different at every position.
- Combinations of substituents envisioned by this invention are preferably those that result in the formation of stable or chemically feasible compounds.
- stable refers to compounds that are not substantially altered when subjected to conditions to allow for their production, detection, and, in certain embodiments, their recovery, purification, and use for one or more of the purposes disclosed herein.
- amino acid further includes analogues, derivatives, and congeners of any specific amino acid referred to herein, as well as C-terminal or N-terminal protected amino acid derivatives (e.g., modified with an N-terminal or C-terminal protecting group).
- amino acid includes both D- and L- amino acids. Hence, an amino acid which is identified herein by its name, three letter or one letter symbol and is not identified specifically as having the D or L configuration, is understood to assume any one of the D or L configurations. In certain embodiments, peptides containing only L-amino acids are contemplated.
- Naturally occurring amino acids are identified throughout by the conventional three-letter and/or one-letter abbreviations, corresponding to the trivial name of the amino acid, in accordance with the following list: Alanine (Ala), Arginine (Arg), Asparagine (Asn), Aspartic acid (Asp), Cysteine (Cys), Glutamic acid (Glu), Glutamine (Gin), Glycine (Gly), Histidine (His), Isoleucine (He), Leucine (Leu), Lysine (Lys), Methionine (Met), Phenylalanine (Phe), Proline (Pro), Serine (Ser), Threonine (Thr), Tryptophan (Trp), Tyrosine (Tyr), and Valine (Vai).
- the abbreviations are accepted in the peptide art and are recommended by the IUPAC-IUB commission in biochemical nomenclature.
- structures depicted herein are also meant to include all isomeric (e.g., enantiomeric, diastereomeric, and geometric (or conformational)) forms of the structure; for example, the R and S configurations for each asymmetric center, (Z) and (E) double bond isomers, and (Z) and (E) conformational isomers. Therefore, single stereochemical isomers as well as enantiomeric, diastereomeric, and geometric (or conformational) mixtures of the present compounds are within the scope of the invention.
- structures depicted herein are also meant to include compounds that differ only in the presence of one or more isotopically enriched atoms.
- compounds having the present structures except for the replacement of hydrogen by deuterium or tritium, or the replacement of a carbon by a n C- or 13 C- or 14 C -enriched carbon are within the scope of this invention.
- crystalline forms of the compounds of the invention and salts thereof are also within the scope of the invention.
- the compounds of the invention may be isolated in various amorphous and crystalline forms, including without limitation forms which are anhydrous, hydrated, non-solvated, or solvated.
- Example hydrates include hemihydrates, monohydrates, dihydrates, and the like.
- the compounds of the invention are anhydrous and non-solvated.
- anhydrous is meant that the crystalline form of the compound contains essentially no bound water in the crystal lattice structure, i.e., the compound does not form a crystalline hydrate.
- crystalline form is meant to refer to a certain lattice configuration of a crystalline substance.
- Different crystalline forms of the same substance typically have different crystalline lattices (e.g., unit cells) which are attributed to different physical properties that are characteristic of each of the crystalline forms.
- different lattice configurations have different water or solvent content.
- the different crystalline lattices can be identified by solid state characterization methods such as by X-ray powder diffraction (PXRD). Other characterization methods such as differential scanning calorimetry (DSC), thermogravimetric analysis (TGA), dynamic vapor sorption (DVS), solid state NMR, and the like further help identify the crystalline form as well as help determine stability and solvent/water content.
- Crystalline forms of a substance include both solvated (e.g., hydrated) and nonsolvated (e.g., anhydrous) forms.
- a hydrated form is a crystalline form that includes water in the crystalline lattice.
- Hydrated forms can be stoichiometric hydrates, where the water is present in the lattice in a certain water/molecule ratio such as for hemihydrates, monohydrates, dihydrates, etc. Hydrated forms can also be non-stoichiometric, where the water content is variable and dependent on external conditions such as humidity.
- the compounds of the invention are substantially isolated.
- substantially isolated is meant that a particular compound is at least partially isolated from impurities.
- a compound of the invention comprises less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 2.5%, less than about 1%, or less than about 0.5% of impurities.
- Impurities generally include anything that is not the substantially isolated compound including, for example, other crystalline forms and other substances.
- RNA unlike DNA, folds into a multitude of secondary and tertiary structures. This structural diversity has impeded the development of ligands that can sequence-specifically target this biomolecule.
- the present invention is in the development of a scaffold for duplex RNA recognition is based on surprising and unexpected discovery of synthetic helical fork motifs that can be engineered to sequence-specifically target dsRNA.
- the present invention is based upon the inventor’s discovery that protein tertiary structure mimic may offer an opportunity to specifically target dsRNA.
- Home et al. “Proteomimetics as Protein-Inspired Scaffolds with Defined Tertiary Folding Patterns,” Nat. Chem. 12:331-337 (2020); Sadek et al., “Modulation of Virus-Induced NF-KB Signaling by NEMO Coiled Coil Mimics,” Nat. Comniun. 11 : 1786 (2020); Watkins et al., “Protein-Protein Interactions Mediated by Helical Tertiary Structure Motifs,” J. Am. Chem. Soc.
- the helix dimer represents the simplest helical tertiary structure and forms the basis of important DNA recognition motifs in the bZIP and bHLH families ( Figure 1A).
- Both transcription factor families feature a Y-shaped scissor-grip or helical fork model of DNA recognition (Vinson et al., “Scissors-Grip Model for DNA Recognition by a Family of Leucine Zipper Proteins,” Science 246:911-916 (1989); Oneil et al., “Design of DNA- Binding Peptides Based on the Leucine Zipper Motif,” Science 249:774-778 (1990), which are hereby incorporated by reference in their entirety).
- a synthetic helical fork (HF) motif may provide a generalized scaffold to target dsRNA. Unlike a single a-helical motif, an HF scaffold would have the advantage of clamping neighboring major grooves with two helices.
- RNA major groove is narrower and deeper than B-form DNA’s and often requires bulges to accommodate a-helices (Battiste et al., “Alpha Helix- RNA Major Groove Recognition in an HIV- 1 Rev Peptide RRE RNA Complex,” Science 273: 1547-1551 (1996); Wu et al., “Structural Basis for Recognition of the AGNN Tetraloop RNA Fold by the Double-Stranded RNA-Binding Domain of Rntlp RNase III,” Proc. Natl. Acad. Sci. USA 101 :8307-8312 (2004), which are hereby incorporated by reference in their entirety).
- a key impediment in developing HF mimics as dsRNA binders is that natural examples of dsRNA bound by HF motifs are scarce.
- bZIP and bHLH family proteins are not known to target RNA.
- a search of the literature and structural data repositories yielded one example of a Y-shaped helical motif from tomato aspermy virus protein 2b (TAV2b) bound to an siRNA duplex ( Figure 2A) (Chen et al., “Structural Basis for RNA-Silencing Suppression by Tomato Aspermy Virus Protein 2b,” Embo Reports 9:754-760 (2008), which is hereby incorporated by reference in its entirety).
- the basic leucine zipper (bZIP) and basic helix-loop-helix (bHLH) families of DNA binding proteins serve as inspirations for the proteomimetic design of the present disclosure.
- the helix dimer represents the simplest helical tertiary structure and forms the basis of important DNA recognition motifs in the bZIP and bHLH families ( Figures 14A-C).
- Both transcription factor families feature a Y-shaped scissor-grip or helical fork model of DNA recognition (Vinson et al., "Scissors-Grip Model for DNA Recognition by a Family of Leucine Zipper Proteins," Science 246:911-916 (1989); Oneil et al., “Design ofDNA-Binding Peptides Based on the Leucine Zipper Motif," Science 249:774-778 (1990), which are hereby incorporated by reference in their entirety). It was rationalized that a synthetic helical fork (HF) motif may provide a generalized scaffold to target dsRNA. Unlike a single a-helical motif, a helix fork scaffold would have the advantage of clamping neighboring major grooves with two helices.
- HF helical fork
- the DNA-binding region of Max emerge from a rigid helix-loop region ( Figure 14B-iv) to complex with DNA.
- the leucine zipper dimerization domain imposes an angle restriction that hinders complex formation with RNA.
- dimeric a-helices that bind RNA There are limited examples of dimeric a-helices that bind RNA; the search of the literature and structural data repositories yielded one example of a Y-shaped a-helical motif from tomato aspermy virus protein 2b (TAV2b) bound to an siRNA duplex (Chen et al., "Structural Basis for RNA-Silencing Suppression by Tomato Aspermy Virus Protein 2b," Embo Reports 9:754-760 (2008), which is hereby incorporated by reference in its entirety).
- TAV2b a-helices mimic the trajectory of Rev a-helix binding geometry in A-form RNA and adopt the wider entry angle required for RNA complexation.
- Max dimer could be engineered to bind RNA if the angle between the major groove binding Max a-helices was modified by replacement of the leucine zipper domain with artificial crosslinkers.
- the DNA binding basic domain of Max was crosslinked with a flexible GlyGlyCys-benzyl crosslinker to span the ⁇ 16 A distance between major grooves ( Figures 15A-15B); the N-terminus was modified with cysteine residues and crosslinked the two a-helices with 1,3-dibenzyl bromide to obtain crosslinked helical fork (CHF) from the Max sequence (CHF-Max).
- CHF- Max The chemical structure CHF- Max is depicted in Figure 16.
- the present disclosure provides an encodable scaffold that engages the major groove of double-stranded RNA.
- the present disclosure also provides the compounds that may mimic TAV2b, pharmaceutical compositions comprising these compounds, and methods for treating certain diseases in a subject in need of such treatment. Also provided in the present disclosure methods of binding double-stranded RNA.
- Minimal motif that mimics an RNA binding viral protein TAV2b consists of two segmented a-helical regions, defined as al and a2, that homodimerize in the major groove of the siRNA duplex ( Figure 2A).
- the binding motif observed in the two al helices closely resembles that of the bZIP transcription factor.
- a crosslinked helical fork (CHF) scaffold has been designed that can mimic al’s affinity for its cognate RNA sequence.
- the minimal peptide sequence RKRHKLNR (SEQ ID NO: 3) was grafted from al onto the scaffold. While the native peptide-RNA complex features several electrostatic contacts with the negatively charged RNA backbone, it contains two hydrogen bonding interactions on the Hoogsteen face of two guanine bases in the duplex (5’-GC-3’) sequence ( Figure 2B).
- the base specific hydrogen-bonding in the major groove is critical for specificity in protein-nucleic acids recognition.
- an encodable scaffold that engages the major groove of double-stranded RNA.
- the term “encodable” refers to a scaffold that can be encoded with the correct hydrogen-bonding functionality to contact RNA bases on the Hoogsteen or Watson - Crick faces.
- encodable scaffold refers to functional groups appended onto the scaffold that allows precise and predetermined targeting of given RNA duplex.
- the scaffold comprises a first domain DI and a second domain D2, wherein the first domain and the second domain are connected by a covalent linker L, and where the scaffold spatially fits in the major groove of double-stranded RNA.
- the scaffold binds to the major groove of the doublestranded RNA using base-specific bonding selected from the group consisting of Hoogsteen hydrogen bonding and electrostatic contact.
- the first domain DI and the second domain D2 are each helical oligomers or helix mimics.
- the first domain DI and the second domain D2 each comprise an oligomer comprising 2 to 20 units.
- DI can comprise an oligomer comprising 2 to 20 units, 2 to 19 units, 2 to 18 units, 2 to 17 units, 2 to 16 units, 2 to 15 units, 2 to 14 units, 2 to 13 units, 2 to 12 units, 2 to 11 units, or 2 to 10 units.
- DI can comprise an oligomer comprising 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 units.
- D2 can comprise an oligomer comprising 2 to 20 units, 2 to 19 units, 2 to 18 units, 2 to 17 units, 2 to 16 units, 2 to 15 units, 2 to 14 units, 2 to 13 units, 2 to 12 units, 2 to 11 units, or 2 to 10 units.
- D2 can comprise an oligomer comprising 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 units.
- DI and the D2 each comprise an oligomer comprising 2 to 20 units, wherein each oligomer is independently selected from a peptide, a nucleic acid, and an oligomer comprising repeating C1-C20 alkyl, C5-C20 aryl, C1-C20 heteroalkyl, C2-C20 heteroaryl, C3-C20 cycloalkyl, or C1-C20 heterocycloalkyl units, and combinations thereof.
- the DI and the D2 each comprise an oligomer comprising 2 to 10 oxopiperazine units.
- the DI and the D2 each comprise an oligomer comprising 2, 3, 4, 5, 6, 7, 8, 9, or 10 oxopiperazine units.
- the DI and the D2 can each comprise a peptide consisting of 7 to 10 amino acids.
- DI can comprise a peptide consisting of 7 to 10 amino acids.
- DI can comprise a peptide consisting of 7, 8, 9, or 10 amino acids.
- D2 can comprise a peptide consisting of 7 to 10 amino acids.
- D2 can comprise a peptide consisting of 7, 8, 9, or 10 amino acids.
- the DI and the D2 each comprise a peptide consisting of 8 or 9 amino acids.
- the DI and the D2 each comprise a peptide having a formula R-aa2-aa3-H-aa5-aa6-aa7-aa8, wherein R is arginine or a compound having the formula: wherein X is CH2 or NH, and n is 0, 1, or 2; and H is histidine.
- the DI and the D2 each comprise a peptide having a formula G-R-aa2-aa3-H-aa5-aa6-aa7-aa8, wherein G is glycine, H is histidine, and R is as defined above.
- the first domain DI and a second domain D2 can consist of both natural and non-natural amino acids.
- the DI and the D2 each comprise a peptide comprising at least one non-natural amino acid, at least two non-natural amino acids, at least three non-natural amino acids, or at least four non-natural amino acids.
- the DI and the D2 each comprise a peptide comprising at least one non-natural amino acid.
- the DI and the D2 each comprise an a-helical peptide or peptidomimetic.
- This special distance can be from about 2A to about 30 A, from about 5A to about 25 A, or from about 7 A to about 20 A.
- the spatial distance between the DI and the D2 is from about 5 A to about 25 A.
- the spatial distance between aal of DI and aal of D2 is about 5A to about 15 A, and the spatial distance between aa8 of DI and aa8 of D2 is about 15A to about 25 A.
- This angle can be between about 40° and about 80°, between about 50° and about 70°, between about 55° and about 65°, between about 60° and about 65°, or between about 63° and about 64°.
- the angle between the DI and the D2 can be about 55°, about 56°, about 57°, about 58°, about 59°, about 60°, about 61°, about 62°, about 63°, about 64°, about 65°, about 66°, about 67°, about 68°, about 69°, or about 70°.
- the angle between the first domain DI and the second domain D2 is between about 50° and about 70°. In another embodiment, the angle between DI and D2 is about 60° to about 65°. In yet another embodiment, the angle between DI and D2 is about 63.7°.
- the first domain DI and the second domain D2 are connected by a covalent linker L.
- This linker L can have a length of about 2A to about 30 A, from about 5 A to about 25 A, from about 5 A to about 20 A, about 5 A to about 15 A, or about 7 A to about 10 A. In at least one embodiment, the linker L is about 5 A to about 15 A in length.
- linker L can be a reversible linker or an irreversible linker.
- the linker L comprises a combination of a C1-C20 alkyl and a C5-C20 aryl.
- the linker L comprises a cis-stilbene, azobenzene, metaxylene, l,3-di( 1H-l,2,3-triazol-l-yl)benzene, l,3-di(l//-l,2,3-triazol-4-yl)benzene, or 1,4- diphenyl- 1H-l,2,3-triazole.
- the linker L has a structure selected from: wherein Q is independently at each occurrence selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain, and represents the point of attachment to DI and D2.
- the reversible linker comprises a disulfide (-S-S-) bond.
- the encodable scaffold has a structure according to Scheme
- the present disclosure provides a compound having a structure according to Formula (I), Formula (II), Formula (III), Formula (IV), Formula (V), Formula (VI), or Formula (VII):
- Q is independently at each occurrence selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain;
- the compound has the structure of Formula (I).
- the compound has the structure of Formula (II).
- the compound has the structure of Formula (III).
- the compound has the structure of Formula (IV). [000121] In some embodiments, the compound has the structure of Formula (V).
- the compound has the structure of Formula (VI).
- the compound has the structure of Formula (VII).
- Formula (II), Formula (III), Formula (IV), Formula (V), or Formula (VI) can be the same or different.
- One Q can be independently selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain and another Q can be independently selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain.
- Q groups are the same. In one embodiment, Q is -S- at each occurrence.
- Ri and R12 are each independently selected from alanine, [3- alanine, glycine, phenylalanine, tyrosine, lysine, aspartic acid, an acetyl group, and dimer combinations thereof, each of which is optionally substituted with fluorescein.
- R11 and R22 are each -NH2, -OH, -NHR*, -OR*, -SH, -SR*, or -NR*-N(R*)2, where R* is independently at each occurrence selected from hydrogen, a C5-C20 aryl, and a C1-C20 alkyl.
- R2 and R13 are each hydrogen.
- R3 and R14 are each a side chain of arginine.
- R4 and R15 are each independently selected from a side chain of lysine and a side chain of alanine.
- R5 and R16 are each a side chain of arginine.
- R6 and R17 are each independently selected from a side chain of histidine and a side chain of alanine.
- R7 and Ris are each independently selected from a side chain of lysine and a side chain of alanine.
- R8 and R19 are each a side chain of leucine.
- R9 and R20 are each a side chain of asparagine.
- Rio and R21 are each a side chain of arginine.
- each of R2 and R13 is hydrogen
- each of R3 and R14 is a side chain of arginine
- each of R4 and R15 is a side chain of lysine
- each of R5 and R16 is a side chain of arginine
- each of Re and R17 is a side chain of histidine
- each of R7 and Ris is a side chain of lysine
- each of R8 and R19 is a side chain of leucine
- each of R9 and R20 is a side chain of asparagine
- each of Rio and R21 is a side chain of arginine
- R11 and R22 are each NH2
- Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-P-alanine-glycine.
- each of R2 and R13 is hydrogen
- each of R3 and R14 is a side chain of arginine
- each of R4 and R15 is a side chain of alanine
- each of R5 and R16 is a side chain of arginine
- each of R6 and R17 is a side chain of histidine
- each of R7 and Ris is a side chain of alanine
- each of Rs and R19 is a side chain of leucine
- each of R9 and R20 is a side chain of asparagine
- each of Rio and R21 is a side chain of arginine
- R11 and R22 are each NH2
- Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-P-alanine-glycine.
- each of R2 and R13 is hydrogen
- each of R3 and R14 is a side chain of alanine
- each of R4 and R15 is a side chain of lysine
- each of R5 and R16 is a side chain of arginine
- each of R6 and R17 is a side chain of alanine
- each of R7 and Ris is a side chain of lysine
- each of R8 and R19 is a side chain of leucine
- each of R9 and R20 is a side chain of asparagine
- each of Rio and R21 is a side chain of arginine
- R11 and R22 are each NH2
- Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-P-alanine-glycine.
- each amino acid side chain can be an L- or D- isomer. In some embodiments, each amino acid side chain is the L-isomer.
- a lysine residue and an aspartic acid residue are cyclized as a lactam macrocycle.
- the peptide is selected from Table 1.
- the compound has a structure as shown in Table 2 or a pharmaceutically acceptable salt thereof.
- Table 2 Chemical Structure of Peptides Table discloses SEQ ID NOS 4-6, 10, 9, 12, 11, 14, 13, 15-16, 7-8, 20, 19, 17-18, and 21-22, respectively, in order of appearance.
- the compound has a structure
- the present disclosure provides pharmaceutical compositions comprising the compound of the present disclosure.
- compositions of the disclosure may be formulated with suitable carriers, excipients, and other agents that provide improved transfer, delivery, tolerance, and the like.
- suitable carriers excipients, and other agents that provide improved transfer, delivery, tolerance, and the like.
- a multitude of appropriate formulations may be found in the formulary known to all pharmaceutical chemists: Remington's Pharmaceutical Sciences, Mack Publishing Company, Easton, PA.
- formulations include, for example, powders, pastes, ointments, jellies, waxes, oils, lipids, lipid (cationic or anionic) containing vesicles (such as LIPOFECTINTM, Life Technologies, Carlsbad, CA), DNA conjugates, anhydrous absorption pastes, oil-in-water and water-in-oil emulsions, emulsions carbowax (polyethylene glycols of various molecular weights), semi-solid gels, and semi-solid mixtures containing carbowax.
- vesicles such as LIPOFECTINTM, Life Technologies, Carlsbad, CA
- DNA conjugates such as LIPOFECTINTM, Life Technologies, Carlsbad, CA
- DNA conjugates such as LIPOFECTINTM, Life Technologies, Carlsbad, CA
- DNA conjugates such as LIPOFECTINTM, Life Technologies, Carlsbad, CA
- DNA conjugates such as LIPOFECTINTM, Life Technologies, Carlsbad, CA
- the dose of a compound administered to a patient may vary depending upon the age and the size of the patient, target disease, conditions, route of administration, and the like.
- the suitable dose is typically calculated according to body weight or body surface area.
- intravenously administer the compound of the present disclosure normally at a single dose of about 0.01 to about 20 mg/kg body weight, more preferably about 0.02 to about 7, about 0.03 to about 5, or about 0.05 to about 3 mg/kg body weight.
- the frequency and the duration of the treatment may be adjusted.
- Effective dosages and schedules for administering a compound may be determined empirically; for example, patient progress may be monitored by periodic assessment, and the dose adjusted accordingly. Moreover, interspecies scaling of dosages may be performed using well-known methods in the art (e.g., Mordenti et al., 1991, Pharmacent. Res. 8:1351).
- Various delivery systems are known and may be used to administer the pharmaceutical composition of the disclosure, e.g., encapsulation in liposomes, microparticles, microcapsules, recombinant cells capable of expressing the mutant viruses, receptor mediated endocytosis (see, e.g., Wu et al., 1987, J. Biol. Chem. 262:4429-4432).
- Methods of introduction include, but are not limited to, intradermal, intramuscular, intraperitoneal, intravenous, subcutaneous, intranasal, epidural, and oral routes.
- composition may be administered by any convenient route, for example by infusion or bolus injection, by absorption through epithelial or mucocutaneous linings (e.g., oral mucosa, rectal and intestinal mucosa, etc.) and may be administered together with other biologically active agents. Administration may be systemic or local.
- epithelial or mucocutaneous linings e.g., oral mucosa, rectal and intestinal mucosa, etc.
- Administration may be systemic or local.
- a pharmaceutical composition of the present disclosure may be delivered subcutaneously or intravenously with a standard needle and syringe.
- a pen delivery device can be used to deliver a pharmaceutical composition of the present disclosure.
- Such a pen delivery device may be reusable or disposable.
- a reusable pen delivery device generally utilizes a replaceable cartridge that contains a pharmaceutical composition. Once all of the pharmaceutical composition within the cartridge has been administered and the cartridge is empty, the empty cartridge may readily be discarded and replaced with a new cartridge that contains the pharmaceutical composition. The pen delivery device may then be reused.
- a disposable pen delivery device there is no replaceable cartridge. Rather, the disposable pen delivery device comes prefilled with the pharmaceutical composition held in a reservoir within the device. Once the reservoir is emptied of the pharmaceutical composition, the entire device is discarded.
- the pharmaceutical composition may be delivered in a controlled release system.
- a pump may be used (see Langer, supra; Sefton, 1987, CRC Crit. Ref. Biomed. Eng. 14:201).
- polymeric materials may be used; see, Medical Applications of Controlled Release, Langer and Wise (eds.), 1974, CRC Pres., Boca Raton, Florida.
- a controlled release system may be placed in proximity of the composition’s target, thus requiring only a fraction of the systemic dose (see, e.g., Goodson, 1984, in Medical Applications of Controlled Release, supra, vol. 2, pp. 115-138).
- the injectable preparations may include dosage forms for intravenous, subcutaneous, intracutaneous and intramuscular injections, drip infusions, etc. These injectable preparations may be prepared by methods publicly known. For example, the injectable preparations may be prepared, e.g., by dissolving, suspending or emulsifying the compound or its salt described above in a sterile aqueous medium or an oily medium conventionally used for injections.
- aqueous medium for injections there are, for example, physiological saline, an isotonic solution containing glucose and other auxiliary agents, etc., which may be used in combination with an appropriate solubilizing agent such as an alcohol (e.g., ethanol), a polyalcohol (e.g., propylene glycol, polyethylene glycol), a nonionic surfactant [e.g., polysorbate 80, HCO-50 (polyoxyethylene (50 mol) adduct of hydrogenated castor oil)], etc.
- an alcohol e.g., ethanol
- a polyalcohol e.g., propylene glycol, polyethylene glycol
- a nonionic surfactant e.g., polysorbate 80, HCO-50 (polyoxyethylene (50 mol) adduct of hydrogenated castor oil
- oily medium there are employed, e.g., sesame oil, soybean oil, etc., which may be used in combination with a solubilizing agent such as benzyl benzoate, benzyl alcohol, etc.
- a solubilizing agent such as benzyl benzoate, benzyl alcohol, etc.
- the pharmaceutical compositions for oral or parenteral use described above are prepared into dosage forms in a unit dose suited to fit a dose of the active ingredients.
- dosage forms in a unit dose include, for example, tablets, pills, capsules, injections (ampoules), suppositories, etc.
- the compounds disclosed herein are useful, inter alia, for the treatment, prevention and/or amelioration of a disease, disorder or condition in need of such treatment.
- the present disclosure provides a method of treating a repeat expansion disorder in a subject in need of such treatment comprising administering to the subject a therapeutically effective amount of a compound according to the disclosure, or the composition comprising any compound according to the present disclosure.
- the compounds disclosed herein are useful for treating a noncoding repeat expansion disorder.
- the disorder is selected from the group consisting of Baratela-Scott syndrome; CANVAS, cerebellar ataxia, neuropathy and vestibular areflexia syndrome, myotonic dystrophy type 1 ; DM2, myotonic dystrophy type 2, progressive myoclonus epilepsy type 1 (Unverricht-Lundborg disease), familial adult myoclonic epilepsy, Fuchs endothelial corneal dystrophy type 3, fragile XE syndrome; FRDA, Friedreich ataxia, frontotemporal dementia/amyotrophic lateral sclerosis, Fragile X-associated premature ovarian infertility, fragile X syndrome, fragile X-associated tremor ataxia syndrome, global developmental delay, progressive ataxia, elevated glutamine, neuronal intranuclear inclusion disease, oculopharyngodistal myopathy type 1, oculopharyn
- the compounds disclosed herein are useful for treating a coding repeat expansion disorder.
- the disorder is selected from the group consisting of brachydactyly and cleidocranial dysplasia (BCCD), blepharophimosis, ptosis and epicanthus inversus (BPES), congenital central hypoventilation syndrome (CCHS), dentatorubral- pallidoluysian atrophy (DRPLA), early infantile epileptic encephalopathy type 1 (EIEE1), Huntington disease, Huntington disease-like 2, hand-foot-genital syndrome, holoprosencephaly type 5, mental retardation with isolated growth hormone deficiency (MRGH), oculopharyngeal muscular dystrophy, spinal and bulbar muscular atrophy, synpolydactyly type 1, and spinocerebellar ataxia.
- BCCD brachydactyly and cleidocranial dysplasia
- BPES congenital central hypoventilation syndrome
- CCHS congen
- the repeat is a repeat RNA selected from CUG, CAG, CAA, GAA, CGG, CCG, GCG, GCA, GCC, GCU, AUUCU, AUUGU, CAGCUG, CUGCAG, or CCCCGCCCCGCG (SEQ ID NO: 25).
- the repeat is CUG.
- the compounds disclosed herein are useful for treating a Huntington’s disease in a subject in need of such treatment.
- the subject is human.
- the subject is a mammal selected from the group consisting of a cat, a dog, a cow, a horse, a sheep, and a pig.
- the present disclosure provides a method of binding double-stranded RNA comprising contacting the double-stranded RNA with a compound disclosed herein.
- the double-stranded RNA is contacted with the compound in a cell.
- the cell is a bacterial cell.
- the cell is a mammalian cell.
- the cell is a human cell.
- the cell is cancer cell, a brain cell, a connective tissue cell, or a muscle cell.
- Fmoc-protected amino acids, chemicals, and solvents were purchased from Alfa Aesar, Chem-Impex International, Sigma Aldrich, Acros Organics, VWR, Tokyo Chemical Industry, Ambeed, and Matrix Scientific. Standard desalted oligonucleotides were purchased from Sigma Aldrich and used without further purification unless otherwise noted. Cy5-r(CUG)io RNA was ordered from Sigma Aldrich as HPLC purified. ATP, [y- 32 P] was purchased from PerkinElmer. T4 Polynucleotide Kinase (lOU/pL) and RNase T1 were purchased from ThermoFisher.
- Synthesized peptides were purified on Thermofisher Scientific Ultimate 3000 HPLC by reverse-phase high- performance liquid chromatography (RP-HPLC) using preparative Cis column (Phenomenex Luna® 5pM Cis, 100 A, 100 x 21.2 mm). Purity was assessed on Agilent 1260 Infinity series RP- HPLC with an analytical Cis column (Xterra® RP18 3.5 pM, 2.1 x 150 mm) using a gradient of 5-95% acetonitrile/water with 0.1% TFA. Traces were monitored at 220 nm.
- Resin was subsequently drained using vacuum filtration and washed with DMF (3x), DCM (3x), and DMF (3x).
- DMF 3x
- DCM 3x
- DMF 3x
- DIPEA 10 eq.
- Peptides was cleaved from resin by stirring in 4 mL of TFA/thioanisole/anisole/l,2-ethanedithiol (90:5:2:3 v/v) for 3 hours at room temperature. Cleavage solution was then filtered through a cotton plug syringe and concentrated via rotary evaporation. The concentrated peptide was precipitated with 10 mL of cold diethyl ether, pelleted by centrifugation, and decanted for a total of 3 times. The pellet was dried by a slow stream of Nife) and dissolved in water and minimal acetonitrile.
- a 4 mb solution containing bromoacetic acid (10 eq.), DIC (20 eq.), and HOBt (10 eq.) in DMF was prepared and stirred for 15 minutes.
- the solution was transferred to the resin and placed on a nutator for 2 hours.
- the resin was drained and washed with DMF (3x), MeOH (3x), and DCM (3x).
- Resin was resuspended in 4 mb of DMF, added with allylamine (20 eq ), and placed on a nutator for 20 minutes.
- the reaction was drained and washed with DMF (3x), MeOH (3x), and DCM (3x).
- the peptide was further elongated by two amino acids followed with an Fmoc-deprotection.
- a separate 20 mb solution of 4-pentenoic acid (5 eq.), DIC (5 eq.), and HOBt (5 eq.) in DMF that was previously stirred for 15 minutes was transferred to the resin and shaken on a nutator for 2 hours.
- the resin was washed with DMF (3x), MeOH (3x), DCM (3x), and Et2O (3x) and dried overnight in a vacuum desiccator.
- the resin was purged with nitrogen, added with 20 mol% Hovey da-Grubbs 2 nd Generation catalyst and 4 mb of anhydrous 1,2-di chloroethane.
- Ring-closing metathesis was performed by heating the mixture to 120 °C for 15 minutes at 150 W.
- the reaction was drained and washed with DMF (3x), MeOH (3x), and DCM (3x).
- 4-Methyltrityl group was then selectively removed by 3 mL additions of TFA/TIPS/DCM (1 :5:94) for 10 minutes at room temperature for a total of 8 rounds.
- the resin was then washed with 15% DIPEA in DMF (2x), DMF (3x), and DCM (3x).
- the peptide was fluorescently labeled by addition of a 4 mL solution of fluorescein isothiocyanate (3 eq.) and DIPEA (5 eq.) in DMF.
- reaction vessel was wrapped in aluminum foil and placed on a nutator overnight in the dark. Resin was drained and washed with DMF (3x), MeOH (3x), DCM (3x), and Et2O. HBS-1 was cleaved from resin and purified under the same conditions from “Synthesis of Peptide- 1 and CHF Monomeric Strands.”
- the flask was covered with aluminum foil and the suspension was left to stir at 85 °C for 24 hours.
- the reaction was then cooled to room temperature, transferred to a clean 50 mL round bottom flask, and diluted with 25 mL of MeOH. While the resin settled to the bottom, the supernatant that contained starting material was transferred to waste.
- the flask was added with another 25 mL of MeOH and the supernatant was removed. A total of 3 rounds was performed.
- CA-stilbene diol la (100 mg, 0.416 mmol) was added to a 10 mL round-bottom flask, charged with N2( g ) and equipped with a stir bar. 4.16 mL of anhydrous THF was added and the solution was left to stir on ice for 10 minutes.
- PBn (86 pL, 0.916 mmol) was added in dropwise addition and the reaction was stirred for 2 hours at room temperature.
- a second addition of PBr? (86 pL, 0.916 mmol) was added in the same manner and stirred for a following 2 hours. The reaction was cooled in an ice bath before quenching with 2 mL of saturated sodium bicarbonate.
- RT Total concentration of oligonucleotide
- Bindings were performed in a 96-well plate containing 5 nM Cy5-r(CUG)io and 2-fold serially diluted CHF-2, beginning at 200 pM, in 50 mM Tris, 50 mM KC1, 1 mM MgCh, 0.05% Tween- 20, pH 7.4. Sample plate was covered with aluminum foil and was incubated at room temperature for 1 hour before loading into premium coated capillary tubes.
- the instrument used the following parameters: excitation power at 10%, MST power at medium, phase 1 duration at 3 seconds, phase
- RT Total concentration of oligonucleotide
- LST Total concentration of fluorescein-labeled peptide
- FSB Fraction of bound fluorescein-labeled peptide
- Circular dichroism (CD) measurements were taken on a Jasco J-1500 Circular Dichroism spectrometer equipped with a temperature controller using 1 mm length cells and a scan speed of 5 nm/min at 298 K. The spectra were generated from an average of 10 scans with baseline subtractions. Samples were prepared in 10 mM Tris and 10 mM KF, pH 7.4 and were incubated for 30 minutes at room temperature before taking a reading.
- the reaction mixture was incubated for 1 hour at 37 °C.
- the reaction was then cooled to room temperature and quenched with 1 volume of (25:24:1) phenol/chloroform/isoamyl alcohol (Acros Organics).
- the solution was vortexed for 30 seconds and centrifuge at 14,000 rpm for 30 minutes.
- the aqueous layer was isolated and transferred to a clean microcentrifuge tube.
- the precipitant was centrifuged at 14,000 rpm, 4 °C for 30 minutes, and the supernatant was removed after.
- the pellet was washed with 3 volumes of 70% cold ethanol in water, centrifuge at 14,000 rpm, 4 °C for 30 minutes, and had the supernatant removed for a total of 3 times.
- the pellet was dried on a speedvac vacuum concentrator, resuspended in 10 pL of IX RNA loading dye, and purified on a denatured 20% polyacrylamide gel containing 8 M urea. Purified 32 P-labeled RNA was refolded by heating to 95 °C for 1 minute and slowly cooled to room temperature for 1 hour.
- Example 11 Electrophoretic Mobility Shift Assay
- Bindings were conducted in autoclave microcentrifuge tubes.
- 1 pL of freshly 32 P-labeled ssRNA 1 (13,500 cpm) was mixed with 2-fold serially diluted CHF-2 (starting at 100 pM) in 50 mM Tris, 50 mM KC1, 1 mM MgCh, pH 7.4.
- Samples were incubated at room temperature for 2 hours and was added with 10 pL of 50% glycerol.
- 10 pL of sample was loaded onto a 10% native polyacrylamide gel and ran at 120 V for 15 minutes. The gel was dried, autoradiographed, and imaged on Amersham Typhoon gel imager.
- Hydroxyl radical cleavages were carried out in autoclaved microcentrifuge tubes. Each tube contained 2 pL of freshly 32 P-labeled RNA 1 (12,400 cpm), 2-fold serially diluted CHF- 2 (starting at 1 mM final concentration), 2 pL of buffer (10 mM Tris, 10 mM KC1, 1 mM MgCh, pH 7.4), and diluted up to 6 pL of DEPC-treated water. The samples were incubated at room temperature for 1 hour.
- RNA 1 was added with 4.75 pL of hydrolysis buffer (50 mM NaHCCh, 1 mM EDTA, pH 9.2; ThermoFisher) in an autoclaved microcentrifuge tube. The RNA was heated to 90 °C for 10 minutes and cooled on ice. 1 volume of 2X RNA loading dye was added and 9 pL of sample was loaded onto the denatured polyacrylamide gel.
- hydrolysis buffer 50 mM NaHCCh, 1 mM EDTA, pH 9.2; ThermoFisher
- RNA 1 was added with 8.5 pL IX Sequencing Buffer (20 mM sodium citrate, 1 mM EDTA, 7M urea, pH 5.0; ThermoFisher) in an autoclaved microcentrifuge tube. The RNA was heated to 90 °C for 30 seconds and cool to room temperature for 1 hour. 1 pL of RNase T1 (lOOU/pL; ThermoFisher) was added and left to incubate at room temperature for 10 minutes. 1 volume of 2X RNA loading dye was added and 10 pL of sample was loaded onto a denatured 20% polyacrylamide gel containing 8 M urea.
- IX Sequencing Buffer 20 mM sodium citrate, 1 mM EDTA, 7M urea, pH 5.0; ThermoFisher
- Example 13 Results and Discussion of Examples 1-12
- TAV2b consists of two segmented a-helical regions, defined as al and a2, that homodimerize in the major groove of the siRNA duplex ( Figure 2A).
- the binding motif observed in the two al helices closely resembles that of the bZIP transcription factor. Therefore, a crosslinked helical fork (CHF) scaffold that can mimic al’s affinity for its cognate RNA sequence was designed.
- the minimal peptide sequence RKRHKLNR SEQ ID NO: 3 was grafted from al onto the scaffold.
- TAV2b/siRNA interaction has also inspired other efforts to develop constrained peptide mimics as RNA ligands.
- Kuepper et al. “Constrained Peptides Mimic a Viral Suppressor of RNA Silencing,” Nucleic Acids Res. 49: 12622-12633 (2021), which is hereby incorporated by reference in its entirety, has recently showed that a constrained sequence consisting of the al and a2 helical regions can recognize the cognate siRNA with high affinity.
- Cis-stilbene provides an optimally-spaced turn segment to link two parallel helices (Erdelyi et al., “A New Tool in Peptide Engineering: A Photoswitchable Stilbene-Type BetaHairpin Mimetic,” Chemistry-a European Journal 12:403-412 (2006), which is hereby incorporated by reference in its entirety).
- TAV2b-al Flu The higher binding affinity of TAV2b-al Flu to RNA is not surprising as the longer peptide contains an additional 8 cationic residues not present in CHF-l Flu , which likely form nonspecific electrostatic interactions with the oligonucleotide phosphate groups.
- Table 3 List of Oligonucleotides with Associated Sequence Length and Molecular Weight (MW) Table discloses SEQ ID NOS 26-33, respectively, in order of appearance.
- Cy5 Cyanine-5 labeled on 5 ’-end of r(CUG)io RNA [000184] It was hypothesized that if CHF-l Flu contacts the Hoogsteen face in the 5’-GC-3’ stretch, this peptide may bind to a CUG repeat sequence featuring multiple 5’-GC-3’ regions with higher affinity.
- the CUG repeat RNA, r(CUG)io, and RNA 1 differ by one nucleotide in spacing between the two sets of diagonal G’s recognized by the peptide dimer - the 5’-GC-3’ stretch is spaced by 5 bases in RNA 1 while GC stretches are spaced by four bases in r(CUG)io.
- CHF-l Flu contains a flexible region attached to cis-stilbene, which was predicted to allow recognition of a designed hairpin featuring CUG repeats.
- the Dimeric Helix is Required for High Affinity Duplex RNA Recognition
- a conformational constraint - a hydrogen bond surrogate - was incorporated in the backbone of the peptide to promote a-helicity (Wang et al., “Evaluation of Biologically Relevant Short Alpha-Helices Stabilized by a Main-Chain Hydrogen-Bond Surrogate,” J. Am. Chem. Soc. 128:9248-9256 (2006), which is hereby incorporated by reference in its entirety).
- HBS-1 did not lead to an enhancement in hairpin oligonucleotide recognition.
- Melamine group was conjugated to the peptide to obtain Peptide-Mel. Peptide-Mel was tested to determine whether addition of melamine group would provide enhanced affinity.
- RNA Binding Specificity of Crosslinked Helix Forks can be Tuned
- the potential of CHF-l Flu to discriminate between two closely related RNA sequences that differ by one contact nucleotide was determined and the potential of this ligand to differentiate between r(CUG)io and r(CUG)Mut was tested.
- r(CUG)wut is a single mismatch control in which one target guanine nucleotide is changed to an adenine.
- CHF-2 Flu was synthesized in which both lysine residues were replaced with alanine.
- CHF-2 F,U proved to be >50-fold more selective for r(CUG)io than to the single base mismatch sequence r(CUG)Mut.
- CHF-2 Flu also showed decreased binding for RNA- 1, DNA-1, and d(CTG)io hairpins suggesting potential non-specific interactions between CHF- l Flu and these oligonucleotides ( Figures 5A-B).
- a fluorescence polarization assay with fluorescein-labeled peptides was utilized for binding affinity analysis.
- the CHF scaffold was rationally designed to bind double-stranded RNA.
- hydroxyl radical footprinting with 5’-32P -radiolabeled RNA 1 and CHF-2 was performed.
- the cleavage pattern of RNA 1 on a denaturing electrophoresis gel was resolved ( Figure 6).
- the footprinting pattern indicated that regions (i), (ii), and (iv) of the hairpin RNA were not cleaved in the presence of CHF-2, as expected from binding of CHF-2 to the stem region of the hairpin RNA ( Figure 6).
- the footprinting result revealed that CHF-2 did not bind at the apical tetraloop, region (iii).
- the potential of CHF-2 to bind a singlestranded analog of RNA 1 was also evaluated using an electrophoretic mobility shift assay ( Figures 8 and 12). This analysis confirmed that CHF-2 does not recognize a single-stranded RNA.
- CHF-2 was probed via circular dichroism (CD).
- CD circular dichroism
- the CHF dimer adopted an a-helical structure in complex with duplex RNA.
- CHF-2 was weakly helical in aqueous solution - as would be expected of two short unconstrained peptides (Kallenbach et al., In Circular Dichroism and the Conformational Analysis of Biomolecules,' Fasman, G. D., Ed.; Plenum Press: New York, p 201-259 (1996), which is hereby incorporated by reference in its entirety).
- a dsRNA-binding scaffold, dubbed crosslinked helical fork (CHF), that mimics a protein tertiary structure often engaged in double-stranded DNA recognition was rationally designed.
- CHF crosslinked helical fork
- the specificity and affinity of the scaffold for dsRNA can be tuned through judicious modifications of the amino acid sequence, net charge, and crosslinker.
- broad potential for the CHF scaffold was envision capable of recognizing hairpin regions in RNA tertiary structure. For this potential to be met, the current CHF design needed to be improved.
- the model dictates helical motifs binding in the RNA major groove.
Landscapes
- Chemical & Material Sciences (AREA)
- Organic Chemistry (AREA)
- Genetics & Genomics (AREA)
- Biochemistry (AREA)
- Biophysics (AREA)
- General Health & Medical Sciences (AREA)
- Health & Medical Sciences (AREA)
- Medicinal Chemistry (AREA)
- Molecular Biology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Life Sciences & Earth Sciences (AREA)
- Peptides Or Proteins (AREA)
- Medicines That Contain Protein Lipid Enzymes And Other Medicines (AREA)
Abstract
The present invention relates to proteomimetic scaffolds. More specifically, the present invention is directed to proteomimetic scaffold and their uses for sequence-specific recognition and binding of duplex RNA. Compositions and dosage forms comprising the peptidomimetic scaffolds of the present invention are disclosed. Methods of using the proteomimetic scaffolds to target double-stranded RNA and to treat diseases and disorders are also disclosed.
Description
PROTEOMIMETIC SCAFFOLDS AND USES THEREOF
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63/521,234, filed June 15, 2023, the disclosure of which is herein incorporated by reference in its entirety.
STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT
[0002] This invention was made with government support under Grant R35 GM130333 awarded by the National Institute of Health (NIH). The government has certain rights in the invention.
SEQUENCE LISTING
[0003] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on June 10, 2024, is named 243735_000348_SL.xml and is 72,706 bytes in size.
FIELD OF INVENTION
[0004] The present invention relates to proteomimetic scaffolds. More specifically, the present invention is directed to proteomimetic scaffolds and their uses for sequence-specific recognition and binding of duplex RNA.
BACKGROUND OF THE INVENTION
[0005] Sequence-specific recognition of nucleic acids by synthetic ligands has been a focal point of vigorous research efforts for several decades. These efforts have yielded pyrrole-imidazole polyamides (Dervan, P. B. “Molecular Recognition of DNA by Small Molecules,” Bioorg. Med. Chem. 9:2215-35 (2001)) and zinefinger proteins (Wolfe et al., “DNA Recognition by Cys2His2 Zinc Finger Proteins,” Annual Review of Biophysics and Biomolecular Structure 29: 183-212 (2000)) as two textbook examples of scaffolds that afford sequence-specific recognition of DNA (Figure 1A). Efforts to recognize RNA have been complicated by the multitude of secondary and
tertiary structures adopted by this biopolymer (Ganser et al., “The Roles of Structural Dynamics in the Cellular Functions of RNAs,” Nature Reviews Molecular Cell Biology 20, 474-489 (2019); Batey et al., “Tertiary Motifs in RNA Structure and Folding,” Angew. Chem. Int. Ed. 38:2326- 2343 (1999)). Base-specific recognition of RNA is possible by antisense strategies (Crooke et al., “Antisense Technology: An Overview and Prospectus,” Nature Reviews Drug Discovery 20 ATI - 453 (2021); Rinaldi et al., “Antisense Oligonucleotides: the Next Frontier for Treatment of Neurological Disorders,” Nature Reviews Neurology 14:9-21 (2018)), including synthetic oligomers such as morpholinos (Corey et al., “Morpholino Antisense Oligonucleotides: Tools for Investigating Vertebrate Development,” Genome Biology llSy.xQvxosNS 1015.1-1015.3 (2001)) and peptide nucleic acids (PNAs) (Nielsen, P. E., “Peptide Nucleic Acids (PNA) in Chemical Biology and Drug Discovery,” Chemistry & Biodiversity 7 : 786-804 (2010)), and classes of small molecules (Tor, Y., “Targeting RNA with Small Molecules,” ChemBiochem 4:998-1007 (2003); Thomas et al., “Targeting RNA with Small Molecules,” Chem. Rev. 108:1171-1224 (2008); Warner et al., “Principles for Targeting RNA with Drug-Like Small Molecules,” Nature Reviews Drug Discovery 17:547-558 (2018); Meyer et al., “Small Molecule Recognition of Disease-Relevant RNA Structures,” Chem. Soc. Rev. 49:7167-7199 (2020); Boer et al., “Chemical Modulation of Pre-mRNA Splicing in Mammalian Systems,” ACS Chem. Biol. 15:808-818 (2020); Falese et al., “Targeting RNA with Small Molecules: From Fundamental Principles Towards the Clinic,” Chem. Soc. Rev. 50:2224-2243 (2021); Arambula et al., “A Simple Ligand that Selectively Targets CUG Trinucleotide Repeats and Inhibits MBNL Protein Binding,” Proc. Natl. Acad. Sci. 106: 16068- 16073 (2009)) that target RNA are growing, but a synthetic scaffold that allows atom-specific engineering to sequence-specifically target a double-stranded RNA motif is not yet available. Given the increasing importance of RNA as a therapeutic target, synthetic ligands that offer specific recognition properties hold significant promise.
[0006] Synthetic strategies to target RNA have ranged from aminoglycosides and their derivatives (Alper et al., “Probing the Specificity of Aminoglycoside-Ribosomal RNA Interactions with Designed Synthetic Analogs,” J. Am. Chem. Soc. 120: 1965-1978 (1998)), small molecules emerging from screens (Martin et al., “Screening Strategies for Identifying RNA- and Ribonucleoprotein-Targeted Compounds,” Trends Pharmacol. Sci. 42:758-771 (2021); Zafferani et al., “Small Molecule Targeting of Biologically Relevant RNA Tertiary and Quaternary Structures,” Cell Chem. Biol. 28:594-609 (2021); Gareiss et al., “Dynamic Combinatorial
Selection of Molecules Capable of Inhibiting the (CUG) Repeat RNA-MBNL1 Interaction In Vitro: Discovery of Lead Compounds Targeting Myotonic Dystrophy (DM1),” J. Am. Chem. Soc. 130: 16254-16261 (2008)), and multivalent intercalators (Guelev et al., “Peptide Bis-Intercalator Binds DNA Via Threading Mode with Sequence Specific Contacts in the Major Groove,” Chemistry & Biology 8:415-425 (2001); Carlson et al., “Preferred RNA Binding Sites for a Threading Intercalator Revealed by In Vitro Evolution,” Chemistry & Biology 10:663-672 (2003)) or hydrogen bonding scaffolds with potential to target repeat sequences (Pushechnikov et al., “Rational Design of Ligands Targeting Triplet Repeating Transcripts That Cause RNA Dominant Disease: Application to Myotonic Muscular Dystrophy Type 1 and Spinocerebellar Ataxia Type 3,” J. Am. Chem. Soc. 131 :9767-9779 (2009)).
[0007] Researchers have focused on mimicry of RNA-binding peptides and proteins as starting points for synthetic scaffolds. Structural efforts to characterize HIV RNAs provided two seductive examples for potential mimicry: the Rev peptide adopts an a-helical conformation to target RRE RNA (Battiste et al., “Alpha Helix- RNA Major Groove Recognition in an HIV-1 Rev Peptide RRE RNA Complex,” Science 273: 1547-1551 (1996)) and Tat can adopt a P-hairpin motif to bind TAR RNA (Figure IB) (Davidson et al., “Simultaneous Recognition of HIV-1 TAR RNA Bulge and Loop Sequences by Cyclic Peptide Mimics of Tat Protein,” Proc. Natl. Acad. Sci. USA 106: 11931-11936 (2009)).
[0008] Several studies to develop u-helix (Harrison et al., “Downsizing Human, Bacterial, and Viral Proteins to Short Water-Stable Alpha Helices that Maintain Biological Potency,” Proc. Natl. Acad. Sci. USA 107:11686-11691 (2010); Mills et al., “An a-Helical Peptidomimetic Inhibitor of the HIV-1 Rev-RRE Interaction,” J. Am. Chem. Soc. 128:3496-3497 (2006)) and flhairpin (Walker et al., “Design of RNA-Targeting Macrocyclic Peptides,” Rna Recognition 623:339-372 (2019)) mimics as specific RNA binding reagents have been detailed, but these efforts have failed to translate to a recognition code akin to a zinc-finger (ZNF) protein-based modular platform for DNA (Sera et al., “Rational Design of Artificial Zinc-Finger Proteins Using a Nondegenerate Recognition Code Table,” Biochemistry 41:7074-7081 (2002)). ZNF, and many other DNA-binding proteins, form specific side-chain interactions with the Hoogsteen surface of DNA bases (Neidle, S., “Principles of Protein-DNA Recognition,” Principles of Nucleic Acid Structure 249-282 (2008)). An adenine base can be recognized by carboxamide side chains of asparagine and glutamine, while the guanidine group of arginine forms specific hydrogen bonding
contacts with guanine (Figure 1C) (Seeman et al., “Sequence-Specific Recognition of Double Helical Nucleic Acids by Proteins,” Proc. Natl. Acad. Sci. USA 73:804-808 (1976)). An a-helix or P-hairpin motif, in principle, provides a similar opportunity for dsRNA recognition, but these protein secondary structure scaffolds are limited in the number of side chains they can project into the RNA grooves, which in turn lowers their potential as a general platform for sequence-specific nucleic acid binding.
[0009] Thus, there exists an unmet need for scaffolds for duplex RNA recognition, and in particular scaffolds for sequence-specific targeting of double-stranded RNA.
SUMMARY OF THE INVENTION
[00010] Various non-limiting aspects and embodiments of the invention are described below.
[00011] In one aspect, the present disclosure provides an encodable scaffold that engages the major groove of double-stranded RNA.
[00012] In another aspect, the present disclosure provides a compound having a structure according to Formula (I), Formula (II), Formula (III), Formula (IV), Formula (V), Formula (VI), or Formula (VII):
Formula (VII), or a pharmaceutically acceptable salt thereof, wherein:
Q is independently at each occurrence selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain;
Ri and R12 are independently selected from hydrogen, an amino acid, an amino acid dimer, each of which amino acid is optionally substituted with a fluorescent tag, a C1-C20 alkyl, a C5-C20 aryl, -N(R*)2, -(C=O)R*, -(C=S)R*, -(C=O)NR*, -(C=S)NR*, -CO2R*, -(C=S)OR*, -(C=O)SR*, and =C(R*)2;
R2, R3, R4, R5, R6, R7, R8, R9, Rio, R11, R12, R13, R14, R15, R16, R17, Ris, R19, R20, R21, and R22 are independently selected from hydrogen, a natural amino acid side chain, a nonnatural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, - OR*, and combinations thereof, where any two of R2, R3, R4, R5, R6, R7, R8, R9, R10, R11, R12, R13, R14, R15, R16, R17, Ris, R19, R20, R21, and R22 optionally bond together to make a macrocycle, and where R* is independently at each occurrence selected from hydrogen, a C5-C20 aryl, and a C1-C20 alkyl.
[00013] In another aspect, the present disclosure provides a method of binding doublestranded RNA. This method comprises contacting the double-stranded RNA with the compound according to the present disclosure, the composition according to the present disclosure, or the dosage form according to the present disclosure.
[00014] In another aspect, the present disclosure provides a method of treating a repeat expansion disorder in a subject in need of such treatment. This method comprises administering to the subject a therapeutically effective amount of the compound according to the present disclosure, the composition according to the present disclosure, or the dosage form according to the present disclosure.
[00015] In another aspect, the present disclosure provides a method for treating a Huntington’s disease in a subject in need of such treatment. This method comprises administering to the subject a therapeutically effective amount of the compound according to the present disclosure, the composition according to the present disclosure, or the dosage form according to the present disclosure.
[00016] These and other aspects of the present invention will become apparent to those skilled in the art after a reading of the following detailed description of the invention, including the appended claims.
BRIEF DESCRIPTION OF THE DRAWINGS
[00017] Figures 1A-1C depict nucleic acid recognition by small molecule and proteins. Fig. 1A shows DNA recognition by hairpin polyamide, pdb 3omj; zinc-finger protein, pdb Imey; and basic-leucine zipper, pdb Ifos. Fig. IB shows RNA recognition by Rev a-helix peptide, pdb letf, and cyclic Tat b-hairpin mimetic, pdb 2kdq. Fig. 1C shows Hoogsteen hydrogen bonding between arginine and guanine, and asparagine or glutamine with adenine.
[00018] Figures 2A-2B depict the design of the crosslinked helical fork (CHF) scaffold. Fig. 2A shows that from wild-type TAV2b protein, al helix was retained and truncated to an 8- residue sequence containing R26-R33 (ribbon). The individual al-derived peptides were covalently crosslinked to obtain the CHF scaffold. Fig. 2B shows side-chain interactions of R26- R33 with duplex RNA from the native complex, pdb 2zi0. Black dashes represent side-chain interaction with the RNA phosphate backbone. Wedges represent Hoogsteen hydrogen bonds between amino acid side chain to nucleobases, dotted box. Figure discloses SEQ ID NO: 3.
[00019] Figure 3A shows binding affinities of TAV2b-ulFlu and CHF-lFlu for designed oligonucleotide hairpins according to certain embodiments of the present invention, RNA 1 and DNA 1. TAV2b-alFlu and CHF-lFlu sequences are listed in brackets. Highlights on helices and amino acid sequence represent residues involved in Hoogsteen recognition. Nle = Norleucine. Ri
= Fluorescein-b-Alanine. R2 = Fluorescein-b-Alanyl-glycine. R3 = Acetyl -tyrosine. All peptide C- termini are carboxamides. Figure discloses SEQ ID NOS 36 and 40, respectively, in order of appearance. Figure 3B depicts designed hairpin oligonucleotides. Oligonucleotide sequences are listed in Table 3. Figure discloses SEQ ID NOS 26 and 27, respectively, in order of appearance. [00020] Figures 4A-4B outline a comparison of monomeric and dimeric peptide recognition of hairpin oligonucleotides. Figure 4A shows the sequence and secondary structure of oligonucleotides. Schematics were generated with NUPACK. Square box indicates guanine nucleobases engaged in Hoogsteen recognition. Figure discloses SEQ ID NOS 26-29, respectively, in order of appearance. Figure 4B shows binding affinities of fluorescein-labeled peptides for different hairpin oligonucleotides as measured in a fluorescence polarization assay. Figure discloses SEQ ID NOS 4, 37, and 5, respectively, in order of appearance. Ac = Acetyl. Flu = Fluorescein. X = 4-pentenoic acid. G* = A’-allylglycine. O = Melamine. All peptide C-termini are carboxamides.
[00021] Figures 5A-5B show analysis of specific and non-specific contacts in CHF’RNA recognition. Figure 5A shows that CHF-1 contains two solvent exposed lysine residues (K27 and K30) along with residues that potentially engage the RNA Hoogsteen surface; substitution of these cationic residues with alanine provided CHF-2. CHF-3 represents a negative control in which the nucleotide-engaging residues (R26 and H29) have been substituted with alanine. Models were generated with UCSF ChimeraX using pdb 2zi0. Binding affinities of fluorescein-labeled CHF compounds as determined in a fluorescence polarization assay. Bar graph of affinities represented in log(Ka). Figure discloses SEQ ID NOS 3, 38, and 39, respectively, in order if appearance. Figure 5B shows Kd values of CHF compounds. N.D. = not determined.
[00022] Figure 6 depicts hydroxyl radical footprinting of 5 ’-32 P-radiolabeled RNA 1 with CHF-2. Lane 1, 5’-32 P-radiolabeled RNA 1; lane 2, NaHCOs hydrolysis ladder; lane 3, RNase T1 standard; Lane 4, hydroxyl radical ( OH), lane 5-20, 2-fold serial dilution of 1 mM to 0.03 pM CHF-2. Footprint signatures are labeled in brackets (i) to (iv). Footprints corresponding to RNA 1 sequence are denoted with Roman numerals and dashed boxes. Overlay of CHF-2 corresponds to binding site for RNA 1. Figure discloses SEQ ID NO: 26.
[00023] Figures 7A-7C show an assessment of conformational changes in peptide (CHF- 2) and r(CUG)io upon complex formation. Figure 7A shows circular dichroism spectra of 30 pM CHF-2 (black) and signal of 7 pM r(CUG)io subtracted from CD spectrum of 10 pM CHF-2 and
7 nM r(CUG)io complex (purple). Figure 7B shows 7 pM r(CUG)io, 10 pM CHF-2 and 7 pM r(CUG)io complex and 4 pM of d(CTG)io. All measurements were conducted in 10 mM Tris and 10 mM KF, pH 7.4. Figure 7C shows a comparison of major grooves between TAV2b:siRNA complex (pdb 2zi0) and a canonical Bform DNA (pdb 3bse). Two overlaid views of the bound siRNA and DNA are shown. Dashed lines indicate major grooves.
[00024] Figure 8 shows secondary structures of ssRNA 1, RNA 1, DNA 1, r(CUG)io, d(CTG)io, and r(CUG)Mut as predicted and generated from NUPACK (Zadeh et al., “NUPACK: Analysis and Design of Nucleic Acid Systems,” J. Comput. Chem. 32(1): 170-173 (2011), which is hereby incorporated by reference in its entirety) and RNAfold (Lorenz et al., “ViennaRNA Package 2.0,” Algorithms Mol. Biol. 6:26 (2011), which is hereby incorporated by reference in its entirety). Black boxes indicate Hoogsteen recognition sites. Grey boxes represent guanine to adenine mutation. Figure discloses SEQ ID NOS 26-30, 32 and 33, respectively, in order of appearance.
[00025] Figure 9A shows analytical HPLC traces of peptides according to certain embodiments of the invention. Figure 9B shows analytical HPLC traces of Synthetic Analogues of CHF-2 according to certain embodiments of the invention.
[00026] Figure 10 shows fluorescence polarization assay binding measurements between peptide compounds according to embodiments of the present disclosure and oligonucleotides.
[00027] Figure 11 shows microscale thermophoresis binding measurement between CHF- 2 and 5 nM Cy5-r(CUG)io.
[00028] Figure 12 depicts an electrophoretic mobility gel shift assay for the direct binding of32P-ssRNA 1 and CHF-2.
[00029] Figure 13A describes direct binding of CHF-2Flu. Direct binding was performed using Fluorescence Polarization Assay between CHF-2Flu and listed oligonucleotides. Figure discloses SEQ ID NOS 28, 32 and 33, respectively, in order of appearance. Figure 13B describes binding of “Stapled” Analogue of CHF-2. Figure 13C describes binding of “D- isomer/RetroInverso” Analogue of CHF-2. Figure 13D describes binding of the “OHM” Analogue of CHF-2.
[00030] Figures 14A-14C show analysis of dsDNA and dsRNA recognition by Max homodimer, a basic helix loop motif. Figure 14A shows that Max homodimer binds to fluorescein- labeled Ebox DNA hairpin sequence (5 ’-FAM-CAC CAC GTG GTT TTT ACC ACG TGG TG-
3 ’ (SEQ ID NO: 1)) with high affinity (Ka = 0.04 ± 0.02 pM) in a fluorescence polarization assay, but shows no affinity for the analogous RNA hairpin sequence (5 ’-FAM-CAC CAC GUG GUU UUU ACC ACG UGG UGG ’ (SEQ ID NO: 2)). Figure 14B shows comparison of dsDNA and dsRNA major groove interactions with a-helices: (i) positioning of base pairs within major grooves; (ii-iv) structural analysis of DNA in complex with Max homodimer (PDB code Ihlo) and RNA in complex with Rev peptide (PDB code letf) and TAV2b (PDB code 2zi0) reveals that the angle of entry of dimeric oc -helical domains into dsDNA and dsRNA diverge. Figure 14C shows that replacement of the leucine zipper region in Max homodimer with a flexible linker leads to CHF-Max (Figures 15 and 16). CHF-Max binding to Ebox DNA and RNA was evaluated in a fluorescence polarization assay. Ka(CHF-Max/Ebox DNA): 0.69 ± 0.21 pM; Ka(CHF-Max/Ebox RNA): 0.48 ± 0.22 pM.
[00031] Figures 15A-15B show the distance between truncated helical dimers in DNA and RNA major grooves. Figure 15A shows crosslinking distance of 16 A for CHF-Max. Figure 15A was generated from PDB code Ihlo. Figure 15B shows crosslinking distance of 8 A for TAV2b- alFlu and 19 A for the minimal CHF constructs. Figure 15B was generated using PDB code 2zi0. [00032] Figure 16 shows the chemical structure of CHF-Max. Figure discloses SEQ ID NOS 23 and 24, respectively, in order of appearance.
DETAILED DESCRIPTION
[00033] Detailed embodiments of the present invention are disclosed herein; however, it is to be understood that the disclosed embodiments are merely illustrative of the invention that may be embodied in various forms. In addition, each of the examples given in connection with the various embodiments of the invention is intended to be illustrative, and not restrictive. Therefore, specific structural and functional details disclosed herein are not to be interpreted as limiting, but merely as a representative basis for teaching one skilled in the art to variously employ the present invention.
[00034] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[00035] As used in this specification and the appended claims, the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise. Thus, for
example, a reference to “a method” includes one or more methods, and/or steps of the type described herein and/or which will become apparent to those persons skilled in the art upon reading this disclosure.
[00036] The terms “treat” or “treatment” of a state, disorder or condition include: (1) preventing, delaying, or reducing the incidence and/or likelihood of the appearance of at least one clinical or sub-clinical symptom of the state, disorder or condition developing in a subject that may be afflicted with or predisposed to the state, disorder or condition but does not yet experience or display clinical or subclinical symptoms of the state, disorder or condition; or (2) inhibiting the state, disorder or condition, i.e., arresting, reducing or delaying the development of the disease or a relapse thereof or at least one clinical or sub-clinical symptom thereof; or (3) relieving the disease, i.e., causing regression of the state, disorder or condition or at least one of its clinical or sub-clinical symptoms. The benefit to a subject to be treated is either statistically significant or at least perceptible to the patient or to the physician.
[00037] A “subject” or “patient” or “individual” or “animal”, as used herein, refers to humans, veterinary animals (e.g., cats, dogs, cows, horses, sheep, pigs, etc.) and experimental animal models of diseases (e.g., mice, rats). In a preferred embodiment, the subject is a human.
[00038] As used herein the term “effective” applied to dose or amount refers to that quantity of a compound or pharmaceutical composition that is sufficient to result in a desired activity upon administration to a subject in need thereof. Note that when a combination of active ingredients is administered, the effective amount of the combination may or may not include amounts of each ingredient that would have been effective if administered individually. The exact amount required will vary from subject to subject, depending on the species, age, and general condition of the subject, the severity of the condition being treated, the particular drug or drugs employed, the mode of administration, and the like.
[00039] The phrase “pharmaceutically acceptable”, as used in connection with compositions of the invention, refers to molecular entities and other ingredients of such compositions that are physiologically tolerable and do not typically produce untoward reactions when administered to a mammal (e.g., a human). Preferably, as used herein, the term "pharmaceutically acceptable" means approved by a regulatory agency of the Federal or a state government or listed in the U.S. Pharmacopeia or other generally recognized pharmacopeia for use in mammals, and more particularly in humans.
[00040] Ranges can be expressed herein as from “about” or “approximately” one particular value and/or to “about” or “approximately” another particular value. When such a range is expressed, another embodiment includes from the one particular value and/or to the other particular value.
[00041] By “comprising” or “containing” or “including” is meant that at least the named compound, element, particle, or method step is present in the composition or article or method, but does not exclude the presence of other compounds, materials, particles, or method steps, even if the other such compounds, material, particles, or method steps have the same function as what is named.
[00042] Compounds of the present invention include those described generally herein, and are further illustrated by the classes, subclasses, and species disclosed herein. As used herein, the following definitions shall apply unless otherwise indicated. For purposes of this invention, the chemical elements are identified in accordance with the Periodic Table of the Elements, CAS version, Handbook of Chemistry and Physics, 75th Ed. Additionally, general principles of organic chemistry are described in “Organic Chemistry”, Thomas Sorrell, University Science Books, Sausalito: 1999, and “March's Advanced Organic Chemistry”, 5th Ed., Ed.: Smith, M.B. and March, J., John Wiley & Sons, New York: 2001 , the entire contents of which are hereby incorporated by reference.
[00043] The term “aliphatic” or “aliphatic group”, as used herein, means a straight-chain (i.e., unbranched) or branched, substituted or unsubstituted hydrocarbon chain that is completely saturated or that contains one or more units of unsaturation, or a monocyclic hydrocarbon, bicyclic hydrocarbon, or tricyclic hydrocarbon that is completely saturated or that contains one or more units of unsaturation, but which is not aromatic (also referred to herein as “carbocycle,” “cycloaliphatic” or “cycloalkyl”), that has a single point of attachment to the rest of the molecule. Unless otherwise specified, aliphatic groups contain 1—30 aliphatic carbon atoms. In some embodiments, aliphatic groups contain 1-20 aliphatic carbon atoms. In other embodiments, aliphatic groups contain 1-10 aliphatic carbon atoms. In still other embodiments, aliphatic groups contain 1-6 aliphatic carbon atoms, and in yet other embodiments, aliphatic groups contain 1, 2, 3, or 4 aliphatic carbon atoms. Suitable aliphatic groups include, but are not limited to, linear or branched, substituted or unsubstituted alkyl, alkenyl, alkynyl groups and hybrids thereof such as (cycloalkyl)alkyl, (cycloalkenyl)alkyl or (cycloalkyl)alkenyl.
[00044] The term “cycloaliphatic,” as used herein, refers to saturated or partially unsaturated cyclic aliphatic monocyclic, bicyclic, or polycyclic ring systems, as described herein, having from 3 to 14 members, wherein the aliphatic ring system is optionally substituted as defined above and described herein. Cycloaliphatic groups include, without limitation, cyclopropyl, cyclobutyl, cyclopentyl, cyclopentenyl, cyclohexyl, cyclohexenyl, cycloheptyl, cycloheptenyl, cyclooctyl, cyclooctenyl, norbomyl, adamantyl, and cyclooctadienyl. In some embodiments, the cycloalkyl has 3-6 carbons. The terms “cycloaliphatic,” may also include aliphatic rings that are fused to one or more aromatic or nonaromatic rings, such as decahydronaphthyl or tetrahydronaphthyl, where the radical or point of attachment is on the aliphatic ring. In some embodiments, a carbocyclic group is bicyclic. In some embodiments, a 'carbocyclic group is tricyclic. In some embodiments, a carbocyclic group is polycyclic. In some embodiments, “cycloaliphatic” (or “carbocycle” or “cycloalkyl”) refers to a monocyclic C3-C6 hydrocarbon, or a Cs-Cio bicyclic hydrocarbon that is completely saturated or that contains one or more units of unsaturation, but which is not aromatic, that has a single point of attachment to the rest of the molecule, or a C9-C16 tricyclic hydrocarbon that is completely saturated or that contains one or more units of unsaturation, but which is not aromatic, that has a single point of attachment to the rest of the molecule.
[00045] As used herein, the term “alkyl” is given its ordinary meaning in the art and may include saturated aliphatic groups, including straight-chain alkyl groups, branched-chain alkyl groups, cycloalkyl (alicyclic) groups, alkyl substituted cycloalkyl groups, and cycloalkyl substituted alkyl groups. In certain embodiments, a straight chain or branched chain alkyl has about 1-20 carbon atoms in its backbone (e.g., C1-C20 for straight chain, C2-C20 for branched chain), and alternatively, about 1-10 carbon atoms, or about 1 to 6 carbon atoms. In some embodiments, a cycloalkyl ring has from about 3-20 carbon atoms in their ring structure where such rings are monocyclic or bicyclic, and alternatively about 5, 6 or 7 carbons in the ring structure. In some embodiments, an alkyl group may be a lower alkyl group, wherein a lower alkyl group comprises 1-4 carbon atoms (e.g., C1-C4 for straight chain lower alkyls).
[00046] As used herein, the term “alkenyl” refers to an alkyl group, as defined herein, having one or more double bonds.
[00047] As used herein, the term “alkynyl” refers to an alkyl group, as defined herein, having one or more triple bonds.
[00048] The term “heteroalkyl” is given its ordinary meaning in the art and refers to alkyl groups as described herein in which one or more carbon atoms is replaced with a heteroatom (e.g., oxygen, nitrogen, sulfur, and the like). Examples of heteroalkyl groups include, but are not limited to, alkoxy, poly(ethylene glycol)-, alkyl -substituted amino, tetrahydrofuranyl, piperidinyl, morpholinyl, etc.
[00049] The term “heterocycloalkyl” is given its ordinary meaning in the art and refers to cycloalkyl groups as described herein in which one or more carbon atoms is replaced with a heteroatom (e.g., oxygen, nitrogen, sulfur, and the like). Examples of heterocycloalkyl groups include, but are not limited to, tetrahydrofuranyl, piperidinyl, morpholinyl, etc.
[00050] The term “aryl” used alone or as part of a larger moiety as in “aralkyl,” “aralkoxy,” or “aryloxyalkyl,” refers to monocyclic or bicyclic ring systems having a total of five to fourteen ring members, wherein at least one ring in the system is aromatic and wherein each ring in the system contains 3 to 7 ring members. The term “aryl” may be used interchangeably with the term “aryl ring.” In certain embodiments of the present invention, “aryl” refers to an aromatic ring system which includes, but not limited to, phenyl, biphenyl, naphthyl, binaphthyl, anthracyi and the like, which may bear one or more substituents. Also included within the scope of the term “aryl,” as it is used herein, is a group in which an aromatic ring is fused to one or more nonaromatic rings, such as indanyl, phthalimidyl, naphthimidyl, phenanthridinyl, or tetrahydronaphthyl, and the like.
[00051] The term “arylene” means a group obtained by removal of a hydrogen atom from an aryl group. Non-limiting examples of arylene include phenylene and naphthylene.
[00052] The terms “heteroaryl” and “heteroar-,” used alone of as part of a larger moiety, e.g., “heteroaralkyl,” or “heteroaralkoxy,” refer to groups having 5 to 10 ring atoms (i.e., monocyclic or bicyclic), in some embodiments 5, 6, 9, or 10 ring atoms. In some embodiments, such rings have 6, 10, or 147t electrons shared in a cyclic array; and having, in addition to carbon atoms, from one to five heteroatoms. The term “heteroatom” refers to nitrogen, oxygen, or sulfur, and includes any oxidized form of nitrogen or sulfur, and any quatemized form of a basic nitrogen. Heteroaryl groups include, without limitation, thienyl, furanyl, pyrrolyl, imidazolyl, pyrazolyl, triazolyl, tetrazolyl, oxazolyl, isoxazolyl, oxadiazolyl, thiazolyl, isothiazolyl, thiadiazolyl, pyridyl, pyridazinyl, pyrimidinyl, pyrazinyl, indolizinyl, purinyl, naphthyridinyl, and pteridinyl. In some embodiments, a heteroaryl is a heterobiaryl group, such as bipyridyl and the like. The terms
“heteroaryl” and “heteroar-”, as used herein, also include groups in which a heteroaromatic ring is fused to one or more aryl, cycloaliphatic, or heterocyclyl rings, where the radical or point of attachment is on the heteroaromatic ring. Nonlimiting examples include indolyl, isoindolyl, benzothienyl, benzofuranyl, dibenzofuranyl, indazolyl, benzimidazolyl, benzthiazolyl, quinolyl, isoquinolyl, cinnolinyl, phthalazinyl, quinazolinyl, quinoxalinyl, 4H — quinolizinyl, carbazolyl, acridinyl, phenazinyl, phenothiazinyl, phenoxazinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, and pyrido[2,3-b]-l,4-oxazin-3(4H)-one. A heteroaryl group may be monocyclic, bicyclic, tricyclic, tetracyclic, and/or otherwise polycyclic. The term “heteroaryl” may be used interchangeably with the terms “heteroaryl ring,” “heteroaryl group,” or “heteroaromatic,” any of which terms include rings that are optionally substituted. The term “heteroaralkyl” refers to an alkyl group substituted by a heteroaryl, wherein the alkyl and heteroaryl portions independently are optionally substituted.
[00053] As used herein, the terms “heterocycle,” “heterocyclyl,” “heterocyclic radical,” and “heterocyclic ring” are used interchangeably and refer to a stable 5- to 7-membered monocyclic or 7-10-membered bicyclic heterocyclic moiety that is either saturated or partially unsaturated, and having, in addition to carbon atoms, one or more, preferably one to four, heteroatoms, as defined above. When used in reference to a ring atom of a heterocycle, the term “nitrogen” includes a substituted nitrogen.
[00054] A heterocyclic ring can be attached to its pendant group at any heteroatom or carbon atom that results in a stable structure and any of the ring atoms can be optionally substituted. Examples of such saturated or partially unsaturated heterocyclic radicals include, without limitation, tetrahydrofuranyl, tetrahydrothiophenyl pyrrolidinyl, piperidinyl, pyrrolinyl, tetrahydroquinolinyl, tetrahydroisoquinolinyl, decahydroquinolinyl, oxazolidinyl, piperazinyl, dioxanyl, dioxolanyl, diazepinyl, oxazepinyl, thiazepinyl, morpholinyl, and quinuclidinyl. The terms “heterocycle,” “heterocyclyl,” “heterocyclyl ring,” “heterocyclic group,” “heterocyclic moiety,” and “heterocyclic radical,” are used interchangeably herein, and also include groups in which a heterocyclyl ring is fused to one or more aryl, heteroaryl, or cycloaliphatic rings, such as indolinyl, 3H-indolyl, chromanyl, phenanthridinyl, or tetrahydroquinolinyl. A heterocyclyl group may be monocyclic, bicyclic, tricyclic, tetracyclic, and/or otherwise polycyclic. The term “heterocyclylalkyl” refers to an alkyl group substituted by a heterocyclyl, wherein the alkyl and heterocyclyl portions independently are optionally substituted.
[00055] As used herein, the term “partially unsaturated” refers to a ring moiety that includes at least one double or triple bond. The term “partially unsaturated” is intended to encompass rings having multiple sites of unsaturation, but is not intended to include aryl or heteroaryl moieties, as herein defined.
[00056] The term “heteroatom” means one or more of oxygen, sulfur, nitrogen, phosphorus, or silicon (including, any oxidized form of nitrogen, sulfur, phosphorus, or silicon; the quaternized form of any basic nitrogen or; a substitutable nitrogen of a heterocyclic ring.
[00057] The term “unsaturated,” as used herein, means that a moiety has one or more units of unsaturation.
[00058] The term “halogen” means F, Cl, Br, or I; the term “halide” refers to a halogen radical or substituent, namely -F, -Cl, -Br, or -I.
[00059] As described herein, compounds of the invention may contain “optionally substituted” moieties. In general, the term “substituted,” whether preceded by the term “optionally” or not, means that one or more hydrogens of the designated moiety are replaced with a suitable substituent. Unless otherwise indicated, an “optionally substituted” group may have a suitable substituent at each substitutable position of the group, and when more than one position in any given structure may be substituted with more than one substituent selected from a specified group, the substituent may be either the same or different at every position. Combinations of substituents envisioned by this invention are preferably those that result in the formation of stable or chemically feasible compounds. The term “stable,” as used herein, refers to compounds that are not substantially altered when subjected to conditions to allow for their production, detection, and, in certain embodiments, their recovery, purification, and use for one or more of the purposes disclosed herein.
[00060] The term "amino acid" further includes analogues, derivatives, and congeners of any specific amino acid referred to herein, as well as C-terminal or N-terminal protected amino acid derivatives (e.g., modified with an N-terminal or C-terminal protecting group). Furthermore, the term “amino acid” includes both D- and L- amino acids. Hence, an amino acid which is identified herein by its name, three letter or one letter symbol and is not identified specifically as having the D or L configuration, is understood to assume any one of the D or L configurations. In certain embodiments, peptides containing only L-amino acids are contemplated.
[00061] Naturally occurring amino acids are identified throughout by the conventional three-letter and/or one-letter abbreviations, corresponding to the trivial name of the amino acid, in accordance with the following list: Alanine (Ala), Arginine (Arg), Asparagine (Asn), Aspartic acid (Asp), Cysteine (Cys), Glutamic acid (Glu), Glutamine (Gin), Glycine (Gly), Histidine (His), Isoleucine (He), Leucine (Leu), Lysine (Lys), Methionine (Met), Phenylalanine (Phe), Proline (Pro), Serine (Ser), Threonine (Thr), Tryptophan (Trp), Tyrosine (Tyr), and Valine (Vai). The abbreviations are accepted in the peptide art and are recommended by the IUPAC-IUB commission in biochemical nomenclature.
[00062] Unless otherwise stated, structures depicted herein are also meant to include all isomeric (e.g., enantiomeric, diastereomeric, and geometric (or conformational)) forms of the structure; for example, the R and S configurations for each asymmetric center, (Z) and (E) double bond isomers, and (Z) and (E) conformational isomers. Therefore, single stereochemical isomers as well as enantiomeric, diastereomeric, and geometric (or conformational) mixtures of the present compounds are within the scope of the invention.
[00063] Unless otherwise stated, all tautomeric forms of the compounds of the invention are within the scope of the invention.
[00064] Additionally, unless otherwise stated, structures depicted herein are also meant to include compounds that differ only in the presence of one or more isotopically enriched atoms. For example, compounds having the present structures except for the replacement of hydrogen by deuterium or tritium, or the replacement of a carbon by a nC- or 13C- or 14C -enriched carbon are within the scope of this invention.
[00065] It is also to be understood that the mention of one or more method steps does not preclude the presence of additional method steps or intervening method steps between those steps expressly identified. Similarly, it is also to be understood that the mention of one or more components in a device or system does not preclude the presence of additional components or intervening components between those components expressly identified.
[00066] Unless otherwise stated, all crystalline forms of the compounds of the invention and salts thereof are also within the scope of the invention. The compounds of the invention may be isolated in various amorphous and crystalline forms, including without limitation forms which are anhydrous, hydrated, non-solvated, or solvated. Example hydrates include hemihydrates, monohydrates, dihydrates, and the like. In some embodiments, the compounds of the invention are
anhydrous and non-solvated. By "anhydrous" is meant that the crystalline form of the compound contains essentially no bound water in the crystal lattice structure, i.e., the compound does not form a crystalline hydrate.
[00067] As used herein, "crystalline form" is meant to refer to a certain lattice configuration of a crystalline substance. Different crystalline forms of the same substance typically have different crystalline lattices (e.g., unit cells) which are attributed to different physical properties that are characteristic of each of the crystalline forms. In some instances, different lattice configurations have different water or solvent content. The different crystalline lattices can be identified by solid state characterization methods such as by X-ray powder diffraction (PXRD). Other characterization methods such as differential scanning calorimetry (DSC), thermogravimetric analysis (TGA), dynamic vapor sorption (DVS), solid state NMR, and the like further help identify the crystalline form as well as help determine stability and solvent/water content.
[00068] Crystalline forms of a substance include both solvated (e.g., hydrated) and nonsolvated (e.g., anhydrous) forms. A hydrated form is a crystalline form that includes water in the crystalline lattice. Hydrated forms can be stoichiometric hydrates, where the water is present in the lattice in a certain water/molecule ratio such as for hemihydrates, monohydrates, dihydrates, etc. Hydrated forms can also be non-stoichiometric, where the water content is variable and dependent on external conditions such as humidity.
[00069] In some embodiments, the compounds of the invention are substantially isolated. By "substantially isolated" is meant that a particular compound is at least partially isolated from impurities. For example, in some embodiments a compound of the invention comprises less than about 50%, less than about 40%, less than about 30%, less than about 20%, less than about 15%, less than about 10%, less than about 5%, less than about 2.5%, less than about 1%, or less than about 0.5% of impurities. Impurities generally include anything that is not the substantially isolated compound including, for example, other crystalline forms and other substances.
[00070] RNA, unlike DNA, folds into a multitude of secondary and tertiary structures. This structural diversity has impeded the development of ligands that can sequence-specifically target this biomolecule.
[00071] The present invention is in the development of a scaffold for duplex RNA recognition is based on surprising and unexpected discovery of synthetic helical fork motifs that can be engineered to sequence-specifically target dsRNA.
[00072] The present invention is based upon the inventor’s discovery that protein tertiary structure mimic may offer an opportunity to specifically target dsRNA. (Home et al., “Proteomimetics as Protein-Inspired Scaffolds with Defined Tertiary Folding Patterns,” Nat. Chem. 12:331-337 (2020); Sadek et al., “Modulation of Virus-Induced NF-KB Signaling by NEMO Coiled Coil Mimics,” Nat. Comniun. 11 : 1786 (2020); Watkins et al., “Protein-Protein Interactions Mediated by Helical Tertiary Structure Motifs,” J. Am. Chem. Soc. 137: 11622-11630 (2015), which are hereby incorporated by reference in their entirety). The helix dimer represents the simplest helical tertiary structure and forms the basis of important DNA recognition motifs in the bZIP and bHLH families (Figure 1A). Both transcription factor families feature a Y-shaped scissor-grip or helical fork model of DNA recognition (Vinson et al., “Scissors-Grip Model for DNA Recognition by a Family of Leucine Zipper Proteins,” Science 246:911-916 (1989); Oneil et al., “Design of DNA- Binding Peptides Based on the Leucine Zipper Motif,” Science 249:774-778 (1990), which are hereby incorporated by reference in their entirety).
[00073] It has been hypothesized by the present inventors that a synthetic helical fork (HF) motif may provide a generalized scaffold to target dsRNA. Unlike a single a-helical motif, an HF scaffold would have the advantage of clamping neighboring major grooves with two helices. Although, several examples of a-helical domains binding the RNA major groove are known, the RNA major groove is narrower and deeper than B-form DNA’s and often requires bulges to accommodate a-helices (Battiste et al., “Alpha Helix- RNA Major Groove Recognition in an HIV- 1 Rev Peptide RRE RNA Complex,” Science 273: 1547-1551 (1996); Wu et al., “Structural Basis for Recognition of the AGNN Tetraloop RNA Fold by the Double-Stranded RNA-Binding Domain of Rntlp RNase III,” Proc. Natl. Acad. Sci. USA 101 :8307-8312 (2004), which are hereby incorporated by reference in their entirety).
[00074] A key impediment in developing HF mimics as dsRNA binders is that natural examples of dsRNA bound by HF motifs are scarce. For example, bZIP and bHLH family proteins are not known to target RNA. A search of the literature and structural data repositories yielded one example of a Y-shaped helical motif from tomato aspermy virus protein 2b (TAV2b) bound to an siRNA duplex (Figure 2A) (Chen et al., “Structural Basis for RNA-Silencing Suppression by
Tomato Aspermy Virus Protein 2b,” Embo Reports 9:754-760 (2008), which is hereby incorporated by reference in its entirety).
[00075] The basic leucine zipper (bZIP) and basic helix-loop-helix (bHLH) families of DNA binding proteins serve as inspirations for the proteomimetic design of the present disclosure. The helix dimer represents the simplest helical tertiary structure and forms the basis of important DNA recognition motifs in the bZIP and bHLH families (Figures 14A-C). Both transcription factor families feature a Y-shaped scissor-grip or helical fork model of DNA recognition (Vinson et al., "Scissors-Grip Model for DNA Recognition by a Family of Leucine Zipper Proteins," Science 246:911-916 (1989); Oneil et al., "Design ofDNA-Binding Peptides Based on the Leucine Zipper Motif," Science 249:774-778 (1990), which are hereby incorporated by reference in their entirety). It was rationalized that a synthetic helical fork (HF) motif may provide a generalized scaffold to target dsRNA. Unlike a single a-helical motif, a helix fork scaffold would have the advantage of clamping neighboring major grooves with two helices.
[00076] The challenge with reengineering these DNA binding proteins as RNA binders is that these motifs are not generally known to bind dsRNA. The studies started by analyzing why bHLH may recognize duplex DNA but not RNA (Figure 14A). First, it was confirmed that Max»Max homodimer recognizes a fluorescein-labeled hairpin DNA motif, referred as Ebox DNA, that encompasses the Max binding site but not the analogous RNA hairpin, Ebox RNA (Figure 14A and Table 1). Max homodimer binds Ebox DNA with high affinity (KD = 0.04 ± 0.02 pM) in a fluorescence polarization assay but fails to show any binding to RNA at up to 0.5 pM concentration.
[00077] Comparison of the A-form and B-form configurations reveals the difference between RNA versus DNA recognition by bHLH proteins in their respective major grooves. The geometries of the B-form versus A-form helices dictate that base pairs lie at the center of helical DNA major groove but towards the edge of the minor groove in the RNA duplex (Figure 14B-i). The Hoogsteen surface in DNA major groove can be readily accessed by side chains projecting from the a-helix backbone but a similar hydrogen-bonding contact with the nucleobase requires the a-helix to bury itself deeper into the RNA major groove. This critical aspect in the a-helix- nucleobase recognition of DNA and RNA major grooves is illustrated by complexes of Max -DNA and Rev-RNA (Figure 14B-ii). The Rev a-helix is buried into the major groove as compared to the Max a-helix.
[00078] The relative orientation of the a -helix positioned into the A-form major groove is distinct from that of the DNA B-form axis. This geometrical difference is most clearly observed when two a-helices are positioned into neighboring major grooves of RNA and DNA. A set of non-parallel dimeric a-helices would adopt inverted acute angles to access contiguous DNA as opposed to RNA major grooves (Figure 14B-iii). The DNA-binding region of Max emerge from a rigid helix-loop region (Figure 14B-iv) to complex with DNA. However, the leucine zipper dimerization domain imposes an angle restriction that hinders complex formation with RNA. There are limited examples of dimeric a-helices that bind RNA; the search of the literature and structural data repositories yielded one example of a Y-shaped a-helical motif from tomato aspermy virus protein 2b (TAV2b) bound to an siRNA duplex (Chen et al., "Structural Basis for RNA-Silencing Suppression by Tomato Aspermy Virus Protein 2b," Embo Reports 9:754-760 (2008), which is hereby incorporated by reference in its entirety). TAV2b a-helices mimic the trajectory of Rev a-helix binding geometry in A-form RNA and adopt the wider entry angle required for RNA complexation.
[00079] Without wishing to be bound by theory, it is postulated that Max dimer could be engineered to bind RNA if the angle between the major groove binding Max a-helices was modified by replacement of the leucine zipper domain with artificial crosslinkers. The DNA binding basic domain of Max was crosslinked with a flexible GlyGlyCys-benzyl crosslinker to span the ~16 A distance between major grooves (Figures 15A-15B); the N-terminus was modified with cysteine residues and crosslinked the two a-helices with 1,3-dibenzyl bromide to obtain crosslinked helical fork (CHF) from the Max sequence (CHF-Max). The chemical structure CHF- Max is depicted in Figure 16. Titration of CHF-Max with Ebox DNA and Ebox RNA shows that zipper-less synthetic motif binds RNA and DNA sequences with similar affinity (KD ~ 0.5 - 0.7 pM; respectively) in a fluorescence polarization assay (Figure 14C). These results support the hypothesis that the rigidity of the zipper domain restricts the entry angle required for RNA binding - a restriction that can be overcome with a flexible linker.
[00080] Synthetic mimics of TAV2b were prepared by the present inventors to determine their potential to recognize duplex RNA. These studies suggest that the helix fork motif, reminiscent of the bZZP and bHLH DNA-binding proteins offers a lead scaffold to develop sequence- and structure-specific RNA binding motifs (Figure 2A).
[00081] According to the foregoing objective and others, the present disclosure provides an encodable scaffold that engages the major groove of double-stranded RNA. The present disclosure also provides the compounds that may mimic TAV2b, pharmaceutical compositions comprising these compounds, and methods for treating certain diseases in a subject in need of such treatment. Also provided in the present disclosure methods of binding double-stranded RNA.
Compounds of the Invention
[00082] Minimal motif that mimics an RNA binding viral protein TAV2b consists of two segmented a-helical regions, defined as al and a2, that homodimerize in the major groove of the siRNA duplex (Figure 2A). The binding motif observed in the two al helices closely resembles that of the bZIP transcription factor.
[00083] A crosslinked helical fork (CHF) scaffold has been designed that can mimic al’s affinity for its cognate RNA sequence. The minimal peptide sequence RKRHKLNR (SEQ ID NO: 3) was grafted from al onto the scaffold. While the native peptide-RNA complex features several electrostatic contacts with the negatively charged RNA backbone, it contains two hydrogen bonding interactions on the Hoogsteen face of two guanine bases in the duplex (5’-GC-3’) sequence (Figure 2B). The base specific hydrogen-bonding in the major groove is critical for specificity in protein-nucleic acids recognition.
[00084] In one aspect, provided herein is an encodable scaffold that engages the major groove of double-stranded RNA.
[00085] As used herein the term “encodable” refers to a scaffold that can be encoded with the correct hydrogen-bonding functionality to contact RNA bases on the Hoogsteen or Watson - Crick faces.
[00086] As used herein the term “encodable scaffold” refers to functional groups appended onto the scaffold that allows precise and predetermined targeting of given RNA duplex.
[00087] In at least one embodiment, the scaffold comprises a first domain DI and a second domain D2, wherein the first domain and the second domain are connected by a covalent linker L, and where the scaffold spatially fits in the major groove of double-stranded RNA.
[00088] In at least one embodiment, the scaffold binds to the major groove of the doublestranded RNA using base-specific bonding selected from the group consisting of Hoogsteen hydrogen bonding and electrostatic contact.
[00089] In some embodiments, the first domain DI and the second domain D2 are each helical oligomers or helix mimics.
[00090] According to the present disclosure, the first domain DI and the second domain D2 each comprise an oligomer comprising 2 to 20 units. DI can comprise an oligomer comprising 2 to 20 units, 2 to 19 units, 2 to 18 units, 2 to 17 units, 2 to 16 units, 2 to 15 units, 2 to 14 units, 2 to 13 units, 2 to 12 units, 2 to 11 units, or 2 to 10 units. For example, DI can comprise an oligomer comprising 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 units. D2 can comprise an oligomer comprising 2 to 20 units, 2 to 19 units, 2 to 18 units, 2 to 17 units, 2 to 16 units, 2 to 15 units, 2 to 14 units, 2 to 13 units, 2 to 12 units, 2 to 11 units, or 2 to 10 units. For example, D2 can comprise an oligomer comprising 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 units.
[00091] In at least one embodiment, DI and the D2 each comprise an oligomer comprising 2 to 20 units, wherein each oligomer is independently selected from a peptide, a nucleic acid, and an oligomer comprising repeating C1-C20 alkyl, C5-C20 aryl, C1-C20 heteroalkyl, C2-C20 heteroaryl, C3-C20 cycloalkyl, or C1-C20 heterocycloalkyl units, and combinations thereof.
[00092] In some embodiments, the DI and the D2 each comprise an oligomer comprising 2 to 10 C1-C20 heterocycloalkyl units, wherein each unit is substituted with a side chain, wherein each side chain is independently at each occurrence selected from hydrogen, a natural amino acid side chain, a non-natural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof, where R* is independently at each occurrence selected from hydrogen and a C1-C20 alkyl.
[00093] In some embodiments, the DI and the D2 each comprise an oligomer comprising 2 to 10 oxopiperazine units. For example, the DI and the D2 each comprise an oligomer comprising 2, 3, 4, 5, 6, 7, 8, 9, or 10 oxopiperazine units.
[00094] According to the present disclosure, the DI and the D2 can each comprise a peptide consisting of 7 to 10 amino acids. DI can comprise a peptide consisting of 7 to 10 amino acids. For example, DI can comprise a peptide consisting of 7, 8, 9, or 10 amino acids. D2 can comprise a peptide consisting of 7 to 10 amino acids. For example, D2 can comprise a peptide consisting of 7, 8, 9, or 10 amino acids.
[00095] In some embodiments, the DI and the D2 each comprise a peptide consisting of 7 to 10 amino acids, wherein each amino acid side chain is independently selected from hydrogen, a natural amino acid side chain, a non-natural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof, wherein R* is independently at each occurrence selected from hydrogen and a C1-C20 alkyl.
[00096] In at least one embodiment, the DI and the D2 each comprise a peptide consisting of 8 or 9 amino acids.
[00097] In some embodiments, the DI and the D2 each comprise a peptide consisting of 8 amino acids having a formula aal-aa2-aa3-aa4-aa5-aa6-aa7-aa8, wherein aal, aa2, aa3, aa4, aa5, aa6, aa7, and aa8 are each an amino acid having a side chain independently selected from hydrogen, a natural amino acid side chain, a non-natural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof.
[00098] In some embodiments, the DI and the D2 each comprise a peptide having a formula R-aa2-aa3-H-aa5-aa6-aa7-aa8, wherein R is arginine or a compound having the formula:
wherein X is CH2 or NH, and n is 0, 1, or 2; and H is histidine.
[00099] In some embodiments, the DI and the D2 each comprise a peptide consisting of 9 amino acids having a formula aa0-aal-aa2-aa3-aa4-aa5-aa6-aa7-aa8, wherein aaO, aal, aa2, aa3, aa4, aa5, aa6, aa7, and aa8 are each an amino acid having a side chain independently selected from hydrogen, a natural amino acid side chain, a non-natural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof.
[000100] In at least one embodiment, the DI and the D2 each comprise a peptide having a formula G-R-aa2-aa3-H-aa5-aa6-aa7-aa8, wherein G is glycine, H is histidine, and R is as defined above.
[000101] According to the present disclosure, the first domain DI and a second domain D2 can consist of both natural and non-natural amino acids. For example, the DI and the D2 each
comprise a peptide comprising at least one non-natural amino acid, at least two non-natural amino acids, at least three non-natural amino acids, or at least four non-natural amino acids.
[000102] In at least one embodiment, the DI and the D2 each comprise a peptide comprising at least one non-natural amino acid.
[000103] In some embodiments, the DI and the D2 each comprise an a-helical peptide or peptidomimetic.
[000104] According to the present disclosure, there is a spatial distance between the first domain DI and a second domain D2. This special distance can be from about 2A to about 30 A, from about 5A to about 25 A, or from about 7 A to about 20 A. In at least one embodiment, the spatial distance between the DI and the D2 is from about 5 A to about 25 A.
[000105] In some embodiments, the spatial distance between aal of DI and aal of D2 is about 5A to about 15 A, and the spatial distance between aa8 of DI and aa8 of D2 is about 15A to about 25 A.
[000106] According to the present disclosure, there is an angle between the first domain DI and the second domain D2. This angle can be between about 40° and about 80°, between about 50° and about 70°, between about 55° and about 65°, between about 60° and about 65°, or between about 63° and about 64°. For example, the angle between the DI and the D2 can be about 55°, about 56°, about 57°, about 58°, about 59°, about 60°, about 61°, about 62°, about 63°, about 64°, about 65°, about 66°, about 67°, about 68°, about 69°, or about 70°.
[000107] In one embodiment, the angle between the first domain DI and the second domain D2 is between about 50° and about 70°. In another embodiment, the angle between DI and D2 is about 60° to about 65°. In yet another embodiment, the angle between DI and D2 is about 63.7°. [000108] According to the present disclosure, the first domain DI and the second domain D2, are connected by a covalent linker L. This linker L can have a length of about 2A to about 30 A, from about 5 A to about 25 A, from about 5 A to about 20 A, about 5 A to about 15 A, or about 7 A to about 10 A. In at least one embodiment, the linker L is about 5 A to about 15 A in length.
[000109] According to the present disclosure, linker L can be a reversible linker or an irreversible linker.
[000110] In some embodiments, the linker L comprises a C1-C20 alkyl, a C5-C20 aryl, a Ci- C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, each of which
may be optionally substituted by one or more of halo, N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof.
[000111] In some embodiments, the linker L comprises a combination of a C1-C20 alkyl and a C5-C20 aryl.
[000112] In some embodiments, the linker L comprises a cis-stilbene, azobenzene, metaxylene, l,3-di( 1H-l,2,3-triazol-l-yl)benzene, l,3-di(l//-l,2,3-triazol-4-yl)benzene, or 1,4- diphenyl- 1H-l,2,3-triazole.
[000113] In some embodiments, the linker L has a structure selected from:
wherein Q is independently at each occurrence selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain, and
represents the point of attachment to DI and D2.
[000114] In some embodiments, the reversible linker comprises a disulfide (-S-S-) bond.
[000115] In some embodiments, the encodable scaffold has a structure according to Scheme
(I):
[000116] In another aspect, the present disclosure provides a compound having a structure according to Formula (I), Formula (II), Formula (III), Formula (IV), Formula (V), Formula (VI), or Formula (VII):
Q is independently at each occurrence selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain;
Ri and R12 are independently selected from hydrogen, an amino acid, an amino acid dimer, each of which amino acid is optionally substituted with a fluorescent tag, a C1-C20 alkyl, a C5-C20 aryl, -N(R*)2, -(C=O)R*, -(C=S)R*, -(C=O)NR*, -(C=S)NR*, -CO2R*, -(C=S)OR*, -(C=O)SR*, and =C(R*)2;
R2, R3, R4, R5, R6, R7, R8, R9, Rio, R11, R12, R13, R14, R15, R16, R17, Ris, R19, R20, R21, and R22 are independently selected from hydrogen, a natural amino acid side chain, anon-natural amino acid side chain, a Ci-C2o alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof, where any two of R2, R3, R4, R5, Re, R7, R8, R9, R10, R11, R12, R13, R14, R15, R16, R17, R18, R19, R20, R2I, and R22 optionally bond together to make a macrocycle, and where R* is independently at each occurrence selected from hydrogen, a C5-C20 aryl, and a C1-C20 alkyl.
[0001 17] In some embodiments, the compound has the structure of Formula (I).
[000118] In some embodiments, the compound has the structure of Formula (II).
[000119] In some embodiments, the compound has the structure of Formula (III).
[000120] In some embodiments, the compound has the structure of Formula (IV).
[000121] In some embodiments, the compound has the structure of Formula (V).
[000122] In some embodiments, the compound has the structure of Formula (VI).
[000123] In some embodiments, the compound has the structure of Formula (VII).
[000124] According to the present disclosure, Q group in the compounds of Formula (I),
Formula (II), Formula (III), Formula (IV), Formula (V), or Formula (VI) can be the same or different. One Q can be independently selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain and another Q can be independently selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain.
[000125] In some embodiments, Q groups are the same. In one embodiment, Q is -S- at each occurrence.
[000126] In some embodiments, Ri and R12 are each independently selected from alanine, [3- alanine, glycine, phenylalanine, tyrosine, lysine, aspartic acid, an acetyl group, and dimer combinations thereof, each of which is optionally substituted with fluorescein.
[000127] In some embodiments, R11 and R22 are each -NH2, -OH, -NHR*, -OR*, -SH, -SR*, or -NR*-N(R*)2, where R* is independently at each occurrence selected from hydrogen, a C5-C20 aryl, and a C1-C20 alkyl.
[000128] In some embodiments, R2 and R13 are each hydrogen.
[000129] In some embodiments, R3 and R14 are each a side chain of arginine.
[000130] In some embodiments, R4 and R15 are each independently selected from a side chain of lysine and a side chain of alanine.
[000131] In some embodiments, R5 and R16 are each a side chain of arginine.
[000132] In some embodiments, R6 and R17 are each independently selected from a side chain of histidine and a side chain of alanine.
[000133] In some embodiments, R7 and Ris are each independently selected from a side chain of lysine and a side chain of alanine.
[000134] In some embodiments, R8 and R19 are each a side chain of leucine.
[000135] In some embodiments, R9 and R20 are each a side chain of asparagine.
[000136] In some embodiments, Rio and R21 are each a side chain of arginine.
[000137] In some embodiments, each of R2 and R13 is hydrogen, each of R3 and R14 is a side chain of arginine, each of R4 and R15 is a side chain of lysine, each of R5 and R16 is a side chain of arginine, each of Re and R17 is a side chain of histidine, each of R7 and Ris is a side chain of lysine,
each of R8 and R19 is a side chain of leucine, each of R9 and R20 is a side chain of asparagine, each of Rio and R21 is a side chain of arginine, R11 and R22 are each NH2, and Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-P-alanine-glycine.
[000138] In some embodiments, each of R2 and R13 is hydrogen, each of R3 and R14 is a side chain of arginine, each of R4 and R15 is a side chain of alanine, each of R5 and R16 is a side chain of arginine, each of R6 and R17 is a side chain of histidine, each of R7 and Ris is a side chain of alanine, each of Rs and R19 is a side chain of leucine, each of R9 and R20 is a side chain of asparagine, each of Rio and R21 is a side chain of arginine, R11 and R22 are each NH2, and Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-P-alanine-glycine. [000139] In some embodiments, each of R2 and R13 is hydrogen, each of R3 and R14 is a side chain of alanine, each of R4 and R15 is a side chain of lysine, each of R5 and R16 is a side chain of arginine, each of R6 and R17 is a side chain of alanine, each of R7 and Ris is a side chain of lysine, each of R8 and R19 is a side chain of leucine, each of R9 and R20 is a side chain of asparagine, each of Rio and R21 is a side chain of arginine, R11 and R22 are each NH2, and Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-P-alanine-glycine.
[000140] According to the present disclosure, each amino acid side chain can be an L- or D- isomer. In some embodiments, each amino acid side chain is the L-isomer.
[000141] In some embodiments, a lysine residue and an aspartic acid residue are cyclized as a lactam macrocycle.
[000142] In some embodiments, the peptide is selected from Table 1.
Table 1. List of Exemplary Peptide Compounds According to the Present Invention, Sequence, Calculated Mass, Observed Mass, and Net Charge
Ac = acetyl; Flu = fluorescein; d> = melamine; X = 4-pentenoic acid; G* = A-allylglycine; Ap = beta-alanine; Nle = Norleucine; K* = lysine; and D* = aspartic acid. K* and D* are cyclized as a lactam macrocycle. X and G* are cyclized as an olefin macrocycle. The compounds were purified by HPLC (Figures 9A-9B).
[000143] In some embodiments, the compound has a structure as shown in Table 2 or a pharmaceutically acceptable salt thereof.
Table 2. Chemical Structure of Peptides Table discloses SEQ ID NOS 4-6, 10, 9, 12, 11, 14, 13, 15-16, 7-8, 20, 19, 17-18, and 21-22, respectively, in order of appearance.
[000144] In some embodiments, the compound has a structure
ID NOS 34 and 35, respectively, in order of appearance)), or a pharmaceutically acceptable salt thereof.
Therapeutic Formulations, Administration and Uses
[000145] The present disclosure provides pharmaceutical compositions comprising the compound of the present disclosure.
[000146] The compositions of the disclosure may be formulated with suitable carriers, excipients, and other agents that provide improved transfer, delivery, tolerance, and the like. A multitude of appropriate formulations may be found in the formulary known to all pharmaceutical chemists: Remington's Pharmaceutical Sciences, Mack Publishing Company, Easton, PA. These formulations include, for example, powders, pastes, ointments, jellies, waxes, oils, lipids, lipid (cationic or anionic) containing vesicles (such as LIPOFECTIN™, Life Technologies, Carlsbad, CA), DNA conjugates, anhydrous absorption pastes, oil-in-water and water-in-oil emulsions, emulsions carbowax (polyethylene glycols of various molecular weights), semi-solid gels, and semi-solid mixtures containing carbowax. See also Powell et al. "Compendium of excipients for parenteral formulations" PDA (1998) J Pharm Sci Technol 52:238-311.
[000147] The dose of a compound administered to a patient may vary depending upon the age and the size of the patient, target disease, conditions, route of administration, and the like. The suitable dose is typically calculated according to body weight or body surface area. When a compound of the present disclosure is used for therapeutic purposes in an adult patient, it may be advantageous to intravenously administer the compound of the present disclosure normally at a single dose of about 0.01 to about 20 mg/kg body weight, more preferably about 0.02 to about 7, about 0.03 to about 5, or about 0.05 to about 3 mg/kg body weight. Depending on the severity of the condition, the frequency and the duration of the treatment may be adjusted. Effective dosages
and schedules for administering a compound may be determined empirically; for example, patient progress may be monitored by periodic assessment, and the dose adjusted accordingly. Moreover, interspecies scaling of dosages may be performed using well-known methods in the art (e.g., Mordenti et al., 1991, Pharmacent. Res. 8:1351).
[000148] Various delivery systems are known and may be used to administer the pharmaceutical composition of the disclosure, e.g., encapsulation in liposomes, microparticles, microcapsules, recombinant cells capable of expressing the mutant viruses, receptor mediated endocytosis (see, e.g., Wu et al., 1987, J. Biol. Chem. 262:4429-4432). Methods of introduction include, but are not limited to, intradermal, intramuscular, intraperitoneal, intravenous, subcutaneous, intranasal, epidural, and oral routes. The composition may be administered by any convenient route, for example by infusion or bolus injection, by absorption through epithelial or mucocutaneous linings (e.g., oral mucosa, rectal and intestinal mucosa, etc.) and may be administered together with other biologically active agents. Administration may be systemic or local.
[000149] A pharmaceutical composition of the present disclosure may be delivered subcutaneously or intravenously with a standard needle and syringe. In addition, with respect to subcutaneous delivery, a pen delivery device can be used to deliver a pharmaceutical composition of the present disclosure. Such a pen delivery device may be reusable or disposable. A reusable pen delivery device generally utilizes a replaceable cartridge that contains a pharmaceutical composition. Once all of the pharmaceutical composition within the cartridge has been administered and the cartridge is empty, the empty cartridge may readily be discarded and replaced with a new cartridge that contains the pharmaceutical composition. The pen delivery device may then be reused. In a disposable pen delivery device, there is no replaceable cartridge. Rather, the disposable pen delivery device comes prefilled with the pharmaceutical composition held in a reservoir within the device. Once the reservoir is emptied of the pharmaceutical composition, the entire device is discarded.
[000150] In certain situations, the pharmaceutical composition may be delivered in a controlled release system. In one embodiment, a pump may be used (see Langer, supra; Sefton, 1987, CRC Crit. Ref. Biomed. Eng. 14:201). In another embodiment, polymeric materials may be used; see, Medical Applications of Controlled Release, Langer and Wise (eds.), 1974, CRC Pres., Boca Raton, Florida. In yet another embodiment, a controlled release system may be placed in
proximity of the composition’s target, thus requiring only a fraction of the systemic dose (see, e.g., Goodson, 1984, in Medical Applications of Controlled Release, supra, vol. 2, pp. 115-138). Other controlled release systems are discussed in the review by Langer, 1990, Science 249: 1527-1533. [000151] The injectable preparations may include dosage forms for intravenous, subcutaneous, intracutaneous and intramuscular injections, drip infusions, etc. These injectable preparations may be prepared by methods publicly known. For example, the injectable preparations may be prepared, e.g., by dissolving, suspending or emulsifying the compound or its salt described above in a sterile aqueous medium or an oily medium conventionally used for injections. As the aqueous medium for injections, there are, for example, physiological saline, an isotonic solution containing glucose and other auxiliary agents, etc., which may be used in combination with an appropriate solubilizing agent such as an alcohol (e.g., ethanol), a polyalcohol (e.g., propylene glycol, polyethylene glycol), a nonionic surfactant [e.g., polysorbate 80, HCO-50 (polyoxyethylene (50 mol) adduct of hydrogenated castor oil)], etc. As the oily medium, there are employed, e.g., sesame oil, soybean oil, etc., which may be used in combination with a solubilizing agent such as benzyl benzoate, benzyl alcohol, etc. The injection thus prepared is preferably filled in an appropriate ampoule.
[000152] Advantageously, the pharmaceutical compositions for oral or parenteral use described above are prepared into dosage forms in a unit dose suited to fit a dose of the active ingredients. Such dosage forms in a unit dose include, for example, tablets, pills, capsules, injections (ampoules), suppositories, etc.
Therapeutic Uses of the Compounds
[000153] In another aspect, the compounds disclosed herein are useful, inter alia, for the treatment, prevention and/or amelioration of a disease, disorder or condition in need of such treatment.
[000154] In one aspect, the present disclosure provides a method of treating a repeat expansion disorder in a subject in need of such treatment comprising administering to the subject a therapeutically effective amount of a compound according to the disclosure, or the composition comprising any compound according to the present disclosure.
[000155] In some embodiments, the compounds disclosed herein are useful for treating a noncoding repeat expansion disorder. In some embodiments, the disorder is selected from the group consisting of Baratela-Scott syndrome; CANVAS, cerebellar ataxia, neuropathy and
vestibular areflexia syndrome, myotonic dystrophy type 1 ; DM2, myotonic dystrophy type 2, progressive myoclonus epilepsy type 1 (Unverricht-Lundborg disease), familial adult myoclonic epilepsy, Fuchs endothelial corneal dystrophy type 3, fragile XE syndrome; FRDA, Friedreich ataxia, frontotemporal dementia/amyotrophic lateral sclerosis, Fragile X-associated premature ovarian infertility, fragile X syndrome, fragile X-associated tremor ataxia syndrome, global developmental delay, progressive ataxia, elevated glutamine, neuronal intranuclear inclusion disease, oculopharyngodistal myopathy type 1, oculopharyngodistal myopathy type 2, oculopharyngeal myopathy with leukoencephalopathy type 1, spinocerebellar ataxia, and X-linked dystonia parkinsonism.
[000156] In some embodiments, the compounds disclosed herein are useful for treating a coding repeat expansion disorder. In some embodiments, the disorder is selected from the group consisting of brachydactyly and cleidocranial dysplasia (BCCD), blepharophimosis, ptosis and epicanthus inversus (BPES), congenital central hypoventilation syndrome (CCHS), dentatorubral- pallidoluysian atrophy (DRPLA), early infantile epileptic encephalopathy type 1 (EIEE1), Huntington disease, Huntington disease-like 2, hand-foot-genital syndrome, holoprosencephaly type 5, mental retardation with isolated growth hormone deficiency (MRGH), oculopharyngeal muscular dystrophy, spinal and bulbar muscular atrophy, synpolydactyly type 1, and spinocerebellar ataxia.
[000157] In some embodiments, the repeat is a repeat RNA selected from CUG, CAG, CAA, GAA, CGG, CCG, GCG, GCA, GCC, GCU, AUUCU, AUUGU, CAGCUG, CUGCAG, or CCCCGCCCCGCG (SEQ ID NO: 25). In one embodiment, the repeat is CUG.
[000158] In some embodiments, the compounds disclosed herein are useful for treating a Huntington’s disease in a subject in need of such treatment. In one embodiment, the subject is human. In one embodiment, the subject is a mammal selected from the group consisting of a cat, a dog, a cow, a horse, a sheep, and a pig.
[000159] In one aspect, the present disclosure provides a method of binding double-stranded RNA comprising contacting the double-stranded RNA with a compound disclosed herein. In one embodiment, the double-stranded RNA is contacted with the compound in a cell. In one embodiment, the cell is a bacterial cell. In one embodiment, the cell is a mammalian cell. In one embodiment, the cell is a human cell. In one embodiment, the cell is cancer cell, a brain cell, a connective tissue cell, or a muscle cell.
EXAMPLES
[000160] The following examples illustrate specific aspects of the instant description. The examples should not be construed as limiting, as the examples merely provide specific understanding and practice of the embodiments and their various aspects.
General Information
[000161] Fmoc-protected amino acids, chemicals, and solvents were purchased from Alfa Aesar, Chem-Impex International, Sigma Aldrich, Acros Organics, VWR, Tokyo Chemical Industry, Ambeed, and Matrix Scientific. Standard desalted oligonucleotides were purchased from Sigma Aldrich and used without further purification unless otherwise noted. Cy5-r(CUG)io RNA was ordered from Sigma Aldrich as HPLC purified. ATP, [y-32P] was purchased from PerkinElmer. T4 Polynucleotide Kinase (lOU/pL) and RNase T1 were purchased from ThermoFisher. Abbreviations represents the following chemicals: lithium aluminum hydride (LAH), phosphorus tribromide (PBri), acetic anhydride (AC2O), diisopropylethylamine (DIPEA), /VpV A", A' -tetramethyl -O-(l //-benzotri azol-l-yl)uronium hexafluorophosphate (HBTU), 1- hydroxybenzotriazole (HOBt), N, N ’-diisopropylcarbodiimide (DIC), trifluoroacetic acid (TFA), triisopropylsilane (TIPS), 1,2-ethanedithiol (EDT), dichloromethane (DCM), N,N- dimethylformamide (DMF), methanol (MeOH), tetrahydrofuran (THF), water (H2O), ethyl acetate (EtOAc), magnesium sulfate (MgSCU), diethyl ether (Et2O), and acetonitrile (ACN). Synthesized peptides were purified on Thermofisher Scientific Ultimate 3000 HPLC by reverse-phase high- performance liquid chromatography (RP-HPLC) using preparative Cis column (Phenomenex Luna® 5pM Cis, 100 A, 100 x 21.2 mm). Purity was assessed on Agilent 1260 Infinity series RP- HPLC with an analytical Cis column (Xterra® RP18 3.5 pM, 2.1 x 150 mm) using a gradient of 5-95% acetonitrile/water with 0.1% TFA. Traces were monitored at 220 nm. Exact mass for peptides were determined using Bruker UltrafleXtreme MALDI-TOF (matrix-assisted laser desorption/ionization time-of-flight). All buffers were prepared with DEPC-treated water.
Example 1; Synthesis of Peptide- 1 and CHF Monomeric Strands
[000162] Peptides were synthesized on Rink Amide MB HA resin at a 0.25 pinole scale using standard Fmoc-solid phase peptide chemistry according to Scheme 1. In brief, resin was swelled in DCM for 10 minutes and drain using vacuum filtration. Afterwards, resin underwent iterative cycles of Fmoc-deprotection and Fmoc-protected amino acid coupling, with washes in between. All reactions were conducted at room temperature. For Fmoc-deprotection, 4 mL of 20% piperidine in DMF was added to the resin and placed on a nutator for 30 minutes. Resin was subsequently drained using vacuum filtration and washed with DMF (3x), DCM (3x), and DMF (3x). For coupling, a 4 mL solution of Fmoc-amino acid (5 eq ), HBTU (4.5 eq.), and DIPEA (10 eq.) in DMF was added to the resin and placed on a nutator for 1 hour. Resin was drained and washed with DMF (3x), MeOH (3x), and DCM (3x). After completing desired peptide sequences, the N-terminus was capped with AC2O (10 eq.) in 4 mL DMF for 30 minutes. Peptides was cleaved from resin by stirring in 4 mL of TFA/thioanisole/anisole/l,2-ethanedithiol (90:5:2:3 v/v) for 3 hours at room temperature. Cleavage solution was then filtered through a cotton plug syringe and concentrated via rotary evaporation. The concentrated peptide was precipitated with 10 mL of cold diethyl ether, pelleted by centrifugation, and decanted for a total of 3 times. The pellet was dried by a slow stream of Nife) and dissolved in water and minimal acetonitrile. Crude peptide solution was passed through a 0.45 pm nylon syringe filter and purified on RP-HPLC using a gradient of 10-50% acetonitrile/water with 0.1% TFA over the course of 30 minutes. All purified peptides were characterized on analytical HPLC and MALDI-TOF mass spectrometer and freeze-dried on lyophilization unit.
Example 2; Synthesis of HBS-1
[000163] HBS synthesis was followed according to Scheme 2, as described in Miller et al., “Synthesis of Hydrogen-Bond Surrogate Alpha-Helices as Inhibitors of Protein-Protein Interactions,” Curr. Protoc. Chem. Biol. 6(2): 101-116 (2014), which is hereby incorporated by reference in its entirety. Following amino acid coupling up to the z + 3rd position using standard Fmoc-solid phase peptide synthesis, the resin underwent Fmoc-deprotection with 20% piperidine in DMF and subsequently washed with DMF (3x), DCM (3x), and DMF (3x). In a separate 20 mb scintillation vial, a 4 mb solution containing bromoacetic acid (10 eq.), DIC (20 eq.), and HOBt (10 eq.) in DMF was prepared and stirred for 15 minutes. The solution was transferred to the resin and placed on a nutator for 2 hours. The resin was drained and washed with DMF (3x), MeOH (3x), and DCM (3x). Resin was resuspended in 4 mb of DMF, added with allylamine (20 eq ), and placed on a nutator for 20 minutes. The reaction was drained and washed with DMF (3x), MeOH (3x), and DCM (3x). The peptide was further elongated by two amino acids followed with an Fmoc-deprotection. A separate 20 mb solution of 4-pentenoic acid (5 eq.), DIC (5 eq.), and HOBt (5 eq.) in DMF that was previously stirred for 15 minutes was transferred to the resin and shaken on a nutator for 2 hours. The resin was washed with DMF (3x), MeOH (3x), DCM (3x), and Et2O (3x) and dried overnight in a vacuum desiccator. The resin was purged with nitrogen, added with 20 mol% Hovey da-Grubbs 2nd Generation catalyst and 4 mb of anhydrous 1,2-di chloroethane. Ring-closing metathesis was performed by heating the mixture to 120 °C for 15 minutes at 150 W. The reaction was drained and washed with DMF (3x), MeOH (3x), and DCM (3x). 4-Methyltrityl
group was then selectively removed by 3 mL additions of TFA/TIPS/DCM (1 :5:94) for 10 minutes at room temperature for a total of 8 rounds. The resin was then washed with 15% DIPEA in DMF (2x), DMF (3x), and DCM (3x). The peptide was fluorescently labeled by addition of a 4 mL solution of fluorescein isothiocyanate (3 eq.) and DIPEA (5 eq.) in DMF. The reaction vessel was wrapped in aluminum foil and placed on a nutator overnight in the dark. Resin was drained and washed with DMF (3x), MeOH (3x), DCM (3x), and Et2O. HBS-1 was cleaved from resin and purified under the same conditions from “Synthesis of Peptide- 1 and CHF Monomeric Strands.”
[000164] Synthesis of Peptide-Mel was followed according to Scheme 3, from previous report (Mao et al., “Synthesis of DNA-Binding Peptoids,” Synlett 26(11): 1581-1585 (2015), which is hereby incorporated by reference in its entirety), with modifications. Following standard Fmoc-solid phase peptide synthesis, the N-terminally deprotected resin was transferred to a 10 mL round bottom flask equipped with a stir bar. 2-Chloro-4,6-diamino-l,3,5-triazine (100 eq.) and 5 mL of 0.1M CS2CO3, pH 10, in (1 : 1) DMF/H2O was added. The flask was covered with aluminum foil and the suspension was left to stir at 85 °C for 24 hours. The reaction was then cooled to room temperature, transferred to a clean 50 mL round bottom flask, and diluted with 25 mL of MeOH. While the resin settled to the bottom, the supernatant that contained starting material was transferred to waste. The flask was added with another 25 mL of MeOH and the supernatant was removed. A total of 3 rounds was performed. Following the workup, the resin was cleaved and purified under the same conditions from “Synthesis of Peptide- 1 and CHF Monomeric Strands.” Peptide-Mel was characterized on analytical HPLC and MALDI-TOF mass spectrometer and freeze-dried on lyophilization unit.
Example 4; Synthesis of Fluorescein-Labeled Peptides
[000165] For N-terminal fluorescein labeling, Scheme 4, standard Fmoc solid phase peptide synthesis was followed as previously mentioned. A 4 mL solution containing Fmoc-P-alanine (5 eq.), 4.5 eq. HBTU, and 10 eq. DIPEA in DMF was added to the Fmoc-deprotected resin and placed on a nutator for 1 hour. Reaction was drained and washed with DMF (3x), MeOH (3x), and DCM (3x). Following another Fmoc-deprotection with subsequent washings, a 4 mL solution of fluorescein isothiocyanate (3 eq.) and DIPEA (5 eq.) in DMF was added to the resin. The reaction vessel was wrapped in aluminum foil and placed on a nutator overnight in the dark. Resin was drained and washed with DMF (3x), MeOH (3x), DCM (3x), and Et2O (3x). Fluorescein-labeled peptide was cleaved and purified under the same conditions from “Synthesis of Peptide- 1 and CHF Monomeric Strands.”
[000166] Cysteine-containing peptides (0.5 mM, 2 eq.) and dibromo-m-xylene or cis-stilbene dibromide (0.25 mM, 1 eq.) were dissolved in a solution of (1 :3) acetonitrile/ 20 mM ammonium bicarbonate, pH 8.0. The reaction proceeded for 1 hour at room temperature and was immediately purified on RP-HPLC using a gradient of 10-65% acetonitrile/water with 0.1% TFA over the course of 30 minutes. TAV2b-al and CHF compounds were characterized on analytical HPLC and MALDI-TOF mass spectrometer and freeze-dried by lyophilization.
[000167] An oven-dried 50 mb round-bottomed flask, equipped with a magnetic stir bar and an addition funnel, was charged with N2(g) and placed in an ice bath. 10 mb of anhydrous THF and 1.01 mb of LAH (2 M in THF, 2.02 mmol) were sequentially added. The solution was stirred for
10 minutes to equilibrate. Dimethyl cis-stilbene-4,4’-dicarboxylate 1 (200 mg, 0.675 mmol) was separately dissolved in 7.3 mL of anhydrous THF and transferred into the addition funnel. The solution was dispensed into the round-bottom flask over the course of 15 minutes and left to stir for 6 hours at room temperature. 1 mL of deionized H2O was added to quench the reaction. The precipitated solution was filtered using a Buchner funnel vacuum filtration. After concentrating the filtrate on a rotary evaporator, the product was redissolved in 10 mL of DCM and washed with H2O (2 x 10 mL). The organic layer was dried over Na2SO4 and concentrated under vacuum. The cA-stilbene diol la was isolated as a white solid (140 mg, 86% yield).
[000168] XH NMR (500 MHz, DMSO-cL) 6 7.19 (s, 8H), 6.58 (s, 2H), 5.15 (t, J = 5 Hz, 2H), 4.45 (d, J = 5 Hz, 4H). 13C NMR (125 MHz, DMSO-t/<) S 141.6, 135.3, 129.6, 128.2, 126.4, 62.7.
[000169] CA-stilbene diol la (100 mg, 0.416 mmol) was added to a 10 mL round-bottom flask, charged with N2(g) and equipped with a stir bar. 4.16 mL of anhydrous THF was added and the solution was left to stir on ice for 10 minutes. PBn (86 pL, 0.916 mmol) was added in dropwise addition and the reaction was stirred for 2 hours at room temperature. A second addition of PBr? (86 pL, 0.916 mmol) was added in the same manner and stirred for a following 2 hours. The reaction was cooled in an ice bath before quenching with 2 mL of saturated sodium bicarbonate. The solution was then concentrated on rotary evaporator and redissolved with 10 mL of EtOAc. The organic layer was washed with saturated sodium bicarbonate (2 x 5 mL). The separated organic layer was dried over MgSCL, concentrated, and immediately purified by flash chromatography (5% EtOAc/Hexanes) to afford cA-stilbene dibromide 2 as a white solid (140 mg, 91%).
[000170] JH NMR (400 MHz, CDC13) S 7.27-7.21 (m, 8H), 6.58 (s, 2H), 4.47 (s, 4H). 13C NMR (125 MHz, CDCI3) <5 137.4, 136.8, 130.3, 129.4, 129.2, 33.5.
[000171] Chemical shifts are in agreement with previous report (Xu et al., “Control of the Intramolecular [2+2] Photocycloaddition in a Bis-Stilbene Macrocycle,” J. Org. Chem. 74(13):4874-7 (2009), which is hereby incorporated by reference in its entirety).
Example 7; Fluorescence Polarization Binding Assay
[000172] Relative binding affinities of all fluorescein-labeled peptides for duplex oligonucleotides were determined by direct fluorescence polarization binding assay (Figures 13A- 13D). Experiments were performed in a 96-well plate format and analyzed on a DTX 880 Multimode Detector (Beckman) plate reader at 25 °C with excitation and emission wavelengths set to 485 nm and 525 nm, respectively. All oligonucleotides were folded by heating at 95 °C for
1 minute and slowly cooled to room temperature for 1 hour. Each well contained the addition of 2-fold serially diluted oligonucleotides to 15 nM solution of fluorescein-labeled peptide in 50 mM Tris, 50 mM KC1, 1 mM MgCE, 0.01% Triton-X-100, pH 7.4. Sample plate was covered with aluminum foil and incubated at room temperature for at 1 hour before taking a measurement. Binding affinity values, from experiments conducted in triplicate, were determined by fitting to a sigmoidal dose-response nonlinear regression model on GraphPad Prism 9.
where,
RT = Total concentration of oligonucleotide
LST= Total concentration of fluorescein-labeled peptide FSB = Fraction of bound fluorescein-labeled peptide
Example 8: Microscale Thermophoresis Binding Assay
[000173] Relative binding affinities of CHF-2 for Cy5-r(CUG)io were determined by microscale thermophoresis (MST) assay. Bindings measurements were conducted in premium coated capillary tubes (NanoTemper) on a Monolith NT.115 Pico (NanoTemper). Cy5-r(CUG)io was refolded by heating at 95 °C for 1 minute and slowly cooled to room temperature for 1 hour. Bindings were performed in a 96-well plate containing 5 nM Cy5-r(CUG)io and 2-fold serially diluted CHF-2, beginning at 200 pM, in 50 mM Tris, 50 mM KC1, 1 mM MgCh, 0.05% Tween- 20, pH 7.4. Sample plate was covered with aluminum foil and was incubated at room temperature for 1 hour before loading into premium coated capillary tubes. The instrument used the following parameters: excitation power at 10%, MST power at medium, phase 1 duration at 3 seconds, phase
2 duration at 30 seconds, and phase 3 duration at 1 second. Bindings were performed in triplicate. Dissociation constant was determined by fitting the data to a sigmoidal dose-response nonlinear regression model on GraphPad Prism 9.
where,
RT = Total concentration of oligonucleotide
LST = Total concentration of fluorescein-labeled peptide FSB = Fraction of bound fluorescein-labeled peptide
Example 9: Circular Dichroism
[000174] Circular dichroism (CD) measurements were taken on a Jasco J-1500 Circular Dichroism spectrometer equipped with a temperature controller using 1 mm length cells and a scan speed of 5 nm/min at 298 K. The spectra were generated from an average of 10 scans with baseline subtractions. Samples were prepared in 10 mM Tris and 10 mM KF, pH 7.4 and were incubated for 30 minutes at room temperature before taking a reading.
Example 10: 5’-32P-Radiolabeling of RNA
[000175] To an autoclaved microcentrifuge tube was added with 1 pL of 10 pM RNA, 1 pL of [y-32P] ATP (150 mCi/mL, 6000 Ci/mmole; PerkinElmer), 2 pL of 10X Reaction Buffer (50 mM Tris, 10 mM MgCl2, 5 mM DTT, 0.1 mM spermidine, pH 7.6; ThermoFisher), and 15 pL of DEPC-treated water. The solution was mixed and 1 pL of T4 Polynucleotide Kinase (lOU/pL; ThermoFisher) was added last. The reaction mixture was incubated for 1 hour at 37 °C. The reaction was then cooled to room temperature and quenched with 1 volume of (25:24:1) phenol/chloroform/isoamyl alcohol (Acros Organics). The solution was vortexed for 30 seconds and centrifuge at 14,000 rpm for 30 minutes. The aqueous layer was isolated and transferred to a clean microcentrifuge tube. 0.1 volume of 3 M sodium acetate, pH 5.2, was added and mixed. 3 volumes of 200 proof cold ethanol were added and the RNA was precipitated overnight in -80 °C. The precipitant was centrifuged at 14,000 rpm, 4 °C for 30 minutes, and the supernatant was removed after. The pellet was washed with 3 volumes of 70% cold ethanol in water, centrifuge at 14,000 rpm, 4 °C for 30 minutes, and had the supernatant removed for a total of 3 times. The pellet was dried on a speedvac vacuum concentrator, resuspended in 10 pL of IX RNA loading dye, and purified on a denatured 20% polyacrylamide gel containing 8 M urea. Purified 32P-labeled RNA was refolded by heating to 95 °C for 1 minute and slowly cooled to room temperature for 1 hour.
Example 11: Electrophoretic Mobility Shift Assay
[000176] Bindings were conducted in autoclave microcentrifuge tubes. In a 10 pL binding volume, 1 pL of freshly 32P-labeled ssRNA 1 (13,500 cpm) was mixed with 2-fold serially diluted CHF-2 (starting at 100 pM) in 50 mM Tris, 50 mM KC1, 1 mM MgCh, pH 7.4. Samples were incubated at room temperature for 2 hours and was added with 10 pL of 50% glycerol. 10 pL of sample was loaded onto a 10% native polyacrylamide gel and ran at 120 V for 15 minutes. The gel was dried, autoradiographed, and imaged on Amersham Typhoon gel imager.
Example 12: Hydroxyl Radical Footprinting Assay
[000177] Hydroxyl radical cleavages were carried out in autoclaved microcentrifuge tubes. Each tube contained 2 pL of freshly 32P-labeled RNA 1 (12,400 cpm), 2-fold serially diluted CHF- 2 (starting at 1 mM final concentration), 2 pL of buffer (10 mM Tris, 10 mM KC1, 1 mM MgCh, pH 7.4), and diluted up to 6 pL of DEPC-treated water. The samples were incubated at room temperature for 1 hour. To the walls of each tube were separately added with 1 pL of 2 mM (NH4)2Fe( SO4)2/ 4 mM EDTA, pH 8.0, 1 pL of 10 mM sodium ascorbate, and 1 pL of 0.6% H2O2. The hydroxyl radical cleavage was initiated by combining all three 1 pL spots and mixing into the bulk solution. The reaction was incubated at room temperature for 10 minutes and quenched with 1 pL of 0.1 M thiourea. 1 volume of 2X RNA loading dye was added and 10 pL of sample was loaded onto a denatured 20% polyacrylamide gel containing 8 M urea.
[000178] To generate a hydrolysis ladder: 0.25 pL of 32P-lab eled RNA 1 was added with 4.75 pL of hydrolysis buffer (50 mM NaHCCh, 1 mM EDTA, pH 9.2; ThermoFisher) in an autoclaved microcentrifuge tube. The RNA was heated to 90 °C for 10 minutes and cooled on ice. 1 volume of 2X RNA loading dye was added and 9 pL of sample was loaded onto the denatured polyacrylamide gel.
[000179] To generate an RNase T1 “G” Ladder: 0.5 pL of 32P-labeled RNA 1 was added with 8.5 pL IX Sequencing Buffer (20 mM sodium citrate, 1 mM EDTA, 7M urea, pH 5.0; ThermoFisher) in an autoclaved microcentrifuge tube. The RNA was heated to 90 °C for 30 seconds and cool to room temperature for 1 hour. 1 pL of RNase T1 (lOOU/pL; ThermoFisher) was added and left to incubate at room temperature for 10 minutes. 1 volume of 2X RNA loading
dye was added and 10 pL of sample was loaded onto a denatured 20% polyacrylamide gel containing 8 M urea.
[000180] Gel -loaded samples were run at 2000 V, 50 W for 3 hours on an Apogee Model S2
Sequencing Gel Electrophoresis System. The gel was dried, autoradiographed, and imaged on Amersham Typhoon gel imager.
Example 13: Results and Discussion of Examples 1-12
Minimal Motif That Mimics an RNA Binding Viral Protein
[000181] TAV2b consists of two segmented a-helical regions, defined as al and a2, that homodimerize in the major groove of the siRNA duplex (Figure 2A). The binding motif observed in the two al helices closely resembles that of the bZIP transcription factor. Therefore, a crosslinked helical fork (CHF) scaffold that can mimic al’s affinity for its cognate RNA sequence was designed. The minimal peptide sequence RKRHKLNR (SEQ ID NO: 3) was grafted from al onto the scaffold. While the native peptide-RNA complex features several electrostatic contacts with the negatively charged RNA backbone, it contains two hydrogen bonding interactions on the Hoogsteen face of two guanine bases in the duplex (5’-GC-3’) sequence (Figure 2B). The base specific hydrogen-bonding in the major groove is critical for specificity in protein-nucleic acids recognition (Rohs et al., “Origins of Specificity in Protein-DNA Recognition,” Annual Review of Biochemistry 79:233-269 (2010), which is hereby incorporated by reference in its entirety). The amino acid sequence in the CHF design was investigated to analyze the potential of this scaffold to recognize duplex RNA. The TAV2b/siRNA interaction has also inspired other efforts to develop constrained peptide mimics as RNA ligands. Kuepper et al., “Constrained Peptides Mimic a Viral Suppressor of RNA Silencing,” Nucleic Acids Res. 49: 12622-12633 (2021), which is hereby incorporated by reference in its entirety, has recently showed that a constrained sequence consisting of the al and a2 helical regions can recognize the cognate siRNA with high affinity.
[000182] Prior efforts to develop mimics of bZIP proteins as DNA ligands demonstrated that crosslinked peptide dimers lacking the leucine zipper dimerization domain can retain affinity and specificity for DNA comparable to the native bZIP proteins (Talanian et al., “Sequence-Specific DNA-Binding by a Short Peptide Dimer,” Science 249:769-771 (1990); Cuenoud et al., “Design of a Metallo Bzip-Protein That Discriminates Between Cre and Apl Target Sites - Selection Against Apl,” Proc. Natl. Acad. Sci. USA 90: 1154-1159 (1993); Cuenoud et al., “Altered
Specificity of DNA- Binding Proteins with Transition-Metal Dimerization Domains,” Science 259:510-513 (1993); Morii et al., “Sequence-Specific DNA-Binding by a Geometrically Constrained Peptide Dimer,” J. Am. Chem. Soc. 115: 1150-1151 (1993); Caamano et al., “A Light- Modulated Sequence-Specific DNA-Binding Peptide,” Angew. Chem. Int. Ed. 39:3104-3107 (2000), which are hereby incorporated by reference in their entirety). Encouraged by these results, the TAV2b dimer was modeled and it was hypothesized that a cis-stilbene linker would span the 19 A distance between the minimal homodimeric regions that occupy the adjacent RNA major grooves. Cis-stilbene provides an optimally-spaced turn segment to link two parallel helices (Erdelyi et al., “A New Tool in Peptide Engineering: A Photoswitchable Stilbene-Type BetaHairpin Mimetic,” Chemistry-a European Journal 12:403-412 (2006), which is hereby incorporated by reference in its entirety).
[000183] Fluorescein-linked CHF-1 (CHF-lFlu) was synthesized and its binding to a hairpin RNA sequence derived from TAV2b’s cognate siRNA duplex (RNA 1) was analyzed. A fluorescence polarization (FP) assay was developed to probe the peptide-RNA binding affinities. The 8mer CHF-lFlu dimer bound RNA 1 with a dissociation constant of 212 ± 15 nM; in comparison, the longer 21mer native dimer, TAV2b-alFlu, bound RNA 1 with KD = 17 ± 1 nM (Figures 3A-3B, Table 3). The higher binding affinity of TAV2b-alFlu to RNA is not surprising as the longer peptide contains an additional 8 cationic residues not present in CHF-lFlu, which likely form nonspecific electrostatic interactions with the oligonucleotide phosphate groups.
Table 3. List of Oligonucleotides with Associated Sequence Length and Molecular Weight (MW) Table discloses SEQ ID NOS 26-33, respectively, in order of appearance.
Cy5 = Cyanine-5 labeled on 5 ’-end of r(CUG)io RNA
[000184] It was hypothesized that if CHF-lFlu contacts the Hoogsteen face in the 5’-GC-3’ stretch, this peptide may bind to a CUG repeat sequence featuring multiple 5’-GC-3’ regions with higher affinity. Triplet repeat RNA and DNA sequences have been implicated in several human genetic diseases; the CUG repeat RNA is known to be critical for the pathogenesis of myotonic dystrophy type 1 (Depienne et al., “30 years of Repeat Expansion Disorders: What have we Learned and What are the Remaining Challenges?,” American Journal of Human Genetics 108:764-785 (2021); Mahadevan et al., “Myotonic-Dystrophy Mutation - an Unstable Ctg Repeat in the 3' Untranslated Region of the Gene,” Science 255:1253-1255 (1992); Miller et al., “Recruitment of Human Muscleblind Proteins to (CUG)(n) Expansions Associated with Myotonic Dystrophy,” EMBO J. 19:4439- 4448 (2000), which are hereby incorporated by reference in their entirety). The CUG repeat RNA, r(CUG)io, and RNA 1 differ by one nucleotide in spacing between the two sets of diagonal G’s recognized by the peptide dimer - the 5’-GC-3’ stretch is spaced by 5 bases in RNA 1 while GC stretches are spaced by four bases in r(CUG)io. CHF-lFlu contains a flexible region attached to cis-stilbene, which was predicted to allow recognition of a designed hairpin featuring CUG repeats. As expected from the binding model, CHF-lFlu bound r(CUG)io with improved affinity (Kd = 74 ± 17 nM) compared to RNA 1 (Figures 4A-4B). Conversely, CHF-1F1U bound the DNA analog d(CTG)io with 6- fold lower binding affinity (Kd = 436 ± 79 nM). In contrast to the helix dimer, the parent sequence TAV2b-alFlu recognized d(CTG)io with 6-fold higher affinity than r(CUG)io (Figure 10).
The Dimeric Helix is Required for High Affinity Duplex RNA Recognition
[000185] To ascertain that the dimeric nature of CHF-lFlu is essential to its RNA recognition properties, a monomeric unconstrained peptide (Peptide-1) was prepared. Peptide-1 showed moderate affinity towards r(CUG)io, (Kd = 2.2 ± 0.9 pM), and no appreciable binding for DNA 1. While preference was observed for r(CUG)io, Peptide- 1 also bound to d(CTG)io (Figures 4A-4B) It was hypothesized that the conformational flexibility of the unconstrained peptide may be contributing to its lower affinity (Harrison et al., “Helical Cyclic Pentapeptides Constrain HIV- 1 Rev Peptide for Enhanced RNA Binding,” Tetrahedron 70:7645-7650 (2014), which is hereby incorporated by reference in its entirety). A conformational constraint - a hydrogen bond surrogate - was incorporated in the backbone of the peptide to promote a-helicity (Wang et al., “Evaluation of Biologically Relevant Short Alpha-Helices Stabilized by a Main-Chain Hydrogen-Bond Surrogate,” J. Am. Chem. Soc. 128:9248-9256 (2006), which is hereby incorporated by reference
in its entirety). However, HBS-1 did not lead to an enhancement in hairpin oligonucleotide recognition. Melamine group was conjugated to the peptide to obtain Peptide-Mel. Peptide-Mel was tested to determine whether addition of melamine group would provide enhanced affinity. Melamine has been proposed to engage the U-U noncanonical pair in CUG repeats (Arambula et al., “A Simple Ligand that Selectively Targets CUG Trinucleotide Repeats and Inhibits MBNL Protein Binding,” Proc. Natl. Acad. Sci. 106: 16068-16073 (2009), which is hereby incorporated by reference in its entirety), and improve affinity and recognition for RNA substrates (Wong et al., “Targeting Toxic RNAs that Cause Myotonic Dystrophy Type 1 (DM1) with a Bisamidinium Inhibitor,” J. Am. Chem. Soc. 136:6355-6361 (2014), which is hereby incorporated by reference in its entirety). Unfortunately, the melamine modification also did not result in substantial improvement in hairpin oligonucleotide affinity (Figures 4A-4B). Overall, these results suggested that the dimeric motif has a distinct advantage over the monomeric constructs for duplex RNA targeting.
The RNA Binding Specificity of Crosslinked Helix Forks can be Tuned
[000186] CHF-lFlu bound to RNA 1 and r(CUG)io with nanomolar affinity providing robust ligands for these hairpin RNA oligonucleotides. The overall goal was to develop a scaffold that allows sequence-specific targeting of duplex RNA. This goal required that the designed CHF could detect single base mismatches. The potential of CHF-lFlu to discriminate between two closely related RNA sequences that differ by one contact nucleotide was determined and the potential of this ligand to differentiate between r(CUG)io and r(CUG)Mut was tested. r(CUG)wut is a single mismatch control in which one target guanine nucleotide is changed to an adenine. The sequence of r(CUG)Mut is shown in Figure 8 and contains CUG repeats followed by CUA repeats. Unfortunately, CHF-lFlu bound r(CUG)Mut with roughly the same affinity as r(CUG)io (Figures 5A-5B). Based on this result, it was conjectured that specificity gained by binding of guanine Hoogsteen faces is being countered through non-specific interactions of solvent exposed cationic side chains. The TAV2b al helix contains two lysine residues, K27 and K30, which project away from RNA major groove and into the solvent but could engage in non-specific interactions distorting the binding profile of a minimal scaffold (Figure 5A). To test the impact of these cationic side chains, CHF-2Flu was synthesized in which both lysine residues were replaced with alanine. CHF-2F,U proved to be >50-fold more selective for r(CUG)io than to the single base mismatch sequence r(CUG)Mut. Surprisingly, CHF-2Flu also showed decreased binding for RNA-
1, DNA-1, and d(CTG)io hairpins suggesting potential non-specific interactions between CHF- lFlu and these oligonucleotides (Figures 5A-B). A fluorescence polarization assay with fluorescein-labeled peptides was utilized for binding affinity analysis. To rule out any dye-specific results, the binding interaction between CHF-2 and r(CUG)io was confirmed using microscale thermophoresis (MST) in which r(CUG)io was 5’-labeled with Cy5. In the MST assay, the unlabeled CHF-2 bound Cy5-r(CUG)io with Ka = 3.4 ± 0.6 pM, in agreement with the binding constant obtained from the FP assay (Figure 11).
[000187] The premise of the designed dimeric proteomimetics is that R26 and H29 engage the Hoogsteen face of two base paired guanine nucleotides as part of r(CUG)io recognition. A double alanine mutant, CHF-3Flu, was designed where R26 and H29 of CHF-lFlu were replaced with alanine (Figures 5A-5B). This negative control showed loss of binding affinity to r(CUG)io and other hairpin oligonucleotides tested demonstrating the importance of contacting the Hoogsteen faces of G.
RNA Footprinting Analysis Revealed Site Specific Binding on Duplex RNA
[000188] The CHF scaffold was rationally designed to bind double-stranded RNA. To investigate the CHF binding site and to ascertain that the scaffold preferentially binds the doublestranded region, hydroxyl radical footprinting with 5’-32P -radiolabeled RNA 1 and CHF-2 was performed. The cleavage pattern of RNA 1 on a denaturing electrophoresis gel was resolved (Figure 6). The footprinting pattern indicated that regions (i), (ii), and (iv) of the hairpin RNA were not cleaved in the presence of CHF-2, as expected from binding of CHF-2 to the stem region of the hairpin RNA (Figure 6). Consistent with the design, the footprinting result revealed that CHF-2 did not bind at the apical tetraloop, region (iii). The potential of CHF-2 to bind a singlestranded analog of RNA 1 was also evaluated using an electrophoretic mobility shift assay (Figures 8 and 12). This analysis confirmed that CHF-2 does not recognize a single-stranded RNA.
Circular Dichroism Spectroscopy Indicated Helical Conformation for Crosslinked Dimers in Complex With Hairpin RNA and B-Form Like Conformation for r(CUG)io
[000189] The conformation of CHF-2 in complex with r(CUG)io was probed via circular dichroism (CD). In the binding model, the CHF dimer adopted an a-helical structure in complex with duplex RNA. CHF-2 was weakly helical in aqueous solution - as would be expected of two
short unconstrained peptides (Kallenbach et al., In Circular Dichroism and the Conformational Analysis of Biomolecules,' Fasman, G. D., Ed.; Plenum Press: New York, p 201-259 (1996), which is hereby incorporated by reference in its entirety). Titration of r(CUG)io into a solution of CHF- 2 led to an enhancement in the 222 nm minima - which indicated an increase in the a-helical content of the peptides (Figure 7A). CD spectra of the RNA-peptide complex and the peptide alone (after subtraction of the r(CUG)io signal) are shown in Figure 7A.
[000190] The circular dichroism spectra were also analyzed to interrogate changes in RNA structure upon binding of CHF-2. Prior studies investigating RNA binding by major groovebinding peptides have noted that the RNA adopts a more B-form like conformation upon complexation (Daly et al., “Circular Dichroism Studies of the HIV-1 Rev Protein and its Specific RNA Binding Site,” Biochemistry 29:9791-5 (1990), which is hereby incorporated by reference in its entirety). A similar shift was observed in the circular dichroism spectrum of r(CUG)io upon complexation with CHF-2. Addition of CHF-2 to r(CUG)w resulted in a decrease in the positive 270 nm signal and an increase in the negative 208 nm band (Figures 7A-C). Significantly, the CD data suggested that r(CUG)io adopts a similar conformation to d(CTG)io when in complex with CHF-2. The RNA conformation in the high-resolution structure of TAV2b-siRNA complex was compared with canonical B-form DNA. This analysis showed that while the diameter and helical pitch of the nucleic acid helix differed, the major groove of the bound siRNA aligned well with the B-form DNA (Figure 7C).
[000191] A dsRNA-binding scaffold, dubbed crosslinked helical fork (CHF), that mimics a protein tertiary structure often engaged in double-stranded DNA recognition was rationally designed. In this initial effort, it was shown that the specificity and affinity of the scaffold for dsRNA can be tuned through judicious modifications of the amino acid sequence, net charge, and crosslinker. Given the growing importance of RNA as a therapeutic target, broad potential for the CHF scaffold was envision capable of recognizing hairpin regions in RNA tertiary structure. For this potential to be met, the current CHF design needed to be improved. The model dictates helical motifs binding in the RNA major groove. Synthetic strategies that constrain the peptide helical conformation (Pelay-Gimeno et al., “Structure-Based Design of Inhibitors of Protein-Protein Interactions: Mimicking Peptide Binding Epitopes,” Angew. Chem. Int. Ed. 54:8896-8927 (2015), which is hereby incorporated by reference in its entirety) or replacement of the peptide with small molecule topographical helix mimics (Lao et al ., “Rational Design of Topographical Helix Mimics
as Potent Inhibitors of Protein- Protein Interactions,” J. Am. Chem. Soc. 136:7877-88 (2014), which is hereby incorporated by reference in its entirety) may lead to enhanced affinity, specificity, and cellular activity. Methods for stabilizing or mimicking the a-helices have been previously described and are also being currently studied (Sawyer et al., “Protein Domain Mimics as Modulators of Protein-Protein Interactions,” Acc. Chem. Res. 50: 1313-1322 (2017), which is hereby incorporated by reference in its entirety). It is anticipated that replacement of canonical amino acid side chains with nonnatural functional groups or nucleotides to form triplex motifs would lead to an improved recognition of the RNA Hoogsteen surface (Huang et al., “The a- Helical Peptide Nucleic Acid Concept: Merger of Peptide Secondary Structure and Codified Nucleic Acid Recognition,” J. Am. Chem. Soc. 126:4626-4640 (2004), which is hereby incorporated by reference in its entirety). The CHF scaffold provides a platform for these and other modifications to develop a new class of sequence-encodable ligands for duplex RNA.
[000192] As various changes can be made in the above-described subject matter without departing from the scope and spirit of the present invention, it is intended that all subject matter contained in the above description, or defined in the appended claims, be interpreted as descriptive and illustrative of the present invention. Many modifications and variations of the present invention are possible in light of the above teachings. Accordingly, the present description is intended to embrace all such alternatives, modifications, and variances which fall within the scope of the appended claims.
[000193] All patents, applications, publications, test methods, literature, and other materials cited herein are hereby incorporated by reference in their entirety as if physically present in this specification.
Claims
1. An encodable scaffold that engages the major groove of double-stranded RNA.
2. The scaffold of claim 1, wherein the scaffold comprises a first domain DI and a second domain D2, wherein the first domain and the second domain are connected by a covalent linker L, and wherein the scaffold spatially fits in the major groove of double-stranded RNA.
3. The scaffold of claim 2, wherein the scaffold binds to the major groove of the doublestranded RNA using base-specific bonding selected from the group consisting of Hoogsteen hydrogen bonding and electrostatic contact.
4. The scaffold of any one of claims 1 -3, wherein the DI and the D2 are each helical oligomers or helix mimics.
5. The scaffold of any one of claims 1-4, wherein the DI and the D2 each comprise an oligomer comprising 2 to 20 units, wherein each oligomer is independently selected from a peptide, a nucleic acid, and an oligomer comprising repeating C1-C20 alkyl, C5-C20 aryl, C1-C20 heteroalkyl, C2-C20 heteroaryl, C3-C20 cycloalkyl, or C1-C20 heterocycloalkyl units, and combinations thereof.
6. The scaffold of any one of claims 1-5, wherein the DI and the D2 each comprise an oligomer comprising 2 to 10 C1-C20 heterocycloalkyl units, wherein each unit is substituted with a side chain, wherein each side chain is independently at each occurrence selected from hydrogen, a natural amino acid side chain, a non-natural amino acid side chain, a Ci- C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof, wherein R* is independently at each occurrence selected from hydrogen and a C1-C20 alkyl.
7. The scaffold of any one of claims 1-6, wherein the DI and the D2 each comprise an oligomer comprising 2 to 10 oxopiperazine units.
8. The scaffold of any one of claims 1-5, wherein the DI and the D2 each comprise a peptide consisting of 7 to 10 amino acids, wherein each amino acid side chain is independently selected from hydrogen, a natural amino acid side chain, a non-natural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof,
wherein R* is independently at each occurrence selected from hydrogen and a C1-C20 alkyl.
9. The scaffold of any one of claims 1-5 and 8, wherein the DI and the D2 each comprise a peptide consisting of 8 or 9 amino acids.
10. The scaffold of any one of claims 1-5 and 8-9, wherein the DI and the D2 each comprise a peptide consisting of 8 amino acids having a formula aal-aa2-aa3-aa4-aa5-aa6-aa7-aa8, wherein aal, aa2, aa3, aa4, aa5, aa6, aa7, and aa8 are each an amino acid having a side chain independently selected from hydrogen, a natural amino acid side chain, a non-natural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, - OR*, and combinations thereof.
12. The scaffold of any one of claims 1-5 and 8-11, wherein the DI and the D2 each comprise a peptide consisting of 9 amino acids having a formula aa0-aal-aa2-aa3-aa4-aa5-aa6-aa7- aa8, wherein aaO, aal, aa2, aa3, aa4, aa5, aa6, aa7, and aa8 are each an amino acid having a side chain independently selected from hydrogen, a natural amino acid side chain, a nonnatural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, - OR*, and combinations thereof.
13. The scaffold of any one of claims 1-5 and 8-12, wherein the DI and the D2 each comprise a peptide having a formula G-R-aa2-aa3-H-aa5-aa6-aa7-aa8, wherein G is glycine, R is as defined in claim 11, and H is histidine.
14. The scaffold of any one of claims 1-5 and 8-13, wherein the DI and the D2 each comprise a peptide comprising at least one non-natural amino acid.
15. The scaffold of any one of claims 1-5 and 8-14, wherein the D I and the D2 each comprise an a-helical peptide or peptidomimetic.
16. The scaffold of any one of claims 1 -15, wherein the spatial distance between the DI and the D2 is about 5 A to about 25 A.
17. The scaffold of any one of claims 1-16, wherein the angle between DI and D2 is between about 50° and about 70°.
18. The scaffold of any one of claims 1-17, wherein the angle between DI and D2 is about 60° to about 65°.
19. The scaffold of any one of claims 1-18, wherein the angle between DI and D2 is about 63.7°
20. The scaffold of any one of claims 1-19, wherein the spatial distance between aal of DI and aal of D2 is about 5A to about 15 A, and the spatial distance between aa8 of DI and aa8 of D2 is about 15A to about 25 A.
21. The scaffold of any one of claims 1-20, wherein the linker L is about 5A to about 15 A in length.
22. The scaffold of any one of claims 1-21, wherein the linker L is a reversible linker.
23. The scaffold of claim 22, wherein the reversible linker comprises a disulfide (-S-S-) bond.
24. The scaffold of any one of claims 1-21, wherein the linker L is an irreversible linker.
25. The scaffold of any one of claims 1-21, wherein the linker L comprises a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, each of which may be optionally substituted by one or more of halo, N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof.
26. The scaffold of any one of claims 1-25, wherein the linker L comprises a combination of a C1-C20 alkyl and a C5-C20 aryl.
27. The scaffold of any one of claims 1-26, wherein the linker L comprises a cis-stilbene, azobenzene, meta-xylene, l,3-di(1H-l,2,3-triazol-l-yl)benzene, l,3-di( 1H-l,2,3-triazol-4- yl)benzene, or l,4-diphenyl- 1H-l,2,3-triazole.
28. The scaffold of any one of claims 1-26, wherein the linker L has a structure selected from:
30. A compound having a structure according to Formula (I), Formula (II), Formula (III), Formula (IV), Formula (V), Formula (VI), or Formula (VII):
Formula (VII), or a pharmaceutically acceptable salt thereof, wherein:
Q is independently at each occurrence selected from -O-, -N-, -C-, -S-, arylene, and amino acid side chain;
Ri and R12 are independently selected from hydrogen, an amino acid, an amino acid dimer, each of which amino acid is optionally substituted with a fluorescent tag, a C1-C20 alkyl, a C5-C20 aryl, -N(R*)2, -(C=O)R*, -(C=S)R*, -(C=O)NR*, -(C=S)NR*, -CO2R*, -(C=S)OR*, -(C=O)SR*, and =C(R*)2.
R.2, R3, R4, R5, R6, R7, R8, R9, R10, R11, R12, R13, R14, R15, R16, R17, R18, R19, R20, R21, and R22 are independently selected from hydrogen, a natural amino acid side chain, a non-natural amino acid side chain, a C1-C20 alkyl, a C5-C20 aryl, a C1-C20 heteroalkyl, a C2-C20 heteroaryl, a C3-C20 cycloalkyl, a C1-C20 heterocycloalkyl, -N(R*)2, (-C=O)R*, -CO2R*, -OR*, and combinations thereof, wherein any two of R2, R3, R4, R5, R6, R7, R8, R9, R10, R11, R12, R13, R14, R15, R16, R17, R18, R19, R20, R21, and R22 optionally bond together to make a macrocycle, and wherein R* is independently at each occurrence selected from hydrogen, a C5-C20 aryl, and a Ci- C20 alkyl.
31. The compound of claim 30, wherein the compound has the structure of Formula (I).
32. The compound of claim 30, wherein the compound has the structure of Formula (II).
33. The compound of claim 30, wherein Q is -S- at each occurrence.
34. The compound of any one of claims 30-33, wherein Ri and R12 are each independently selected from alanine, β-alanine, glycine, phenylalanine, tyrosine, lysine, aspartic acid, an acetyl group, and dimer combinations thereof, each of which is optionally substituted with fluorescein.
35. The compound of any one of claims 30-34, wherein R11 and R22 are each -NH2, -OH, - NHR*, -OR*, -SH, -SR*, or -NR*-N(R*)2, wherein R* is independently at each occurrence selected from hydrogen, a C5-C20 aryl, and a C1-C20 alkyl.
36. The compound of any one of claims 30-35, wherein R2 and R13 are each hydrogen.
37. The compound of any one of claims 30-36, wherein R3 and R14 are each a side chain of arginine.
38. The compound of any one of claims 30-37, wherein R4 and R15 are each independently selected from a side chain of lysine and a side chain of alanine.
39. The compound of any one of claims 30-38, wherein R5 and R16 are each a side chain of arginine.
40. The compound of any one of claims 30-39, wherein R6 and R17 are each independently selected from a side chain of histidine and a side chain of alanine.
41. The compound of any one of claims 30-40, wherein R7 and R18 are each independently selected from a side chain of lysine and a side chain of alanine.
42. The compound of any one of claims 30-41, wherein Rs and R19 are each a side chain of leucine.
43. The compound of any one of claims 30-42, wherein R9 and R20 are each a side chain of asparagine.
44. The compound of any one of claims 30-43, wherein R10 and R21 are each a side chain of arginine.
45. The compound of any one of claims 30-44, wherein each of R2 and R13 is hydrogen, each of R3 and R14 is a side chain of arginine, each of R4 and R15 is a side chain of lysine, each of R5 and R16 is a side chain of arginine, each of R6 and R17 is a side chain of histidine, each of R7 and R18 is a side chain of lysine, each of R8 and R19 is a side chain of leucine, each of R9 and R20 is a side chain of asparagine, each of R10 and R21 is a side chain of arginine, R11 and R22 are each NH2, and Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-β-alanine-glycine.
46. The compound of any one of claims 30-44, wherein each of R2 and R13 is hydrogen, each of R3 and R14 is a side chain of arginine, each of R4 and R15 is a side chain of alanine, each
of R5 and Ri6 is a side chain of arginine, each of R6 and R17 is a side chain of histidine, each of R7 and Ris is a side chain of alanine, each of R8 and R19 is a side chain of leucine, each of R9 and R20 is a side chain of asparagine, each of Rio and R21 is a side chain of arginine, R11 and R22 are each NH2, and Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-β-alanine-glycine.
47. The compound of any one of claims 30-44, wherein each of R2 and R13 is hydrogen, each of R3 and R14 is a side chain of alanine, each of R4 and R15 is a side chain of lysine, each of Rs and R16 is a side chain of arginine, each of R6 and R17 is a side chain of alanine, each of R7 and R18 is a side chain of lysine, each of Rs and R19 is a side chain of leucine, each of R9 and R20 is a side chain of asparagine, each of Rio and R21 is a side chain of arginine, R11 and R22 are each NH2, and Ri and R12 are each independently selected from tyrosine, acyl-tyrosine, and fluorescin-P-alanine-glycine.
48. The compound of any one of claims 45-47, wherein each amino acid side chain is the L- isomer.
49. The compound of any one of claims 30-44, wherein a lysine residue and an aspartic acid residue are cyclized as a lactam macrocycle.
51. The compound of any one of claims 1-5 and 8-50 having the structure selected from the group consisting of (SEQ ID NOS 10, 9, 34-35, and 12, 11, 14, 13, and 15-16 disclosed below, respectively, in order of appearance):
53. A pharmaceutical composition comprising the compound of any of claims 1-52 and a pharmaceutically acceptable carrier.
54. A pharmaceutical dosage form comprising the compound of any of claims 1-52 or the pharmaceutical composition of claim 53.
55. A method of binding double-stranded RNA comprising contacted said double- stranded RNA with the compound of any of claims 1-52, or the pharmaceutical composition of claim 53, or the pharmaceutical dosage form of claim 54.
56. The method of claim 55, wherein the contacting is in a cell.
57. The method of claim 56, wherein the cell is a bacterial cell.
58. The method of claim 56, wherein the cell is a viral cell.
59. The method of claim 56, wherein the cell is a mammalian cell.
60. The method of claim 56 or claim 59, wherein the cell is a human cell.
61. The method of any one of claims 56 and 59-60, wherein the cell is a cancer cell, a brain cell, a connective tissue cell, or a muscle cell.
62. A method of treating a repeat expansion disorder in a subject in need of such treatment, the method comprising administering to the subject an effective amount of the compound of any of claims 1-52, or the pharmaceutical composition of claim 53, or the pharmaceutical dosage form of claim 54.
63. The method of claim 62, wherein the disorder is a noncoding repeat expansion disorder.
64. The method of claim 62 or claim 63, wherein the disorder is selected from the group consisting of Baratela-Scott syndrome; CANVAS, cerebellar ataxia, neuropathy and vestibular areflexia syndrome, myotonic dystrophy type 1; DM2, myotonic dystrophy type 2, progressive myoclonus epilepsy type 1 (Unverricht-Lundborg disease), familial adult myoclonic epilepsy, Fuchs endothelial corneal dystrophy type 3, fragile XE syndrome; FRDA, Friedreich ataxia, frontotemporal dementia/amyotrophic lateral sclerosis, Fragile X-associated premature ovarian infertility, fragile X syndrome, fragile X-associated tremor ataxia syndrome, global developmental delay, progressive ataxia, and elevated glutamine, neuronal intranuclear inclusion disease, oculopharyngodistal myopathy type 1, oculopharyngodistal myopathy type 2, oculopharyngeal myopathy with leukoencephalopathy type 1, spinocerebellar ataxia, and X-linked dystonia parkinsonism.
65. The method of claim 62, wherein the disorder is a coding repeat expansion disorder.
66. The method of claim 62 or claim 65, wherein the disorder is selected from the group consisting of brachydactyly and cleidocranial dysplasia (BCCD), blepharophimosis, ptosis and epicanthus inversus (BPES), congenital central hypoventilation syndrome (CCHS), dentatorubral-pallidoluysian atrophy (DRPLA), early infantile epileptic encephalopathy type 1 (EIEE1 ), Huntington disease, Huntington disease-like 2, hand-foot-genital syndrome, holoprosencephaly type 5, mental retardation with isolated growth hormone deficiency (MRGH), oculopharyngeal muscular dystrophy, spinal and bulbar muscular atrophy, synpolydactyly type 1, and spinocerebellar ataxia.
67. The method of any one of claims 62-66, wherein the repeat is a repeat RNA selected from CUG, CAG, CAA, GAA, CGG, CCG, GCG, GCA, GCC, GCU, AUUCU, AUUGU, CAGCUG, CUGCAG, or CCCCGCCCCGCG (SEQ ID NO: 25).
68. The method of claim 67, wherein the repeat is CUG.
69. A method of treating Huntington’s disease in a subject in need of such treatment, the method comprising administering to the subject an effective amount of the compound of any of claims 1-52, or the pharmaceutical composition of claim 53, or the pharmaceutical dosage form of claim 54.
70. The method of any one of claims 62-69, wherein the subject is human.
71. The method of any one of claims 62-69, wherein the subject is a mammal selected from the group consisting of a cat, a dog, a cow, a horse, a sheep, and a pig.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363521234P | 2023-06-15 | 2023-06-15 | |
| US63/521,234 | 2023-06-15 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2024259294A2 true WO2024259294A2 (en) | 2024-12-19 |
| WO2024259294A3 WO2024259294A3 (en) | 2025-05-08 |
Family
ID=93852778
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/US2024/034092 Ceased WO2024259294A2 (en) | 2023-06-15 | 2024-06-14 | Proteomimetic scaffolds and uses thereof |
Country Status (1)
| Country | Link |
|---|---|
| WO (1) | WO2024259294A2 (en) |
-
2024
- 2024-06-14 WO PCT/US2024/034092 patent/WO2024259294A2/en not_active Ceased
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024259294A3 (en) | 2025-05-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10793605B2 (en) | Conformationally-preorganized, miniPEG-containing γ-peptide nucleic acids | |
| STORY et al. | Side‐product formation during cyclization with HBTU on a solid support | |
| JP5199126B2 (en) | Synthesis of glucagon-like peptides | |
| Liu et al. | A synthetic route to human insulin using isoacyl peptides | |
| Shoulders et al. | The aberrance of the 4 S diastereomer of 4-hydroxyproline | |
| JP2001502673A (en) | Chiral peptide nucleic acid | |
| Shin et al. | Impact of Backbone Pattern and Residue Substitution on Helicity in α/β/γ-Peptides | |
| Fowler et al. | Synthesis and characterization of nitroaromatic peptoids: fine tuning peptoid secondary structure through monomer position and functionality | |
| Rump et al. | Cyclotriveratrylene (CTV) as a new chiral triacid scaffold capable of inducing triple helix formation of collagen peptides containing either a native sequence or Pro‐Hyp‐Gly repeats | |
| Doan et al. | Solid‐phase synthesis of C‐terminal azapeptides | |
| Belgi et al. | Alkyne-bridged α-conotoxin Vc1. 1 potently reverses mechanical allodynia in neuropathic pain models | |
| Tian et al. | Continuous solid-phase synthesis and disulfide cyclization of peptide− PNA− peptide chimeras | |
| US20090131306A1 (en) | Chemically modified cyclic peptides containing cell adhesion recognition (car) sequences and uses therefor | |
| WO2024259294A2 (en) | Proteomimetic scaffolds and uses thereof | |
| Otani et al. | Oligomers of β-amino acid bearing non-planar amides form ordered structures | |
| EP2569327B1 (en) | Eif4e binding peptides | |
| US11162083B2 (en) | Peptide based inhibitors of Raf kinase protein dimerization and kinase activity | |
| Ditmangklo et al. | Synthesis of pyrrolidinyl PNA and its site-specific labeling at internal positions by click chemistry | |
| Caporale et al. | Side chain cyclization based on serine residues: synthesis, structure, and activity of a novel cyclic analogue of the parathyroid hormone fragment 1− 11 | |
| Alexander McNamara et al. | Peptides Constrained by an Aliphatic Linkage between Two C α Sites: Design, Synthesis, and Unexpected Conformational Properties of an i,(i+ 4)-Linked Peptide | |
| Boutard et al. | Examination of the active secondary structure of the peptide 101.10, an allosteric modulator of the interleukin‐1 receptor, by positional scanning using β‐amino γ‐lactams | |
| Hurevich et al. | Rational conversion of noncontinuous active region in proteins into a small orally bioavailable macrocyclic drug-like molecule: the HIV-1 CD4: gp120 paradigm | |
| Beck-Sickinger | Synthesis of conformationally restricted peptides | |
| Rai | Retroinverso mimetics of S peptide | |
| JPWO2016129680A1 (en) | CTB-PI polyamide conjugates that activate the expression of specific genes |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| NENP | Non-entry into the national phase |
Ref country code: DE |




































