EP4676545A1 - Synthetic promoters - Google Patents
Synthetic promotersInfo
- Publication number
- EP4676545A1 EP4676545A1 EP24712568.5A EP24712568A EP4676545A1 EP 4676545 A1 EP4676545 A1 EP 4676545A1 EP 24712568 A EP24712568 A EP 24712568A EP 4676545 A1 EP4676545 A1 EP 4676545A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- seq
- sftpb
- vector
- enhancer
- nucleic acid
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
- C12N15/79—Vectors or expression systems specially adapted for eukaryotic hosts
- C12N15/85—Vectors or expression systems specially adapted for eukaryotic hosts for animal cells
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
- A61K48/00—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy
- A61K48/005—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy characterised by an aspect of the 'active' part of the composition delivered, i.e. the nucleic acid delivered
- A61K48/0058—Nucleic acids adapted for tissue specific expression, e.g. having tissue specific promoters as part of a contruct
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2740/00—Reverse transcribing RNA viruses
- C12N2740/00011—Details
- C12N2740/10011—Retroviridae
- C12N2740/15011—Lentivirus, not HIV, e.g. FIV, SIV
- C12N2740/15032—Use of virus as therapeutic agent, other than vaccine, e.g. as cytolytic agent
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2740/00—Reverse transcribing RNA viruses
- C12N2740/00011—Details
- C12N2740/10011—Retroviridae
- C12N2740/15011—Lentivirus, not HIV, e.g. FIV, SIV
- C12N2740/15041—Use of virus, viral particle or viral elements as a vector
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2740/00—Reverse transcribing RNA viruses
- C12N2740/00011—Details
- C12N2740/10011—Retroviridae
- C12N2740/15011—Lentivirus, not HIV, e.g. FIV, SIV
- C12N2740/15041—Use of virus, viral particle or viral elements as a vector
- C12N2740/15043—Use of virus, viral particle or viral elements as a vector viral genome or elements thereof as genetic vector
Definitions
- the present invention relates to nucleic acid cassettes for gene therapy, particularly to promoter and promoter/enhancer combinations for improved expression of transgenes in a lung parenchyma ⁇ specific/preferred manner.
- the invention further relates nucleic acid cassettes comprising said promoters and promoter/enhancer combinations, viral and non ⁇ viral vectors comprising such nucleic acid cassettes, and the use of such nucleic acid cassettes and vectors to increase expression of therapeutic proteins by lung parenchyma cells.
- SFTPB deficiency is a severe monogenic interstitial lung disorder that leads to loss of life in infants as a result of alveolar collapse and respiratory distress syndrome.
- the only curative treatment is thought to be lung transplantation; however, the lack of suitable donor organs makes this a non ⁇ viable option in most circumstances.
- nucleic acids as medicine, or gene therapy, is a promising new treatment modality, both for SFTPB deficiency and other genetic diseases, including genetic diseases of the respiratory tract.
- gene therapies currently in use or under development are not effective at curing diseases is because it is difficult to make sufficient protein to reach the therapeutic threshold needed to treat or cure the disease.
- the present inventors have previously developed a lentiviral vector, which has been pseudotyped with hemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, comprising a promoter and a transgene.
- the backbone of the vector is from a simian immunodeficiency virus (SIV), such as SIV1 or African green monkey SIV (SIV ⁇ AGM).
- SIV ⁇ AGM African green monkey SIV
- the backbone of a viral vector of the invention is from SIV ⁇ AGM.
- the HN and F proteins function, respectively, to attach to sialic acids and mediate cell fusion for vector entry to target cells.
- the present inventors discovered that this specifically F/HN ⁇ pseudotyped lentiviral vector can efficiently transduce airway epithelium, resulting in transgene expression sustained for periods beyond the proposed lifespan of airway epithelial cells. Importantly, the present inventors also found that re ⁇ administration does not result in a loss of efficacy. These features make the vectors of the present invention attractive candidates for treating diseases via their use in expressing therapeutic proteins: (i) within the cells of the respiratory tract; (ii) secreted into the lumen of the respiratory tract; and (iii) secreted into the circulatory system. However, even using this state ⁇ of ⁇ the ⁇ art platform technology, the levels of transgene expressed are at the lower predicted threshold required for clinical efficacy.
- exogenous signal peptides can be used to increase expression and secretion of therapeutic proteins by airway cells.
- exogenous signal peptides it is possible to produce more protein for every copy of a gene therapy vector or transgene that is put into a cell, increasing the dose of therapeutic protein without increasing the amount of gene therapy vector given to a patient.
- exogenous signal peptides is not appropriate for all therapeutic proteins or all conditions. For example, not all therapeutic proteins are secreted, and for some there may be clinical reasons why manipulating the signal peptides is undesirable.
- gene therapy vectors comprising promoters providing high and sustained gene expression in a variety of cell types are preferred, especially in a therapeutic context. For this reason, the inventors previously used a hCEF promoter to drive strong and persistent expression in mouse lung with non ⁇ viral formulations. However for some conditions, expression in a specific tissue or cell type is required to (i) achieve desired therapeutic target, (ii) avoid gene expression ⁇ related toxicities, and (iii) circumvent immune responses to the therapeutic agent stemming from gene expression in undesired cell types.
- Such promoters, cassettes and vectors may be of particular use in the treatment of genetic diseases, particularly genetic respiratory diseases, such as surfactant protein deficiency.
- genetic diseases particularly genetic respiratory diseases, such as surfactant protein deficiency.
- SUMMARY OF THE INVENTION At present, there remains a pressing need for technology that enables high levels of cell ⁇ or tissue ⁇ specific expression of transgenes for gene therapy, including from the inventors’ own lentiviral platform.
- novel promoters that drive cell ⁇ specific expression in the lung parenchyma.
- the present inventors have developed a panel of novel promoter sequences comprising a functional fragment of the human SFTPB gene promoter.
- the inventors have surprisingly shown that the SFTPB promoter fragment of the invention drives increased transgene expression compared with the full ⁇ length SFTPB promoter. Furthermore, the inventors have also surprisingly demonstrated that the SFTPB promoter fragment of the invention can be combined with particular enhancers, including some enhancers which are not associated with cell ⁇ specific expression in the lung parenchyma, to further improve transgene expression. Accordingly, the present invention provides an SFTPB promoter fragment which comprises or consists of at least part of exon 1 of the SFTPB gene, but does not comprise the SFTPB gene start codon. Typically said promoter is less than 800 bases in length, preferably less than 700 bases in length.
- Said promoter may comprise or consist of: (a) SEQ ID NO: 1 or a sequence with at least 80% identity to SEQ ID NO: 1; or (b) bases 81 ⁇ 710 of SEQ ID NO: 2, or a sequence with at least 80% identity to bases 81 ⁇ 710 of SEQ ID NO: 2; wherein optionally (i) said SFTPB promoter fragment further comprises up to 20 bases at the 5’ end, which may optionally correspond to up to 20 bases 5’ to base 81 of SEQ ID NO: 2; and/or (ii) said SFTPB promoter fragment further comprises up to 4 bases at the 3’ end, which may optionally correspond to up to 4 bases 3’ to base 710 of SEQ ID NO: 2.
- the SFTPB promoter fragment may comprise or consist of SEQ ID NO: 1, or a sequence with at least 90% identity to SEQ ID NO: 1.
- the SFTPB promoter fragment of the invention may further comprise a 5’ enhancer.
- Said enhancer may be in (i) the forward, or (ii) the reverse, orientation; preferably wherein the enhancer is in the forward orientation.
- Said enhancer may be selected from a SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer, an ELF3 enhancer, an actin enhancer, an LMO7 enhancer, a SFTPC enhancer or a SFTPB enhancer.
- the enhancer may be selected from a SLC34A2 enhancer, a VEGFA enhancer or a CMV enhancer.
- the enhancer may be selected from (a) an SLC34A2 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 29 ⁇ 33, preferably SEQ ID NO: 31; (b) a VEGFA enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38; (c) a CMV enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11; (d) a SV40 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or 35; (e) an ELF3 enhancer which comprises or consists of a nucleotide sequence
- the SFTPB promoter fragment of the invention may comprise or consist of a nucleic acid sequence of any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 49, or a sequence with at least 90% identity to any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 48.
- the SFTPB promoter fragment of the invention comprises or consists of a nucleic acid sequence of any one of SEQ ID NOs: 46, 47 or 48, or a sequence with at least 90% identity to any one of SEQ ID NOs: 46, 47 or 48.
- the invention also provides a nucleic acid cassette comprising: (a) an SFTPB promoter fragment of the invention; and (b) a transgene.
- Said transgene may encode a therapeutic protein selected from: (a) secreted therapeutic protein selected from: Surfactant Protein B (SP ⁇ B), Surfactant Protein C (SP ⁇ C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF), decorin, an anti ⁇ inflammatory protein (e.g. IL ⁇ 10 or TGF ⁇ ) or monoclonal antibody, an anti ⁇ inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP ⁇ binding cassette sub ⁇ family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.
- a therapeutic protein selected from: (a) secreted therapeutic protein selected from: Surfactant Protein B (SP ⁇ B), Surfactant Protein C (SP ⁇ C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor
- the SFTPB promoter fragment may increase expression of the transgene by lung parenchymal cells, optionally compared with the full ⁇ length SFTPB promoter. Expression of the transgene may be increased by at least 2 ⁇ fold, preferably at least 5 ⁇ fold compared with the full ⁇ length SFTPB promoter.
- the lung parenchymal cells may comprise one or more cell type selected from: alveolar type I epithelial (ATI) cells, alveolar type II epithelial cells (ATII), and/or club cells, preferably ATII and/or ATI cells.
- Expression of the transgene by a nucleic acid cassette or promoter of the invention may be specific to lung parenchymal cells; and/or the ratio of lung expression: nose expression of the transgene by the SFTPB promoter fragment is at least 2:1.
- the invention further provides a gene therapy vector, comprising a nucleic acid cassette of the invention.
- Said gene therapy vector may be a non ⁇ viral vector, wherein optionally: (a) the non ⁇ viral vector is a plasmid; and/or (b) the non ⁇ viral vector is comprised in a cationic liposome, which preferably comprises GL67A.
- Said gene therapy vector may be a viral vector, optionally selected from: a lentiviral vector; an AAV vector; and an adenoviral vector.
- Said lentiviral vector may be pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, optionally from a Sendai virus.
- Said lentiviral vector may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector.
- HIV Human immunodeficiency virus
- SIV Simian immunodeficiency virus
- FV Feline immunodeficiency virus
- EIAV Equine infectious anaemia virus
- Visna/maedi virus vector a Visna/maedi virus vector.
- said lentiviral vector is a SIV vector.
- the invention further provides a method of expressing a therapeutic protein in a target cell, comprising delivering a nucleic acid cassette of the invention or a gene therapy vector of the invention into the target cells.
- Said delivering may comprise integrating said nucleic acid cassette or gene therapy vector into said target cell's genome.
- the invention also provides a gene therapy vector of the invention for use in a method of treating a disease.
- Said disease may be a genetic disease.
- the disease may be: (a) a respiratory disease, particularly a genetic respiratory disease; or (b) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder.
- the disease may be selected from Surfactant Protein B (SP ⁇ B) Deficiency; Surfactant Protein C (SP ⁇ C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1 ⁇ antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID ⁇ 19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia.
- SP ⁇ B Surfactant Protein B
- SP ⁇ C Surfactant Protein C
- Pulmonary surfactant metabolism dysfunction 2 SMDP2
- Pulmonary surfactant metabolism dysfunction 3 SMDP3
- PCD Primary Cilia
- the invention further provides a cell comprising a nucleic acid cassette of the invention or a gene therapy vector of the invention.
- the invention also provides a composition comprising a nucleic acid cassette of the invention or a gene therapy vector of the invention and a pharmaceutically acceptable carrier, diluent or excipient.
- lentiviral vectors such as the lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus as exemplified herein
- HN haemagglutinin ⁇ neuraminidase
- F fusion proteins from a respiratory paramyxovirus
- the invention also provides lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered simultaneously or sequentially with a surfactant.
- HN haemagglutinin ⁇ neuraminidase
- F fusion
- said disease is: (a) a genetic disease; (b) a respiratory disease, particularly a genetic respiratory disease; (c) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder; and/or (d) selected from Surfactant Protein B (SP ⁇ B) Deficiency; Surfactant Protein C (SP ⁇ C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1 ⁇ antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID ⁇ 19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia.
- SP ⁇ B Surfactant Protein B
- SP ⁇ C
- the respiratory paramyxovirus may be a Sendai virus; and/or the lentiviral vector may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector, preferably a SIV vector
- the transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP ⁇ B), Surfactant Protein C (SP ⁇ C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF), decorin, an anti ⁇ inflammatory protein (e.g.
- the promoter may comprise a SFTPB promoter fragment as defined herein; and/or the lentiviral vector may comprise a nucleic acid cassette as defined herein.
- the lentiviral vector may be administered before the surfactant.
- the surfactant may be administered before the lentiviral vector.
- the lentiviral vector and surfactant may be administered simultaneously, optionally wherein the lentiviral vector and surfactant are mixed prior to administration.
- FIG. 2 Graphs quantifying EGFP expression in human HEK293T cells transduced with recombinant SIV lentiviral vectors pseudotyped with VSV ⁇ G and expressing the EGFP transgene from a range of promoters (CMV (SEQ ID NO: 6), hCEF ( (SEQ ID NO: 5), EF1aS ( (SEQ ID NO: 7), PGK ( (SEQ ID NO: 8), mSPB ( (SEQ ID NO: 1), fSPB (SEQ ID NO: 4), mSPC (SEQ ID NO: 9), fSPC(SEQ ID NO: 10)) at a multiplicity of infection (moi) of 0.5 or 1.0.
- CMV SEQ ID NO: 6
- hCEF (SEQ ID NO: 5)
- EF1aS (SEQ ID NO: 7)
- PGK (SEQ ID NO: 8)
- mSPB (SEQ ID NO: 1), fSPB (SEQ
- FIG. 3 Experimental schematics for experiments to test EGFP expression in vivo using lentiviral vectors expressing EGFP from different promoters. A Schematic showing timing of dosing, on Day 0 mice were dosed with lentiviral vectors via nasal instillation and 7 days post ⁇ dosing the mice were culled and lung tissue harvested for cryosections. B Schematic showing treatment groups.
- FIG. 5 Panel of ATII specific genes (and enhancers) generated by interrogation of LungGENS and a tissue expression database.
- Figure 6 Schematics of exemplary mSPB/enhancer constructs. Different lengths (indicated in base pairs (bp)) of the newly identified enhancer sequences (boxes) were sub ⁇ cloned in front of the mSPB (SEQ ID NO: 1) promoter sequence (arrow boxes) in both forward (f) and reverse (r) orientations to generate candidate synthetic promoters expressing the EGFP2ALux reporter transgene.
- FIG. 7 Graph showing expression of EGFP in HEK293T cells by different mSPB/enhancer constructs.
- Figure 8 Graph showing expression of EGFP in murine LA ⁇ 4 cells by different mSPB/enhancer constructs. P values are given where significant expression was observed.
- FIG. 9 Graph showing expression of EGFP in human SALI cells by different mSPB/enhancer constructs. P values are given where significant expression was observed.
- RLU Relative Light Units
- Figure 10 Graph showing expression of EGFP in human SALI cells by lentiviral vectors with different mSPB/enhancers in a secondary screen. P values are given where significant expression was observed.
- CMV SEQ ID NO: 6
- hCEF SEQ ID NO: 5
- mSPB (SEQ ID NO: 1)
- Luciferase expression levels (Relative Light Units (RLU)) were determined and compared with na ⁇ ve (non ⁇ transduced) control cells.
- Figure 11 Graph showing expression of EGFP in HEK293T cells by lentiviral vectors with different mSPB/enhancers in a secondary screen. P values are given where significant expression was observed.
- CMV SEQ ID NO: 6
- hCEF SEQ ID NO: 5
- Luciferase expression levels (Relative Light Units (RLU)) were determined and compared with na ⁇ ve (non ⁇ transduced) control cells.
- Figure 12 Heat map of luciferase expression in mice treated with lentiviral vectors with different mSPB/enhancers.
- CMV SEQ ID NO: 6
- hCEF SEQ ID NO: 5
- mSPB (SEQ ID NO: 1)
- Luciferase signal was observed in the nose and lung areas at all timepoints with the non ⁇ specific CMV[SEQ ID NO: 6] and hCEF[ SEQ ID NO: 5] lung promoter in line with expectations. Luciferase signal in the mSPB[SEQ IDN O: 1] group was overall lower and restricted to the lung, also as expected.
- Figure 13 Graphs showing in vivo luciferase expression by lentiviral vectors with different mSPB/enhancers (A) and the ratio of luciferase expression in the lungs and the nose (B).
- Lentivirus was dosed intranasally (1x10 7 TU per mouse in 100uL TSSM Buffer). On days 2, 7 and 28 post ⁇ dosing, mice were anaesthetised and imaged for luciferase expression.
- Figure 14 Graphs showing luciferase expression by plasmids with different mSPB/enhancers in human SALI cells (A), in vivo (B), and in HEK293T cells (C).
- the enhancer/mSPB promoter constructs expressing Lux reporter transgene were evaluated in the context of a non ⁇ viral formulation, to deliver plasmids containing the CMV[SEQ ID NO: 6], hCEF[SEQ ID NO: 5] and mSPB[SEQ ID NO: 1] promoter sequences as well as the selected enhancer/mSPB promoter constructs: slc2[SEQ ID NO: 31], veg2[SEQ ID NO: 38], sv40[SEQ ID NO: 34], and CMVenh[SEQ ID NO: 11].
- FIG. 15 Representative immunohistochemistry images showing EGFP expression in lung parenchyma sections taken from mice treated with lentiviral vectors expressing EGFP from different enhancer/mSPB constructs. Representative images from (A) CMVenh (Alv ⁇ 01, SEQ ID NO: 46), (B) Slc2 (Alv ⁇ 02, SEQ ID NO: 47) and (C) Vegf2 (Alv ⁇ 03, SEQ ID NO: 48) groups are shown.
- mice were administered D ⁇ luciferin and imaged for luciferase activity in the lung and nasal cavity.
- Signal in regions of interest ROI
- ROI Signal in regions of interest
- AUC area under the curve
- C Average radiance in the lung
- D Specificity for expression in the lung parenchyma was determined by calculating the ratio of signal in the lung (indicative of alveolar and airway cell transduction) to signal in the nasal cavity (indicative of airway cell transduction) 28 days after dosing.
- Representative images of EGFP positive cells in the airway (Aw) and parenchyma (P) are shown after administration of (A) rSIV.F/HN hCEF EGFP or (B) rSIV.F/HN mSP ⁇ B EGFP. Images of lung sections were further analysed (using Visiopharm software) which required manual indication of airways and parenchyma. The percentage of EGFP ⁇ positive cells from the total lung and from the airways was used to (C) estimate the percentage of EGFP ⁇ positive cells observed in the parenchyma from each promoter.
- FIG. 18 The effect of mSPB, Alv ⁇ 1, Alv ⁇ 2 and Alv ⁇ 3 driven SFP ⁇ B expression on transepithelial electrical resistance (TEER) in a Surfactant Air Liquid Interface (SALI) model.
- TEER values from SALI cultures generated from the H441 SP ⁇ B KO cells were lower than the parental H441 cell line at 14 days post airlift constituting a phenotypic defect.
- “capable of interacting” also means interacting
- “capable of cleaving” also means cleaves
- “capable of binding” also means binds and "capable of specifically targeting" also means specifically targets.
- Numeric ranges are inclusive of the numbers defining the range. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within this disclosure.
- “About” may generally mean an acceptable degree of error for the quantity measured given the nature or precision of the measurements. Exemplary degrees of error are within 20 percent (%), typically, within 10%, and more typically, within 5% of a given value or range of values. Preferably, the term “about” shall be understood herein as plus or minus ( ⁇ ) 5%, preferably ⁇ 4%, ⁇ 3%, ⁇ 2%, ⁇ 1%, ⁇ 0.5%, ⁇ 0.1%, of the numerical value of the number with which it is being used.
- the term “consisting essentially of''” refers to those elements required for a given invention. The term permits the presence of elements that do not materially affect the basic and novel or functional characteristic(s) of that invention (i.e. inactive or non ⁇ immunogenic ingredients).
- Embodiments described herein as “comprising” one or more features may also be considered as disclosure of the corresponding embodiments “consisting of” and/or “consisting essentially of” such features. Concentrations, amounts, volumes, percentages and other numerical values may be presented herein in a range format.
- a "vector” or “construct” refers to a macromolecule or complex of molecules comprising a polynucleotide to be delivered to a host cell, either in vitro or in vivo.
- a vector can be a linear or a circular molecule.
- a vector of the invention may be viral or non ⁇ viral.
- lentiviral vectors refers to a common type of non ⁇ viral vector.
- a plasmid is an extra ⁇ chromosomal DNA molecule separate from the chromosomal DNA which is capable of replicating independently of the chromosomal DNA.
- a plasmid is circular and may be double ⁇ stranded.
- the terms "nucleic acid cassette”, “nucleic acid construct”, “expression cassette” and “nucleic acid expression cassette” are used interchangeably to mean a nucleic acid molecule that is capable of directing transcription.
- a nucleic acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter and a nucleic acid sequence to be transcribed.
- a nucleic acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter and a nucleic acid sequence encoding a protein of interest.
- a nucleic acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter, and a nucleic acid encoding a therapeutic protein.
- a nucleic acid cassette may include additional elements, such as an enhancer, and/or a transcription termination signal.
- the terms “transduced” and “modified” are used interchangeably to describe cells which have been modified to express a transgene of interest. Typically the modification occurs through transduction of the cells.
- Titre and “yield” are used interchangeably to mean the amount of viral (e.g. lentiviral, particularly SIV) vector produced by a method of the invention.
- Titre is the primary benchmark characterising manufacturing efficiency, with higher titres generally indicating that more vector is manufactured (e.g. using the same amount of reagents).
- Titre or yield may relate to the number of vector genomes that have integrated into the genome of a target cell (integration titre), which is a measure of “active” virus particles, i.e. the number of particles capable of transducing a cell.
- Transducing units (TU/mL also referred to as TTU/mL) is a biological readout of the number of host cells that get transduced under certain tissue culture/virus dilutions conditions, and is a measure of the number of “active” virus particles.
- the total number of (active+inactive) virus particles may also be determined using any appropriate means, such as by measuring either how much Gag is present in the test solution or how many copies of viral RNA are in the test solution. Assumptions are then made that a viral (e.g. lentivirus, particularly SIV) particle contains either 2000 Gag molecules or 2 viral RNA molecules. Once total particle number and a transducing titre/TU have been measured, a particle:infectivity ratio calculated.
- amino acids are referred to herein using the name of the amino acid, the three ⁇ letter abbreviation or the single letter abbreviation. Unless otherwise indicated, any nucleic acid sequences are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively.
- protein and “polypeptide” are used interchangeably herein to designate a series of amino acid residues, connected to each other by peptide bonds between the alpha ⁇ amino and carboxyl groups of adjacent residues.
- protein refers to a polymer of amino acids, including modified amino acids (e.g., phosphorylated, glycated, glycosylated, etc.) and amino acid analogues, regardless of its size or function.
- modified amino acids e.g., phosphorylated, glycated, glycosylated, etc.
- amino acid analogues regardless of its size or function.
- Protein and polypeptide are often used in reference to relatively large polypeptides, whereas the term “peptide” is often used in reference to small polypeptides, but usage of these terms in the art overlaps.
- protein and “polypeptide” are used interchangeably herein when referring to a gene product and fragments thereof.
- polypeptides or proteins include gene products, naturally occurring proteins, homologs, orthologs, paralogs, fragments and other equivalents, variants, fragments, and analogues of the foregoing.
- polynucleotides refers to any molecule, preferably a polymeric molecule, incorporating units of ribonucleic acid, deoxyribonucleic acid or an analogue thereof.
- the nucleic acid can be either single ⁇ stranded or double ⁇ stranded.
- a single ⁇ stranded nucleic acid can be one nucleic acid strand of a denatured double ⁇ stranded DNA Alternatively, it can be a single ⁇ stranded nucleic acid not derived from any double ⁇ stranded DNA.
- the nucleic acid can be DNA.
- the nucleic acid can be RNA Suitable nucleic acid molecules are DNA, including genomic DNA or cDNA. Other suitable nucleic acid molecules are RNA, including siRNA, shRNA, and antisense oligonucleotides.
- transgene and “gene” are also used interchangeably and both terms encompass fragments or variants thereof encoding the target protein.
- transgenes of the present invention include nucleic acid sequences that have been removed from their naturally occurring environment, recombinant or cloned DNA isolates, and chemically synthesized analogues or analogues biologically synthesized by heterologous systems. Minor variations in the amino acid sequences of the invention are contemplated as being encompassed by the present invention, providing that the variations in the amino acid sequence(s) maintain at least 60%, at least 70%, more preferably at least 80%, at least 85%, at least 90%, at least 95%, and most preferably at least 97% or at least 99% sequence identity to the amino acid sequence of the invention or a fragment thereof as defined anywhere herein.
- homology is used herein to mean identity.
- sequence of a variant or analogue sequence of an amino acid sequence of the invention may differ on the basis of substitution (typically conservative substitution) deletion or insertion. Proteins comprising such variations are referred to herein as variants. Proteins of the invention may include variants in which amino acid residues from one species are substituted for the corresponding residue in another species, either at the conserved or non ⁇ conserved positions. Variants of protein molecules disclosed herein may be produced and used in the present invention. Following the lead of computational chemistry in applying multivariate data analysis techniques to the structure/property ⁇ activity relationships [see for example, Wold, et al. Multivariate data analysis in chemistry. Chemometrics ⁇ Mathematics and Statistics in Chemistry (Ed.: B. Kowalski); D.
- proteins can be derived from empirical and theoretical models (for example, analysis of likely contact residues or calculated physicochemical property) of proteins sequence, functional and three ⁇ dimensional structures and these properties can be considered individually and in combination.
- Amino acids are referred to herein using the name of the amino acid, the three ⁇ letter abbreviation or the single letter abbreviation.
- the term “protein”, as used herein, includes proteins, polypeptides, and peptides.
- amino acid sequence is synonymous with the term “polypeptide” and/or the term “protein”.
- amino acid sequence is synonymous with the term “peptide”.
- the terms "protein” and "polypeptide” are used interchangeably herein.
- the conventional one ⁇ letter and three ⁇ letter codes for amino acid residues may be used.
- the 3 ⁇ letter code for amino acids as defined in conformity with the IUPACIUB Joint Commission on Biochemical Nomenclature (JCBN). It is also understood that a polypeptide may be coded for by more than one nucleotide sequence due to the degeneracy of the genetic code. Amino acid residues at non ⁇ conserved positions may be substituted with conservative or non ⁇ conservative residues. In particular, conservative amino acid replacements are contemplated.
- a “conservative amino acid substitution” is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain.
- conservatively modified variants in a protein of the invention does not exclude other forms of variant, for example polymorphic variants, interspecies homologs, and alleles.
- Non ⁇ conservative amino acid substitutions include those in which (i) a residue having an electropositive side chain (e.g., Arg, His or Lys) is substituted for, or by, an electronegative residue (e.g., Glu or Asp), (ii) a hydrophilic residue (e.g., Ser or Thr) is substituted for, or by, a hydrophobic residue (e.g., Ala, Leu, Ile, Phe or Val), (iii) a cysteine or proline is substituted for, or by, any other residue, or (iv) a residue having a bulky hydrophobic or aromatic side chain (e.g., Val, His, Ile or Trp) is substituted for, or by, one having a smaller side chain (e.g., Ala or Ser) or no side chain (e.g., Gly).
- an electropositive side chain e.g., Arg, His or Lys
- an electronegative residue e.g., Glu or As
- “Insertions” or “deletions” are typically in the range of about 1, 2, or 3 amino acids. The variation allowed may be experimentally determined by systematically introducing insertions or deletions of amino acids in a protein using recombinant DNA techniques and assaying the resulting recombinant variants for activity. This does not require more than routine experiments for a skilled person.
- a “fragment” of a polypeptide comprises at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97% or more of the original polypeptide.
- a fragment may comprise at least 5, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20 or more amino acids of the protein from which it is derived.
- a fragment may be continuous or discontinuous, preferably continuous.
- the polynucleotides of the present invention may be prepared by any means known in the art. For example, large amounts of the polynucleotides may be produced by replication in a suitable host cell.
- the natural or synthetic DNA fragments coding for a desired fragment will be incorporated into recombinant nucleic acid constructs, typically DNA constructs, capable of introduction into and replication in a prokaryotic or eukaryotic cell.
- DNA constructs will be suitable for autonomous replication in a unicellular host, such as yeast or bacteria, but may also be intended for introduction to and integration within the genome of a cultured insect, mammalian, plant or other eukaryotic cell lines.
- the polynucleotides of the present invention may also be produced by chemical synthesis, e.g. by the phosphoramidite method or the tri ⁇ ester method, and may be performed on commercial automated oligonucleotide synthesizers.
- a double ⁇ stranded fragment may be obtained from the single stranded product of chemical synthesis either by synthesizing the complementary strand and annealing the strand together under appropriate conditions or by adding the complementary strand using DNA polymerase with an appropriate primer sequence.
- the term “isolated” in the context of the present invention denotes that the polynucleotide sequence has been removed from its natural genetic milieu and is thus free of other extraneous or unwanted coding sequences (but may include naturally occurring 5' and 3' untranslated regions such as promoters and terminators), and is in a form suitable for use within genetically engineered protein production systems. Such isolated molecules are those that are separated from their natural environment. In view of the degeneracy of the genetic code, considerable sequence variation is possible among the polynucleotides of the present invention.
- Degenerate codons encompassing all possible codons for a given amino acid are set forth below: Amino Acid Codons Degenerate Codon Cys TGC TGT TGY Ser AGC AGT TCA TCC TCG TCT WSN Thr ACA ACC ACG ACT ACN Pro CCA CCC CCG CCT CCN Ala GCA GCC GCG GCT GCN Gly GGA GGC GGG GGT GGN Asn AAC AAT AAY Asp GAC GAT GAY Glu GAA GAG GAR Gln CAA CAG CAR His CAC CAT CAY Arg AGA AGG CGA CGC CGG CGT MGN Lys AAA AAG AAR Met ATG ATG Ile ATA ATC ATT ATH Leu CTA CTC CTG CTT TTA TTG YTN Val GTA GTC GTG GTT GTN Phe TTC TTT TTY Tyr TAC TAT TAY Trp TGG TGG Ter TAA TAG TGA TRR Asn/ Asp RAY Glu
- variant amino acid sequences may encode variant amino acid sequences, but one of ordinary skill in the art can easily identify such variant sequences by reference to the amino acid sequences of the present invention.
- a “variant” nucleic acid sequence has substantial homology or substantial similarity to a reference nucleic acid sequence (or a fragment thereof).
- a nucleic acid sequence or fragment thereof is “substantially homologous” (or “substantially identical”) to a reference sequence if, when optimally aligned (with appropriate nucleotide insertions or deletions) with the other nucleic acid (or its complementary strand), there is nucleotide sequence identity in at least about 70%, 75%, 80%, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or more% of the nucleotide bases. Methods for homology determination of nucleic acid sequences are known in the art.
- a “variant” nucleic acid sequence is substantially homologous with (or substantially identical to) a reference sequence (or a fragment thereof) if the “variant” and the reference sequence they are capable of hybridizing under stringent (e.g. highly stringent) hybridization conditions.
- Nucleic acid sequence hybridization will be affected by such conditions as salt concentration (e.g. NaCl), temperature, or organic solvents, in addition to the base composition, length of the complementary strands, and the number of nucleotide base mismatches between the hybridizing nucleic acids, as will be readily appreciated by those skilled in the art.
- Stringent temperature conditions are preferably employed, and generally include temperatures in excess of 30°C, typically in excess of 37°C and preferably in excess of 45°C.
- Stringent salt conditions will ordinarily be less than 1000 mM, typically less than 500 mM, and preferably less than 200 mM.
- the pH is typically between 7.0 and 8.3.
- Methods of determining nucleic acid percentage sequence identity are known in the art. By way of example, when assessing nucleic acid sequence identity, a sequence having a defined number of contiguous nucleotides may be aligned with a nucleic acid sequence (having the same number of contiguous nucleotides) from the corresponding portion of a nucleic acid sequence of the present invention.
- Tools known in the art for determining nucleic acid percentage sequence identity include Nucleotide BLAST (as described below).
- preferential codon usage refers to codons that are most frequently used in cells of a certain species, thus favouring one or a few representatives of the possible codons encoding each amino acid.
- the amino acid threonine (Thr) may be encoded by ACA, ACC, ACG, or ACT, but in mammalian host cells ACC is the most commonly used codon; in other species, different codons may be preferential.
- Preferential codons for a particular host cell species can be introduced into the polynucleotides of the present invention by a variety of methods known in the art.
- any nucleic acid sequence may be codon ⁇ optimised for expression in a host or target cell.
- the vector genome or corresponding plasmid
- the REV gene or corresponding plasmid
- the fusion protein (F) gene or correspond plasmid
- the hemagglutinin ⁇ neuraminidase (HN) gene or corresponding plasmid, or any combination thereof may be codon ⁇ optimised.
- a “fragment” of a polynucleotide of interest comprises a series of consecutive nucleotides from the sequence of said full ⁇ length polynucleotide.
- a “fragment” of a polynucleotide of interest may comprise (or consist of) at least 600 consecutive nucleotides from the sequence of said polynucleotide (e.g. at least 600, 650, 700, 750, 800 850, 900, or 950 consecutive nucleic acid residues of said polynucleotide).
- a fragment as defined herein retains the same function as the full ⁇ length polynucleotide.
- the terms “decrease”, “reduced”, “reduction”, or “inhibit” are all used herein to mean a decrease by a statistically significant amount.
- the terms “reduce,” “reduction” or “decrease” or “inhibit” typically means a decrease by at least 10% as compared to a reference level (e.g.
- the terms “increased”, “increase”, “enhance”, or “activate” are all used herein to mean an increase by a statically significant amount.
- the terms “increased”, “increase”, “enhance”, or “activate” can mean an increase of at least 25%, at least 50% as compared to a reference level, for example an increase of at least about 50%, or at least about 75%, or at least about 80%, or at least about 90%, at least about 95%, or at least about 98%, or at least about 99%, or at least about 100%, or at least about 250% or more compared with a reference level, or at least about a 1.5 ⁇ fold, or at least about a 2 ⁇ fold, or at least about a 2.5 ⁇ fold, or at least about a 3 ⁇ fold, or at least about a 4 ⁇ fold, or at least about a 5 ⁇ fold or at least about a 10 ⁇ fold increase, or any increase between 1.5 ⁇ fold and 10 ⁇ fold or greater as compared to a reference level.
- an “increase” is an observable or statistically significant increase in such level.
- the terms “individual”, “subject”, and “patient”, are used interchangeably herein to refer to a mammalian subject for whom diagnosis, prognosis, disease monitoring, treatment, therapy, and/or therapy optimisation is desired.
- the mammal can be (without limitation) a human, non ⁇ human primate, mouse, rat, dog, cat, horse, or cow.
- the individual, subject, or patient is a human.
- An “individual” may be an adult, juvenile or infant.
- An “individual” may be male or female.
- a "subject in need" of treatment for a particular condition can be an individual having that condition, diagnosed as having that condition, or at risk of developing that condition.
- a subject can be one who has been previously diagnosed with or identified as suffering from or having a condition in need of treatment or one or more complications or symptoms related to such a condition, and optionally, have already undergone treatment for a condition as defined herein or the one or more complications or symptoms related to said condition.
- a subject can also be one who has not been previously diagnosed as having a condition as defined herein or one or more or symptoms or complications related to said condition.
- a subject can be one who exhibits one or more risk factors for a condition, or one or more or symptoms or complications related to said condition or a subject who does not exhibit risk factors.
- the term “healthy individual” refers to an individual or group of individuals who are in a healthy state, e.g. individuals who have not shown any symptoms of the disease, have not been diagnosed with the disease and/or are not likely to develop the disease e.g. cystic fibrosis (CF) or any other disease described herein).
- CF cystic fibrosis
- Preferably said healthy individual(s) is not on medication affecting CF and has not been diagnosed with any other disease.
- the one or more healthy individuals may have a similar sex, age, and/or body mass index (BMI) as compared with the test individual.
- BMI body mass index
- Application of standard statistical methods used in medicine permits determination of normal levels of expression in healthy individuals, and significant deviations from such normal levels.
- control and “reference population” are used interchangeably.
- SFTPB Surfactant Protein B promoter fragments
- SFTPB Surfactant Protein B
- SFTPB Surfactant Protein B
- the full ⁇ length SFTPB (surfactant protein B; SFTPB) promoter has previously been defined as a sequence that could be amplified from human genomic DNA with the PCR primers ATTTGAGCTCTTCTTTCTGCTGAACCATCG (sense, SEQ ID NO: 49) and TCTTAGATCTGTCAGACAGCTCTGGGTTCC (antisense, SEQ ID NO: 50), wherein the underlines sequences align with GenBank NCBI Reference Sequence: NG_016967.1 (version 1, accessed 18 November 2022) while the additional 5’ sequences provide SacI (sense) and BglII (antisense) restriction enzyme sites.
- the forward primer binds to the sense strand from bases 4845 to 4864 in NG_016967.1.
- the reverse primer binds to the reverse complement of bases 5797 to 5816 in NG_016967.1.
- the primer pair define a 972 bp genomic fragment.
- the 972 bp genomic fragment includes all of exon 1 of the SFTPB gene (bases 5543 to 5565 of NG_016967.1), wherein the A at base 5543 is reported as the starting nucleotide of the mRNA generated by the SFTPB promoter and the initiating ATG at bases 5559 to 5561 of NG_016967.1 encodes the first methionine of pre ⁇ pro ⁇ SFTPB.
- This 972 bp genomic fragment further includes a part of intron 1 from bases 5626 to 5816 of NG_016967.1.
- references herein to a full ⁇ length SFTPB promoter refer specifically to this 972 bp genomic fragment, which is present SEQ ID NO: 2.
- the present inventors have identified and isolated functional fragments of the SFTPB (surfactant protein B) gene promoter.
- SFTPB promoter fragments generated by the inventors are able to drive cell ⁇ specific expression of transgenes in the lung parenchyma.
- the fragments of the SFTPB gene promoter are shorter than the 972 bp genomic fragment previously identified, i.e.
- the SFTPB promoter fragments of the invention may be also referred to as a core SFTPB promoters.
- the SFTPB promoter fragments of the invention surprisingly increase transgene expression compared with the full ⁇ length SFTPB promoter, and can do so in a lung parenchymal cell preferred/specific manner.
- a promoter of the invention is an SFTPB promoter fragment as described herein. Said SFTPB promoter fragment is functional, also as described herein.
- an SFTPB promoter fragment of the invention comprises or consists of a core SFTPB promoter fragment, as described herein, or a variant thereof.
- An SFTPB promoter fragment of the invention may comprise or consist of a fragment of SEQ ID NO: 2, which lack all or part of SFTPB intron 1.
- SFTPB intron 1 begins at base 5626 of NG_016967.1 (corresponding to residue 782 of SEQ ID NO: 2) and corresponds to bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention may comprise part (but not all) of the SFTPB intron 1.
- An SFTPB promoter fragment of the invention may not comprise one or more bases from bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention may comprise fewer than 180, fewer than 170, fewer than 160, fewer than 150, fewer than 140, fewer than 130, fewer than 120, fewer than 110, fewer than 100, fewer than 100, fewer than 90, fewer than 80, fewer than 70, fewer than 60, fewer than 50, fewer than 40, fewer than 30, fewer than 20, or fewer than 10 contiguous bases from 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention does not comprise bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2).
- the SFTPB promoter fragments of the invention do not comprise intron 1 of the SFTPB gene (or any portion thereof).
- an SFTPB promoter fragment of the invention comprises at least part of exon 1 of the SFTPB gene.
- SFTPB exon 1 begins at base 5543 of NG_016967.1 (corresponding to residue 699 of SEQ ID NO: 2) and corresponds to bases 5543 to 5565 of NG_016967.1 (corresponding to residues 699 to 781 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention may comprise the contiguous nucleotide sequence of bases 5543 to 5565 of NG_016967.1 (corresponding to residues 699 to 781 of SEQ ID NO: 2).
- An SFTPB promoter fragment of the invention may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 bases, preferably 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 bases, 3’ to base 5542 of NG_016967.1 (corresponding to residue 698 of SEQ ID NO: 2), and wherein the additional bases correspond to a fragment of exon 1 of the SFTPB gene.
- an SFTPB promoter fragment of the invention comprises 1 base 3’ to base 5542 of NG_016967.1, the additional base corresponds to base 5543 of NG_016967.1 (corresponding to residue 699 of SEQ ID NO: 2)
- the additional bases correspond to bases 5543 to 5544 of NG_016967.1 (corresponding to residues 699 to 700 of SEQ ID NO: 2)
- the additional bases correspond to bases 5543 to 5545 of NG_016967.1 (corresponding to residues 699 to 701 of SEQ ID NO: 2)
- an SFTPB promoter fragment of the invention comprises 4 bases 3’ to base 5542 of NG_016967.1
- the additional bases correspond to bases 5543 to 5546 of NG_016967.1 (corresponding to residues 699 to
- an SFTPB promoter fragment of the invention comprises a portion of SFTPB corresponding to SEQ ID NO: 3.
- An SFTPB promoter fragment of the invention may not comprise the SFTPB gene start codon.
- an SFTPB promoter fragment of the invention may comprise of at least part of exon 1 of the SFTPB gene, but does not comprise the SFTPB gene start codon.
- an SFTPB promoter fragment of the invention typically does not comprise the initiating ATG at bases 5559 to 5561 of NG_016967.1 (corresponding to residues 715 to 717 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention also does not comprise any of the SFTPB gene sequence 3’ of this start codon.
- An SFTPB promoter fragment of the invention may comprise fewer than 900 bases in length, fewer than 850 bases in length, fewer than 800 bases in length, fewer than 750 bases in length, fewer than 700 bases in length, fewer than 650 bases in length, fewer than 640 bases in length, fewer than 630 bases in length, fewer than 620 bases in length, fewer than 610 bases in length or fewer than 600 bases in length.
- an SFTPB promoter fragment of the invention comprises fewer than 800 bases in length.
- a particularly preferred SFTPB promoter fragment of the invention comprises fewer than 700 bases in length.
- An SFTPB promoter fragment of the invention may consist of fewer than 900 bases in length, fewer than 850 bases in length, fewer than 800 bases in length, fewer than 750 bases in length, fewer than 700 bases in length, fewer than 650 bases in length, fewer than 640 bases in length, fewer than 630 bases in length, fewer than 620 bases in length, fewer than 610 bases in length or fewer than 600 bases in length.
- an SFTPB promoter fragment of the invention consists of fewer than 800 bases in length.
- a particularly preferred SFTPB promoter fragment of the invention may consist of fewer than 700 bases in length.
- the exemplified SFTPB promoter fragment of the invention consists of 630 or 635 bases in length.
- An SFTPB promoter fragment of the invention may comprise from 600 bases to 800 bases in length, from 600 bases to 750 bases in length, from 600 bases to 700 bases in length, or from 600 bases to 650 bases in length.
- An SFTPB promoter fragment of the invention may consist of from 600 bases to 800 bases in length, from 600 bases to 750 bases in length, from 600 bases to 700 bases in length, or from 600 bases to 650 bases in length.
- An SFTPB promoter fragment of the invention may comprise from 630 bases to 800 bases in length, from 635 bases to 800 bases in length, from 630 bases to 750 bases in length, from 635 bases to 750 bases in length, from 630 bases to 700 bases in length, from 635 bases to 700 bases in length, from 630 bases to 650 bases in length, or from 635 bases to 650 bases in length.
- An SFTPB promoter fragment of the invention may consist of from 630 bases to 800 bases in length, from 635 bases to 800 bases in length, from 630 bases to 750 bases in length, from 635 bases to 750 bases in length, from 630 bases to 700 bases in length, from 635 bases to 700 bases in length, from 630 bases to 650 bases in length, or from 635 bases to 650 bases in length.
- An SFTPB promoter fragment of the invention may be a fragment of a mammalian or avian, preferably a mammalian SFTPB promoter, i.e. an SFTPB promoter fragment of the invention may be a mammalian or avian SFTPB promoter fragment.
- a mammalian SFTPB promoter fragment may be a human, non ⁇ human primate, mouse, rat, dog, cat, horse, or cow SFTPB promoter fragment.
- the SFTPB promoter fragment is a fragment of a human SFTPB promoter, i.e. preferably the SFTPB promoter fragment of the invention is a human SFTPB promoter fragment.
- an SFTPB promoter fragment of the invention may comprise any functional fragment of SEQ ID NO: 2.
- such an SFTPB promoter fragment of the invention is of a length as described herein.
- An SFTPB promoter fragment of the invention may comprise or consist of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8%sequence identity to a fragment of SEQ ID NO: 2, wherein the promoter retains the function of the SFTPB gene promoter, as defined herein.
- an SFTPB promoter fragment of the invention is of a length as described herein.
- such an SFTPB promoter fragment of the invention may comprise any additional feature (e.g.
- An SFTPB promoter fragment of the invention may comprise a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention may comprise a sequence having at least 90% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
- An SFTPB promoter fragment of the invention may comprise a sequence having at least at least 95% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
- An SFTPB promoter fragment of the invention may comprise the sequence of bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
- An SFTPB promoter fragment of the invention may consist of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity or more to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
- An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 90% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
- An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 95% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
- An SFTPB promoter fragment of the invention may consist of the sequence of bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention derived from SEQ ID NO: 2, such as those described above, may comprise a substitution at residue 683 of SEQ ID NO: 2 (corresponding to base 5527 of NG_016967.1).
- an SFTPB promoter fragment of the invention derived from SEQ ID NO: 2, such as those described above, may comprise substitution of an alanine residue by a cytosine at residue 683 of SEQ ID NO: 2 (corresponding to base 5527 of NG_016967.1), in other words, may comprise an A683C substitution.
- An exemplified SFTPB promoter fragment of the invention is SEQ ID NO: 1.
- an SFTPB promoter fragment which comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to SEQ ID NO: 1.
- an SFTPB promoter fragment of the invention may comprise a sequence having at least 90% identity to SEQ ID NO: 1.
- An SFTPB promoter fragment of the invention may comprise a sequence having at least at least 95% identity to SEQ ID NO: 1.
- an SFTPB promoter fragment of the invention may comprise the sequence of SEQ ID NO: 1.
- the present invention provides an SFTPB promoter fragment which consists of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to SEQ ID NO: 1.
- an SFTPB promoter fragment of the invention may consist of a sequence having at least 90% identity to SEQ ID NO: 1.
- An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 95% identity to SEQ ID NO: 1.
- an SFTPB promoter fragment of the invention may consist of the sequence of SEQ ID NO: 1.
- An SFTPB promoter fragment of the invention may further comprise up to 10 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bases), up to 15 bases, up to 20 bases, up to 25 bases or up to 25 bases at the 5’ end.
- additional bases may preferably be present when said SFTPB promoter fragment (i) comprises a sequence having at least 70% identity to SEQ ID NO: 1; and/or (ii) comprises a sequence having at least 70% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2); as described herein.
- additional bases are present at the 5’ end of an SFTPB promoter fragment of the invention
- said additional bases may be bases which correspond to the corresponding number of bases 5’ to base 4925 of NG_016967.1 (corresponding to residue 81 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention comprises 10 bases 5’ to base 4925 of NG_016967.1
- the additional base corresponds to bases 4915 to 4924 of NG_016967.1 (corresponding to residues 71 to 80 of SEQ ID NO: 2)
- the additional bases correspond to bases 4905 to 4924 of NG_016967.1 (corresponding to residues 61 to 80 of SEQ ID NO: 2), and so on.
- an SFTPB promoter fragment of the invention may further comprise up to 20 bases at the 5’ end, and particularly preferably, the up to 20 additional bases correspond to bases 5’ of residue 4925 of NG_016967.1 (corresponding to residue 81 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention may further comprise up to 10 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bases) bases at the 3’ end.
- Such additional bases may preferably be present when said SFTPB promoter fragment (i) comprises a sequence having at least 70% identity to SEQ ID NO: 1; and/or (ii) comprises a sequence having at least 70% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2); as described herein.
- said additional bases may be bases which correspond to the corresponding number of bases 3’ to base 5554 of NG_016967.1 (corresponding to residue 710 of SEQ ID NO: 2).
- an SFTPB promoter fragment of the invention comprises 10 bases 3’ to base 5554 of NG_016967.1
- the additional base corresponds to bases 5555 to 5564 of NG_016967.1 (corresponding to residues 711 to 720 of SEQ ID NO: 2)
- the additional bases correspond to bases 5555 to 5559 of NG_016967.1 (corresponding to residues 711 to 715 of SEQ ID NO: 2)
- the additional bases correspond to bases 5555 to 5558 of NG_016967.1 (corresponding to residues 711 to 714 of SEQ ID NO: 2), and so on.
- an SFTPB promoter fragment of the invention may further comprise up to 4 bases at the 3’ end, and particularly preferably, the up to 4 additional bases correspond to bases 3’ of residue 5554 of NG_016967.1 (corresponding to residue 710 of SEQ ID NO: 2).
- An SFTPB promoter fragment of the invention may comprise additional sequences at the 5’ and/or 3’ end to facilitate molecular biology applications of said SFTPB promoter fragment.
- one or more restriction enzyme site may be added at the 5’ and/or 3’ end of the sequence. Such restriction enzyme sites may be used to facilitate cloning of the SFTPB promoter fragment into a non ⁇ viral vector (e.g.
- restriction enzyme sites are present at both the 5’ and 3’ end of the SFTPB promoter fragment, each restriction enzyme site may be selected independently, and thus the 5’ and 3’ sites may be the same or different.
- restriction enzyme sites include NheI and BgIII.
- a SFTPB promoter fragment of the invention may have a 5’ BgIII restriction enzyme site (5’AGATCT3’) and/or a 3’ NheI restriction enzyme site (5’GCTAGC3’).
- Enhancer As exemplified herein, the inventors have also modified the SFTPB promoter fragments of the invention to further increase transgene expression levels and/or to increase promoter activity, particularly in the lung parenchyma.
- the inventors modified the SFTPB promoter fragment to include an enhancer sequence.
- the invention further provides an SFTPB promoter fragment which further comprises an enhancer. All disclosure herein to SFTPB promoter fragments of the invention applies equally and without reservation to SFTPB promoter fragment which further comprise an enhancer.
- An enhancer is a cis ⁇ acting DNA sequence which can increase gene transcription.
- An enhancer of the invention may be from about 150 to about 900 bp in length, such as from about 200 to about 900 bp, from about 200 to about 800 bp, from about 300 to about 700 bp, from about 400 to about 800 bp, or from about 300 to about 600 bp in length.
- An enhancer of the invention may be linked to an SFTPB promoter fragments of the invention by a linker.
- Said linker is typically a short DNA sequence, which may be from about 1 to about 50 bp in length, such as from about 1 to about 20 bp, from about 1 to about 10 bp, from about 5 to about 20 bp in length, or from about 5 to about 10 bp in length.
- a linker may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp, particularly 6 bp, in length.
- a non ⁇ limiting example of a linker sequence is given in SEQ ID NO: 51.
- Said linker may comprise or consist of one or more restriction enzyme site, non ⁇ limiting examples of which are described herein.
- a promoter/enhancer combination of the invention may be joined by a linker which comprises or consists of a NheI restriction site and/or a BgIII restriction site, preferably a BgIII restriction site.
- An enhancer of the invention may comprise additional sequences at the 5’ and/or 3’ end.
- one or more restriction enzyme site may be added at the 5’ and/or 3’ end of the enhancer.
- each restriction enzyme site may be selected independently, and thus the 5’ and 3’ sites may be the same or different.
- Said 5’ or 3’ restriction site may comprise all or part of a linker joining the enhancer to the SFTPB promoter fragment of the invention.
- the enhancer may be (i) 5’ to an SFTPB promoter fragment of the invention, or (ii) 3’ to an SFTPB promoter fragment of the invention.
- the enhancer is 5’ to an SFTPB promoter fragment of the invention.
- the (5’ or 3’, preferably 5’) enhancer may be in (i) the forward, or (ii) the reverse orientation.
- the enhancer is in the forward orientation.
- an SFTPB promoter fragment further comprising an enhancer sequence as defined herein has the potential to provide an even greater increase in transgene expression compared with the full ⁇ length SFTPB promoter (i.e. transgene expression increases full ⁇ length SFTPB promoter ⁇ SFTPB promoter fragment of the invention ⁇ SFTPB promoter of the invention further comprising an enhancer).
- Higher expression may be defined as greater than 105%, greater than 110%, greater than 120%, greater than 130%, greater than 140%, greater than 150%, greater than 175%, greater than 200%, greater than 225%, greater than 250%, greater than 300%, greater than 350%, greater than 400% or more of the transgene expression compared with a suitable control, such as the SFTPB promoter fragment alone, or the full ⁇ length SFTPB promoter.
- a suitable control such as the SFTPB promoter fragment alone, or the full ⁇ length SFTPB promoter.
- expression of a transgene by an SFTPB promoter of the invention further comprising an enhancer may be greater than 105%, greater than 110%, greater than 120%, greater than 130%, greater than 140%, greater than 150%, greater than 175%, greater than 200%, greater than 225%, or greater than 250% of the transgene expression when using the SFTPB promoter fragment alone.
- Transgene expression may be quantified at the nucleic acid and/or protein level, and can be quantified by any suitable standard technique known to the person skilled in the art, for example, by real ⁇ time reverse transcription polymerase chain reaction (RT ⁇ qPCR), Western blotting and enzyme ⁇ linked immunosorbent assay or ELISA.
- the inventors also surprisingly found that the (cell ⁇ specific) expression driven by a SFTPB promoter fragment comprising an enhancer sequence as defined herein is comparable to gene expression levels when using ubiquitously used strong promoters (e.g. CMV, hCEF).
- Comparable expression may be defined as at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% of gene expression relative to expression of the same transgene using the same strong promoter (e.g. CMV or hCEF promoter [such as SEQ ID NOs: 6 and 5 respectively]).
- Expression of the transgene using the SFTPB promoter fragment may be higher than expression using a strong promoter (e.g. CMV or hCEF) promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more relative to expression of the same transgene using the same strong promoter (e.g. CMV or hCEF promoter).
- a strong promoter e.g. CMV or hCEF promoter
- the SFTPB promoter fragment comprises a SFTPB promoter fragment as defined above operably linked to an enhancer
- said enhancer may preferably be of viral origin, a lung ⁇ preferred enhancer, a lung ⁇ parenchyma ⁇ preferred enhancer, a pneumocyte ⁇ preferred enhancer, an ATII and club cell ⁇ preferred enhancer, or an ATII cell preferred enhancer, a lung ⁇ specific enhancer, a lung ⁇ parenchyma ⁇ specific enhancer, a pneumocyte ⁇ specific enhancer, an ATII and club cell ⁇ specific enhancer, or an ATII cell specific enhancer.
- the terms “preferred” and “specific” are defined herein.
- the term “preferred” may alternatively or additionally be defined as higher expression in the lung relative to expression in the nose (particularly the nasal cavity), and thus give rise to an increased ratio of lung expression: nose expression, as described herein.
- Expression in the lung and nasal cavity can be determined using an in vivo luciferase reporter assay (e.g., wherein vectors comprising nucleic acid cassettes of the invention are administered to mice, and the relative bioluminescence of the lungs and nasal cavity is quantified).
- the enhancer may be selected from a hB ⁇ actin enhancer, a SLC34A2 (Sodium ⁇ dependent phosphate transport protein 2B) enhancer, a VEGFA (Vascular Endothelial Growth Factor A) enhancer, a CMV (Cytomegalovirus) enhancer, an SV40 (simian virus 40) enhancer, an ELF3 (E74 like ETS transcription factor 3) enhancer, an SFTPC (surfactant protein C) enhancer, an SFTPB enhancer, or a LMO7 (LIM domain 7) enhancer.
- a hB ⁇ actin enhancer a SLC34A2 (Sodium ⁇ dependent phosphate transport protein 2B) enhancer, a VEGFA (Vascular Endothelial Growth Factor A) enhancer, a CMV (Cytomegalovirus) enhancer, an SV40 (simian virus 40) enhancer, an ELF3 (E74 like ETS transcription factor 3) enhancer, an SFTPC
- the enhancer is selected from a SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer.
- the enhancer is a SLC34A2 enhancer.
- the enhancer is a VEGFA enhancer.
- the enhancer is a CMV enhancer.
- the enhancer is a SV40 enhancer.
- the enhancer is an ELF3 enhancer.
- Combinations of enhancers, typically those identified herein, and combinations of one or more preferred enhancer described herein, may be used according to the present invention.
- the CMV enhancer may be in the forwards orientation.
- the CMV enhancer is in the reverse orientation.
- the ELF3 enhancer may be in the forwards orientation.
- the SV40 enhancer may be in the forwards orientation.
- the SV40 enhancer may be in the reverse orientation.
- the VEGFA enhancer may be in the forwards orientation.
- the VEGFA enhancer may be in the reverse orientation.
- the SLC34A2 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
- the SLC34A2 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SQE ID NO: 31.
- the SLC34A2 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
- the SLC34A2 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
- the SLC34A2 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
- the SLC34A2 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
- the VEGFA enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
- the VEGFA enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
- the VEGFA enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
- the VEGFA enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
- the VEGFA enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
- the VEGFA enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
- the CMV enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11.
- the CMV enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11.
- the CMV enhancer may comprise the nucleotide sequence of SEQ ID NO: 5.
- the CMV enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11.
- the CMV enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11.
- the CMV enhancer may consist of the nucleotide sequence of SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11. This exemplary CMV enhancer sequence is CpG ⁇ free.
- the SV40 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 34 or 35.
- the SV40 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 34 or 35.
- the SV40 enhancer may comprise the nucleotide sequence of SEQ ID NO: 34 or 35.
- the SV40 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 34 or 35.
- the SV40 enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or 35.
- the SV40 enhancer may consist of the nucleotide sequence of SEQ ID NO: 34 or 35.
- the ELF3 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 13 to 20.
- the ELF3 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 13 to 20.
- the ELF3 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 13 to 20.
- the ELF3 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 13 to 20.
- the ELF3 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 13 to 20.
- the ELF3 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 13 to 20.
- the actin enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 21 or 22.
- the actin enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 21 or 22.
- the actin enhancer may comprise the nucleotide sequence of SEQ ID NO: 21 or 22.
- the actin enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 21 or 22.
- the actin enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 21 or 22.
- the actin enhancer may consist of the nucleotide sequence of SEQ ID NO: 21 or 22.
- the LMO7 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 23 to 25.
- the LMO7 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 23 to 25.
- the LMO7 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 23 to 25.
- the LMO7 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 23 to 25.
- the LMO7 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 23 to 25.
- the LMO7 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 23 to 25.
- the SFTPC enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 27 or 28.
- the SFTPC enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 27 or 28.
- the SFTPC enhancer may comprise the nucleotide sequence of SEQ ID NO: 27 or 28.
- the SFTPC enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 27 or 28.
- the SFTPC enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 27 or 28.
- the SFTPC enhancer may consist of the nucleotide sequence of SEQ ID NO: 27 or 28.
- the SFTPB enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 26.
- the SFTPB enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 26.
- the SFTPB enhancer may comprise the nucleotide sequence of SEQ ID NO: 26.
- the SFTPB enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 26.
- the SFTPB enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 26.
- the SFTPB enhancer may consist of the nucleotide sequence of SEQ ID NO: 26. Any SFTPB promoter fragment of the invention may be combined with any enhancer of the invention.
- any preferred SFTPB promoter fragment of the invention may be combined with any preferred enhancer (e.g. an SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer).
- any preferred enhancer e.g. an SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer.
- a preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).
- a particularly preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).
- a particularly preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 47.
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 47.
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 47.
- Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 47.
- a preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).
- a particularly preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83).
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83).
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83).
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83).
- a particularly preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 48.
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 48.
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 48.
- Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 48.
- a preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).
- a particularly preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81).
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81).
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81).
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81).
- a particularly preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 46.
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 46.
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 46.
- Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 46.
- a preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).
- Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).
- Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).
- Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).
- a preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).
- Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).
- Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).
- Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).
- a preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%,a t least 98%, at least 99% identity or more to any one of SEQ ID NOs: 41 to 48.
- a preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to any one of SEQ ID NOs: 41 to 48.
- a preferred SFTPB promoter fragment further comprising an enhancer may comprise a nucleic acid sequence of any one of SEQ ID NOs: 41 to 48.
- a preferred SFTPB promoter fragment further comprising an enhancer may consist of a nucleic acid sequence of any one of SEQ ID NOs: 41 to 48.
- a particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%,a t least 98%, at least 99% identity or more to any one of SEQ ID NOs: 46 to 48.
- a particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to any one of SEQ ID NOs: 46 to 48.
- a particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise a nucleic acid sequence of any one of SEQ ID NOs: 46 to 48.
- a particularly preferred SFTPB promoter fragment further comprising an enhancer may consist of a nucleic acid sequence of any one of SEQ ID NOs: 46 to 48.
- Lung parenchyma and specific/preferred expression therein The respiratory system can be divided into airways and lung parenchyma.
- the airways consist of the bronchus, which bifurcates off the trachea and divides into bronchioles and then further into alveoli.
- the parenchyma is responsible for gas exchange and includes the alveoli, alveolar ducts, and terminal and respiratory bronchioles.
- the most prominent structure in the lung parenchyma is the alveolus. Two types of epithelial cell line the alveolus.
- ATI cells exhibit a broad, flattened morphology and cover around 95% of the surface area, whilst the cuboidal alveolar type II cells (ATII cells) line the remainder of the alveolus.
- ATI cells provide a gas exchange interface with the underlying endothelium, whereas ATII cells serve as both progenitors of ATI cells and also play a critical role in maintaining the homeostasis of the alveolus. The latter role is fulfilled by the secretion of surfactant proteins from specialised organelles within ATII cells, so ⁇ called ‘lamellar bodies’, into the alveolar space.
- ATII cells are the only epithelial cell of the lung which synthesise and release all four surfactant proteins A, B, C and D, with surfactant protein C being unique to the ATII cell.
- ATII cells have the following functions: (1) the transepithelial movement of water and ions regulating the volume of the alveolar surface liquid (ASL) preventing alveoli flooding, (2) the expression of immunomodulatory proteins necessary for host defence and the regulation of innate immunity and (3) the regeneration of alveolar epithelium after injury.
- ASL alveolar surface liquid
- surfactant proteins A, B and D are also synthesised by club cells (previously named Clara Cells) founds in the terminal and respiratory bronchioles of humans.
- Club cells are non ⁇ ciliated epithelial cells found mainly in bronchioles as well as basal cells found in large airways. They have been ascribed several protective roles, including airway repair after injury, secretion of anti ⁇ inflammatory and immunomodulatory proteins, and detoxification.
- ATI dysfunction, ATII dysfunction and/or club cell dysfunction or dropout is associated with the pathogenesis of various parenchymal lung diseases. Accordingly, the lung parenchyma may be targeted for treating genetic diseases such as surfactant deficiencies and interstitial lung disease.
- the promoters of the invention drive transgene expression in the lung parenchyma, typically one or more cell type of the lung parenchyma.
- the promoters of the invention drive transgene expression in (i) ATI cells, (ii) ATII cells, (iii) club cells, (iv) ATI cells and ATII cells, (v) ATI cells and club cells, (vi) ATII cells and club cells, or (vii) ATI cells, ATII cells and club cells.
- a SFTPB promoter fragment of the invention is functional. As used herein, the term “functional” may mean that an SFTPB promoter fragment of the invention retains the functionality of the full ⁇ length SFTPB promoter.
- a promoter of the invention may be defined as being capable of expressing a gene of interest (e.g., a transgene) in a cell type which, in a healthy subject, would express SFTPB.
- a gene of interest e.g., a transgene
- the an SFTPB promoter fragment may be used to express a transgene in ATII cells, club cells and/or ATI cells.
- an SFTPB promoter fragment of the invention may be defined as functional as it is capable of preferentially or specifically expressing a gene of interest (e.g., a transgene) in a tissue ⁇ type or cell ⁇ type which, in a healthy subject, would express SP ⁇ B.
- tan SFTPB promoter fragment of the invention may be a lung ⁇ parenchyma preferred promoter.
- An SFTPB promoter fragment of the invention may be a lung ⁇ parenchyma specific promoter.
- An SFTPB promoter fragment of the invention may be a pneumocyte ⁇ preferred promoter, whereby a pneumocyte is defined as any of the specialized cells of the alveoli of the lungs.
- An SFTPB promoter fragment of the invention may be a pneumocyte ⁇ specific promoter.
- An SFTPB promoter fragment of the invention may be an ATII cell ⁇ preferred promoter, a club cell ⁇ preferred promoter and/or an ATI cell ⁇ preferred promoter.
- An SFTPB promoter fragment of the invention may be an ATII cell ⁇ specific promoter, a club cell ⁇ specific expression and/or an ATI cell ⁇ specific promoter.
- An SFTPB promoter fragment of the invention may preferably be an ATII cell ⁇ preferred promoter.
- An SFTPB promoter fragment of the invention may preferably be an ATII cell ⁇ specific promoter.
- Tissue or cell preferred expression may be defined as expression that is higher in said tissue or cell than other tissue or cell types.
- lung ⁇ parenchyma preferred expression may be defined as expression that is significantly higher in the lung ⁇ parenchyma (or one or more cell type therein, as described above) than expression in one or more of: the brain, the eye, the endocrine tissues, the proximal digestive tract, the gastrointestinal tract, liver and gall bladder, pancreas, the kidney and/or urinary bladder, male tissues (i.e., the testis, epididymis, prostate and/or seminal vesicle) and female tissues (i.e., the vagina, breast, cervix, endometrium, fallopian tube, ovary and placenta), muscle tissues, connective and soft tissues, skin, bone marrow and lymphoid tissue.
- male tissues i.e., the testis, epididymis, prostate and/or seminal vesicle
- female tissues i.e., the vagina, breast, cervix, endometrium, fallopian tube, ovary and place
- Lung ⁇ parenchyma preferred expression may be defined as expression that is at least about 5 times greater, at least about 10 times greater, at least about 20 times greater, at least about 50 times greater or at least about 100 or more times greater in the lung parenchyma than one or more of the above reference tissue types, especially the gastrointestinal tract and/or the brain.
- Tissue or cell specific expression may be defined as expression that is at least about 10 times greater, at least about 20 times greater, at least about 50 times greater or at least about 100 or more times greater in said tissue or cell than any other tissue or cell types.
- Preferential and/or specific expression may be assessed at the level of RNA and/or protein expression, preferably protein expression. Expression can be measured by any suitable standard technique known to the person skilled in the art.
- RNA expression levels can be measured by quantitative real ⁇ time PCR.
- Protein expression can be measured by western blotting or immunohistochemistry.
- restricting the expression of transgenes to cells expressing endogenous SFTPB is expected to reduce the effects of off ⁇ target gene expression, overexpression (e.g., toxicity/ER stress/UPR, etc) and/or reduce immune responses.
- the SFTPB promoters of the invention as a consequence of their preferential and/or specific expression in the lung parenchyma, or one or more cell type thereof, have potential clinical benefits as a result of these advantageous properties.
- an SFTPB promoter fragment of the invention may reduce the effects of off ⁇ target gene expression, overexpression, and/or reduce immune responses. Additionally, or alternatively, compared to a hCEF promoter of SEQ ID NO: 5and/or a CMV promoter of SEQ ID NO: 6, an SFTPB promoter fragment of the invention may preferentially or specifically express a transgene in the lung parenchyma, pneumocytes, such as ATII cells, ATI cells, and/or club cells.
- an SFTPB promoter fragment of the invention may increase transgene expression compared with a full ⁇ length SFTPB promoter, such as that described herein.
- an SFTPB promoter fragment of the invention increases transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)).
- An SFTPB promoter fragment of the invention increases transgene expression in lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2 ⁇ fold, at least about 2.5 ⁇ fold, at least about 3 ⁇ fold, at least about 4 ⁇ fold, at least about 5 ⁇ fold, at least about 7.5 ⁇ fold or more.
- the increase in transgene expression by an SFTPB promoter fragment of the invention may be quantified compared with a suitable control, preferably compared with a full ⁇ length SFTPB promoter, such as that described herein.
- an SFTPB promoter fragment of the invention increases transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with expression of the same transgene in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with a full ⁇ length SFTPB promoter, such as that described herein.
- An SFTPB promoter fragment of the invention increases transgene expression in lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2 ⁇ fold, at least about 2.5 ⁇ fold, at least about 3 ⁇ fold, at least about 4 ⁇ fold, at least about 5 ⁇ fold, at least about 7.5 ⁇ fold or more compared with expression of the same transgene in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with a full ⁇ length SFTPB promoter, such as that described herein.
- the SFTPB promoter fragment of the invention increases transgene expression in ATII cells, and optionally one or more additional lung parenchymal cell type as described herein. Again, this increase in expression is preferably compared with expression of the same transgene in the same cell type(s) by a full ⁇ length SFTPB promoter, such as that described herein.
- an SFTPB promoter fragment of the invention may preferentially drive transgene expression in the lung (particularly the lung parenchyma or one or more cell type thereof) compared with transgene expression in the nose (or cells thereof).
- an SFTPB promoter fragment of the invention may preferentially increase transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with transgene expression in the nose (or cells thereof).
- An SFTPB promoter fragment of the invention may preferentially increase transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2 ⁇ fold, at least about 2.5 ⁇ fold, at least about 3 ⁇ fold, at least about 4 ⁇ fold, at least about 5 ⁇ fold, at least about 7.5 ⁇ fold or more compared with transgene expression in the nose (or cells thereof).
- the ratio of lung expression: nose expression by an SFTPB promoter fragment of the invention may be at least about 2:1, such as at least about 2.5:1, at least about 3:1, at least about 4:1, at least about 5:1 or more.
- any disclosure herein in relation to promoters (i.e. SFTPB promoter fragments) of the invention applies equally and without reservation to other aspects of the invention comprising, or relating to, said promoters.
- any disclosure herein in relation to promoters (i.e. SFTPB promoter fragments) of the invention applies equally and without reservation to promoter/enhancer combinations, viral and non ⁇ viral vectors of the invention.
- promoter/enhancer combinations, viral and non ⁇ viral vectors of the invention drive transgene expression in the lung parenchyma, typically one or more cell type of the lung parenchyma.
- the promoters i.e. SFTPB promoter fragments
- promoter/enhancer combinations, viral and non ⁇ viral vectors of the invention drive transgene expression in (i) ATI cells, (ii) ATII cells, (iii) club cells, (iv) ATI cells and ATII cells, (v) ATI cells and club cells, (vi) ATII cells and club cells, or (vii) ATI cells, ATII cells and club cells.
- Nucleic Acid Cassettes The invention also provides a nucleic acid cassette.
- the present invention provides a nucleic acid cassette comprising (a) an SFTPB promoter fragment; and (b) a transgene.
- a transgene may be defined as anucleic acid sequence encoding a therapeutic protein.Thus, the terms “nucleic acid sequence encoding a therapeutic protein” and the term “transgene” may be used interchangeably.
- an SFTPB promoter fragment may increase expression of the transgene by the lung parenchyma (e.g. ATII cells), as defined herein.
- the increase in expression of a transgene by an SFTPB promoter fragment of the invention may be as defined herein, including disclosure of increasing transgene expression using SFTPB promoter fragments and/or SFTPB promoter fragments combined with an enhancer, as described above.
- any disclosure herein in relation to increasing transgene expression using an SFTPB promoter fragment and/or SFTPB promoter fragment combined with an enhancer of the invention applies equally and without reservation to nucleic acid cassettes of the invention.
- the increase in expression of a transgene by an SFTPB promoter fragment of the invention is an increase of at least about 50%, at least about 60%, at least about 70%, at least about 80% or more.
- the increase in expression of a transgene by an SFTPB promoter fragment of the invention is an increase of at least about 50%.
- an SFTPB promoter fragment of the invention may increase expression of the transgene by the lung parenchyma (e.g.
- the SFTPB promoter fragment of the invention may increase expression of the transgene by the lung parenchyma (e.g. ATII cells) relative to expression using the full ⁇ length SFTPB gene promoter (e.g., the 972 bp genomic fragment defined above).
- expression of the transgene by the lung parenchyma (e.g. ATII cells) using the SFTPB promoter fragment in a nucleic acid of the invention may be comparable to the expression of the transgene using a ubiquitously used strong promoter (e.g. CMV or hCEF).
- the expression of the transgene using the SFTPB promoter fragment may be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% of the expression using a CMV promoter, as described herein.
- expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may be higher than expression using a CMV promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more relative to expression of the same transgene using a CMV promoter.
- expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% of the expression using a hCEF promoter.
- expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may be higher than expression using a hCEF promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more relative to expression of the same transgene using a CMV promoter.
- a nucleic acid cassette or vector of the invention enables long ⁇ term transgene expression, resulting in long ⁇ term expression of a transgene, which offers clinical benefits for the expression of therapeutic proteins in patients.
- Long ⁇ term expression means expression of a transgene, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more.
- long ⁇ term expression means expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
- the long ⁇ term expression is typically accompanied by long ⁇ term secretion or long ⁇ term membrane insertion of the (therapeutic) protein encoded by the transgene, depending on whether the (therapeutic) protein is a secreted protein (e.g. SFTPB) or a membrane protein.
- Long ⁇ term secretion means secretion of a (therapeutic) protein encoded by a transgene of the invention, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more.
- long ⁇ term secretion means secretion for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
- Long ⁇ term membrane insertion means that a (therapeutic) protein encoded by a transgene of the invention is inserted into and present in the cell membrane, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more.
- long ⁇ term expression means membrane insertion for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
- a nucleic acid cassette or vector of the invention may drive (increased) long ⁇ lasting expression of a transgene, and/or (preferably and) expression and secretion/membrane insertion of a (therapeutic) protein encoded by said transgene in one or more cell type of the lung parenchyma, as described herein, in vivo in a patient.
- a nucleic acid cassette or vector of the invention drives expression of a transgene, and/or (preferably and) expression and secretion/membrane insertion of a (therapeutic) protein encoded by a transgene of the invention in one or more cell / cell type of the lung parenchyma, as described herein, for at least 45 days, more preferably at least 90 days.
- the nucleic acid of the nucleic acid cassette may be as defined herein.
- the nucleic acid cassette comprise DNA or RNA.
- the nucleic acid cassette is DNA.
- a nucleic acid cassette of the invention may optionally be codon optimised for expression in a particular cell type, for example, eukaryotic cells (e.g.
- codon optimised refers to the replacement of at least one codon within a base polynucleotide sequence with a codon that is preferentially used by the host organism in which the polynucleotide is to be expressed. Typically, the most frequently used codons in the host organism are used in the codon ⁇ optimised polynucleotide sequence. Methods of codon optimisation are well known in the art. It will be understood by a skilled person that numerous different polynucleotides can encode the same polypeptide as a result of the degeneracy of the genetic code.
- a nucleic acid cassette that comprises or consists of the SFTPB promoter fragment and transgene (e.g., encoding a therapeutic protein) of the invention includes all polynucleotide sequences that are degenerate versions of each other and that encode the same amino acid sequence.
- a nucleic acid cassette of the invention preferably comprises an SFTPB promoter fragment comprising an SFTPB promoter fragment operably linked to an enhancer sequence, as described herein.
- a nucleic acid cassette of the invention typically comprises a SFTPB promoter fragment comprising a SFTPB promoter fragment and an enhancer sequence, wherein the SFTPB promoter fragment is operably linked to a nucleic acid sequence comprising or consisting of a transgene (e.g., encoding a therapeutic protein).
- a transgene e.g., encoding a therapeutic protein.
- operably linked it is meant that the SFTPB promoter fragment is configured to express the transgene (e.g., encoding the therapeutic protein).
- the transgene encoding the therapeutic protein may also be linked to a suitable terminator sequence. Suitable terminator sequences are well known in the art.
- the promoter included in the nucleic acid cassettes and vectors of the invention may be specifically selected and/or modified to further refine regulation of expression of the therapeutic gene.
- an SFTPB promoter fragment of the invention may be modified to reduce the number of CpG dinucleotides, or to render the SFTPB promoter fragment CpG ⁇ free.
- a number of (CpG ⁇ free) promoters, and methods for the generation of CpG ⁇ free promoters which are suitable for use in the present invention are described in Pringle et al. (J. Mol. Med. Berl. 2012, 90(12): 1487 ⁇ 96), which is herein incorporated by reference in its entirety.
- the nucleic acid cassettes and vectors of the invention comprise an SFTPB promoter fragment having low or no CpG dinucleotide content.
- Low CpG dinucleotide content may be defined as 10 CpG dinucleotides or less, preferably 5 CpG dinucleotides or less, such as 5, 4, 3, 2 or 1 CpG dinucleotides.
- An SFTPB promoter fragment may have some or all CG dinucleotides replaced with any one of AG, TG or GT.
- the absence (or reduction) of CpG dinucleotides further improves the performance of some nucleic acid cassettes and vectors of the invention, particularly lentiviral (e.g. SIV) vectors of the invention and in particular in situations where it is not desired to induce an immune response against an expressed antigen or an inflammatory response against the delivered expression construct.
- the elimination or reduction of CpG dinucleotides reduces the occurrence of flu ⁇ like symptoms and inflammation which may result from administration of constructs, particularly when administered to the airways.
- the nucleic acid cassettes and vectors of the invention may be modified to allow shut down of gene expression. Standard techniques for modifying the vector in this way are known in the art. As a non ⁇ limiting example, Tet ⁇ responsive promoters are widely used.
- the nucleic acid cassette of the invention (or a vector comprising said cassette) may have an intron positioned between the promoter and the transgene.
- suitable introns are found for example, in UK Application No. 2213936.4, which is herein incorporated by reference in its entirety.
- nucleic acid of the invention is present in a non ⁇ viral vector (e.g. plasmid), the presence of at least one intron between the SFTPB promoter fragment and the transgene may be preferred, for example an intron as described in UK Application No. 2213936.4.
- the nucleic acid cassettes and vectors of the invention may include at least one part of a vector, in particular, regulatory elements.
- the promoter within a nucleic acid cassette of the invention may be used to express more than one polypeptide, including one or more therapeutic protein.
- the nucleic acid cassette may comprise a nucleic acid sequence which, when transcribed, gives rise to multiple polypeptides, for instance a transcript may contain multiple open reading frames (ORFs) and also one or more Internal Ribosome Entry Sites (IRES) to allow translation of ORFs after the first ORF.
- a transcript may be polycistronic, i.e. it may be translated to give a polypeptide which is subsequently cleaved to give a plurality of polypeptides.
- a nucleic acid cassette of the invention may comprise multiple promoters, including multiple SFTPB promoter fragments of the invention, or multiple copies of any specific an SFTPB promoter fragment of the invention, and hence give rise to a plurality of transcripts and hence a plurality of polypeptides, including a plurality of therapeutic proteins.
- Nucleic acid cassettes may, for instance, express one, two, three, four or more polypeptides via a promoter or promoters, including one or more SFTPB promoter fragment of the invention.
- a nucleic acid cassette may comprise one or more translation initiation sequence (TIS).
- Translation initiation plays an important role in mRNA translation, canonically a methionyl tRNA unique for initiation (Met ⁇ tRNAi) identifies the AUG start codon and triggers the downstream translation process.
- Non ⁇ canonical start codons e.g. CUG for valyl ⁇ tRNA
- the nucleic acid cassettes of the present invention may comprise at least one termination signal.
- a “termination signal” or “terminator” is comprised of the DNA sequences involved in specific termination of an RNA transcript by an RNA polymerase. Thus, a termination signal that ends the production of an RNA transcript is contemplated according to the present invention.
- a terminator may be necessary in vivo to achieve desirable message levels.
- a terminator region may also comprise specific DNA sequences that permit site ⁇ specific cleavage of the new transcript so as to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch of about 200 A residues (polyA) to the 3’ end of the transcript. RNA molecules modified with this polyA tail appear to more stable and are translated more efficiently.
- a terminator typically comprises a signal for the cleavage of the RNA, and it is preferred that the terminator signal promotes polyadenylation of the message.
- the terminator and/or polyadenylation site elements can serve to enhance message levels and to minimize read through from the cassette into other sequences.
- Terminators contemplated for use in the invention include any known terminator of transcription described herein or known to one of ordinary skill in the art, including but not limited to, for example, the termination sequences of genes, such as for example the bovine growth hormone terminator or viral termination sequences, such as for example the SV40 terminator.
- the termination signal may be a lack of transcribable or translatable sequence, such as due to a sequence truncation.
- the invention also provides gene therapy vectors comprising a nucleic acid cassette of the invention.
- nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in a given host cell.
- Any appropriate host cell may be used, such as mammalian, bacterial, insect, yeast, and/or plant host cells.
- cell ⁇ free expression systems may be used. Such expression systems and host cells are standard in the art.
- nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in the lung.
- the nucleic acid cassettes and vectors of the invention may be capable of expressing the transgene in the lung parenchyma.
- the nucleic acid cassettes and vectors of the invention may be capable of expressing the transgene in one or more cell type selected from ATII cells, ATI cells, club cells, and/or bronchioalveolar stem cells.
- the nucleic acid cassettes and vectors of the invention are typically capable of expressing the transgene in one or more ATII cells, ATI cells, club cells, bronchioalveolar stem cells in the terminal bronchioles.
- the nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in ATII cells.
- the nucleic acid cassettes and vectors of the invention are capable of expressing the (therapeutic) protein encoded by the transgene in one or more cell/cell type of the lung parenchyma, particularly ATII cells, ATI cells, club cells and/or bronchioalveolar stem cells in the terminal bronchioles, particularly in ATII cells.
- the nucleic acid cassettes of the invention may be made using any suitable process known in the art.
- the nucleic acid cassettes may be made using chemical synthesis techniques.
- the nucleic acid cassettes of the invention may be made using molecular biology techniques.
- Non ⁇ Viral Vectors The present invention also provides a vector comprising a nucleic acid cassette of the invention.
- the vector(s) may be present in the form of a therapeutic composition or formulation.
- the vector may be a non ⁇ viral vector.
- the non ⁇ viral vector(s) may be a DNA vector, such as a DNA plasmid.
- the vector(s) may be an RNA vector, such as a mRNA vector or a self ⁇ amplifying RNA vector.
- the non ⁇ viral vector may be an exosomes or microvesicle (MV).
- the non ⁇ viral (e.g. DNA and/or RNA) vector(s) of the invention may be capable of expression in eukaryotic and/or prokaryotic cells.
- the non ⁇ viral e.g.
- DNA and/or RNA vector(s) are capable of expression in a cell of a subject, for example, a cell of a mammalian or avian subject to be immunised.
- the nucleic acid cassettes and vectors of the invention are capable of expressing a transgene in airway cells, preferably lung parenchymal cells (as described herein).
- a non ⁇ viral vector of the present invention may be a phage vector, such as an AAV/phage hybrid vector as described in Hajitou et al., Cell 2006; 125(2) pp. 385 ⁇ 398; herein incorporated by reference.
- Vector(s) of the present invention e.g.
- non ⁇ viral DNA or RNA vectors may be designed in silico, and then synthesised by conventional polynucleotide synthesis techniques.
- Non ⁇ viral plasmids cannot replicate in the subject to be treated, as they lack the viral genetic material which hijacks the body's normal production machinery. However they are capable of replicating in appropriate host cells, such as yeasts or bacteria including E. coli, and particularly airway cells as defined herein.
- the term "plasmid” as used herein refers to a construction comprised of genetic material designed to direct transformation of a targeted cell.
- the plasmid contains a plasmid backbone.
- a "plasmid backbone” as used herein contains multiple genetic elements positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transfected cells.
- the plasmid backbone can contain one or more unique restriction sites within the backbone.
- the plasmid may be capable of autonomous replication in a defined host or organism such that the cloned sequence is reproduced.
- the plasmid can confer some well ⁇ defined phenotype on the host organism which is either selectable or readily detected.
- the plasmid or plasmid backbone may have a linear or circular configuration.
- the components of a plasmid can contain, but is not limited to, a DNA molecule incorporating: (1) the plasmid backbone; (2) a sequence comprising or consisting of an SFTPB promoter fragment; (3) a transgene sequence encoding a (therapeutic) protein; and optionally (4) additional regulatory elements for transcription, translation, RNA stability and replication.
- the purpose of the plasmid in human gene therapy for the efficient delivery of nucleic acid sequences to, and expression of therapeutic proteins in, a cell or tissue.
- the purpose of the plasmid is to achieve high copy number, avoid potential causes of plasmid instability and provide a means for plasmid selection.
- the nucleic acid cassette contains the necessary elements for expression of the nucleic acid within the cassette. Expression includes the efficient transcription of an inserted gene, nucleic acid sequence, or nucleic acid cassette with the plasmid.
- a DNA plasmid may be CpG ⁇ free, or be optimised to reduce CpG dinucleotides as described herein.
- a DNA plasmid of the invention may be codon ⁇ optimised as described herein. Methods of preparing plasmid DNA are well known in the art. Typically, they are capable of autonomous replication in an appropriate host or producer cell.
- the term "exosome” as used herein refers to an extracellular vesicle formed by exocytosis from a cell of origin.
- An exosome typically comprises a nucleic cassette of the invention. Exosomes may be used to transform a targeted cell.
- the nucleic acid within an exosome may contain one or more genetic elements, which may be positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transformed cells.
- the term "microvesicle” (MV) as used herein refers to an extracellular vesicle formed typically between about 30 to about 1,000 nm in diameter.
- An MV typically comprises a nucleic cassette of the invention. Exosomes may be used to transform a targeted cell.
- the nucleic acid within an MV may contain one or more genetic elements, which may be positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transformed cells.
- Host cells containing (e.g. transformed, transfected, or electroporated with) the plasmid may be prokaryotic or eukaryotic in nature, either stably or transiently transformed, transfected, or electroporated with the plasmid.
- Suitable host cells include bacterial, yeast, fungal, invertebrate, and mammalian cells.
- the host cell is bacterial; more preferably E. coli.
- Host cells can then be used in methods for the large scale production of the plasmid.
- the cells are grown in a suitable culture medium under favourable conditions, and the desired plasmid isolated from the cells, or from the medium in which the cells are grown, by any purification technique well known to those skilled in the art; e.g. see Sambrook et al, supra.
- Any appropriate delivery means can be used to deliver a non ⁇ viral vector (e.g. plasmid) of the invention to a target cell or patient.
- Suitable delivery means are known in the art and within the routine skill of one of ordinary skill in the art.
- Non ⁇ limiting examples include the use of cationic lipids, polymers (e.g. polyethyleneimine and poly ⁇ L ⁇ lysine) and electroporation.
- cationic lipids may be used to deliver non ⁇ viral (e.g. plasmid) vectors of the invention to target cells or to a patient.
- non ⁇ viral e.g. plasmid
- cationic lipids suitable for use according to the invention are GL67A and lipofectamine.
- the cationic lipid mixture GL67A is a mixture of three components ⁇ GL67 (Cholest ⁇ 5 ⁇ en ⁇ 3 ⁇ ol (3 ⁇ ) ⁇ ,3 ⁇ [(3 ⁇ aminopropyl)[4 ⁇ [(3 ⁇ aminopropyl)amino]butyl]carbamate], (CAS Number: 179075 ⁇ 30 ⁇ 0)), DOPE (1,2 ⁇ dioleoyl ⁇ sn ⁇ glycero ⁇ 3 ⁇ phosphoethanolamine) and DMPE ⁇ PEG5000 (1,2 ⁇ Dimyristoyl ⁇ sn ⁇ Glycero ⁇ 3 ⁇ Phosphoethanolamine ⁇ N ⁇ [methoxy (Polyethylene glycol)5000]). These components are formulated at a 1:2:0.05 molar ratio to form GL67A.
- Lipofectamine consists of a 3:1 mixture of DOSPA (2,3 ⁇ dioleoyloxy ⁇ N ⁇ [2(sperminecarboxamido)ethyl] ⁇ N,N ⁇ dimethyl ⁇ 1 ⁇ propaniminium trifluoroacetate) and DOPE.
- Viral vectors The present invention also provides a vector comprising a nucleic acid cassette of the invention.
- the vector(s) may be present in the form of a therapeutic composition or formulation.
- the vector may be a viral vector.
- a viral vector of the invention may be a lentiviral vector, an adeno ⁇ associated virus (AAV) vector, an adenoviral vector, a poxvirus vector, a herpes simplex virus (HSV) vector. Derivatives of these viral vectors, such as lentivirus ⁇ derived particles are also encompassed within the invention.
- adenoviral vectors include human serotypes such as AdHu5, simian serotypes such as ChAd63, ChAdOX1 or ChAdOX2, and other forms.
- Non ⁇ limiting examples of poxvirus vectors include a modified vaccinia Ankara (MVA)).
- ChAdOX1 and ChAdOX2 are disclosed in WO2012/172277 (herein incorporated by reference in its entirety).
- ChAdOX2 is a BAC ⁇ derived and E4 modified AdC68 ⁇ based viral vector.
- Viral vectors are usually non ⁇ replicating or replication impaired vectors, which means that the viral vector cannot replicate to any significant extent in normal cells (e.g. normal human cells), as measured by conventional means – e.g. via measuring DNA synthesis and/or viral titre.
- Non ⁇ replicating or replication impaired vectors may have become so naturally (i.e. they have been isolated as such from nature) or artificially (e.g.
- viral vector is incapable of causing a significant infection in an animal subject, typically in a mammalian subject such as a human or other primate.
- Viral vector(s) of the present invention may be designed in silico, and then synthesised by conventional polynucleotide synthesis techniques.
- the invention relates to retroviral vectors, particularly lentiviral vectors.
- lentivirus refers to a family of retroviruses.
- Retroviral/lentiviral vectors of the invention can integrate into the genome of transduced cells and lead to long ⁇ lasting expression.
- retroviruses suitable for use in the present invention include gammaretroviruses such as murine leukaemia virus (MLV) and feline leukaemia virus (FLV).
- lentiviruses suitable for use in the present invention include Simian immunodeficiency virus (SIV), Human immunodeficiency virus (HIV), Feline immunodeficiency virus (FIV), Equine infectious anaemia virus (EIAV), and Visna/maedi virus.
- a particularly preferred lentiviral vector is an SIV vector (including all strains and subtypes), such as a SIV ⁇ AGM (originally isolated from African green monkeys, Cercopithecus aethiops).
- the retroviral/lentiviral (e.g. SIV) vectors of the present invention are typically pseudotyped with hemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, or with G glycoprotein from Vesicular Stomatitis Virus (G ⁇ VSV).
- HN hemagglutinin ⁇ neuraminidase
- F fusion
- G ⁇ VSV Vesicular Stomatitis Virus
- the lentiviral (e.g. SIV) vectors of the present invention are pseudotyped with HN and F from a respiratory paramyxovirus.
- the respiratory paramyxovirus is a Sendai virus (murine parainfluenza virus type 1).
- the F protein may be a truncated F protein, typically one in which the cytoplasmic domain is truncated.
- the truncated F protein is Fct4, in which 38 amino acids have been truncated from the C ⁇ terminus of the F protein, with 4 amino acids of the F protein cytoplasmic domain being retained.
- the F protein may be a truncated F protein, typically one in which the cytoplasmic domain is truncated.
- the truncated F protein is Fct4, in which 38 amino acids have been truncated from the C ⁇ terminus of the F protein, with 4 amino acids of the F protein cytoplasmic domain being retained.
- a retroviral/lentiviral (e.g. SIV) vector for use according to the invention may be integrase ⁇ competent (IC).
- the lentiviral (e.g. SIV) vector may be integrase ⁇ deficient (ID).
- Viral vectors of the invention, particularly retroviral/lentiviral (e.g. SIV) vectors as described herein may transduce one or more cells types as described herein to achieve long term transgene expression.
- the HN protein may be a truncated and/or chimeric HN protein, typically one in which the cytoplasmic domain is truncated or substituted.
- the HN protein is a chimeric HN protein in which (i) the cytoplasmic domain of the HN is replaced by the cytoplasmic domain of the transmembrane (TMP) protein; or (ii) the cytoplasmic domain of the TMP is added to the cytoplasmic domain of the HN protein.
- TMP transmembrane
- the HN protein may be as described in Kobayashi et al. (J. Virol. (2003) 77(4):2607 ⁇ 2614), which is herein incorporated by reference in its entirety.
- the viral vectors of the invention particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention enable high levels of transgene expression. Together with the increased levels of expression/secretion/membrane insertion of a therapeutic protein resulting from the use of an SFTPB promoter fragment of the invention, these viral vectors typically result in high levels (therapeutic levels) of expression of the transgene, and the (therapeutic) protein encoded by said transgene.
- retroviral/lentiviral vectors of the present invention enable high levels of transgene expression. Together with the increased levels of expression/secretion/membrane insertion of a therapeutic protein resulting from the use of an SFTPB promoter fragment of the invention, these viral vectors typically result in high levels (therapeutic levels) of expression of the transgene, and the (therapeutic) protein encoded by said transgene.
- the transgene to be included in a viral vector of the invention may be modified to facilitate expression.
- the transgene sequence may be in CpG ⁇ depleted /low (or CpG ⁇ fee) and/or codon ⁇ optimised form to facilitate gene expression. Standard techniques for modifying the transgene sequence in this way are known in the art.
- the viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention exhibit enhanced expression of the transgene. Accordingly, the viral vectors of the invention, particularly the retroviral/lentiviral (e.g.
- SIV vectors of the invention are capable of producing long ⁇ lasting, repeatable, high ⁇ level transgene expression, particularly in lung parenchyma without inducing side effects (e.g., an undue immune response).
- the viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention enable long ⁇ term transgene expression, resulting in long ⁇ term expression (and secretion or membrane insertion) of a (therapeutic) protein by cells of the lung parenchyma as described herein.
- Long ⁇ term expression means expression of a transgene gene and/or encoded (therapeutic) protein, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more.
- long ⁇ term expression means expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
- the invention relates to the use of F/HN lentiviral vectors comprising a nucleic acid cassette of the invention, particularly SIV F/HN vectors.
- the nucleic acid cassette comprised in a viral vector of the invention, particularly a retroviral/lentiviral (e.g. SIV) vector of the invention may have no intron positioned between the promoter and the nucleic acid encoding the signal peptide and/or the nucleic acid encoding the therapeutic protein.
- the viral vectors of the invention may be made using any suitable process known in the art.
- retroviral/lentiviral (e.g. SIV) vectors of the invention may be made using the methods disclosed in UK Application No. 2102832.9, which is herein incorporated by reference in its entirety).
- the viral vectors of the invention may comprise a central polypurine tract (cPPT) and/or the Woodchuck hepatitis virus posttranscriptional regulatory elements (WPRE).
- cPPT central polypurine tract
- WPRE Woodchuck hepatitis virus posttranscriptional regulatory elements
- An exemplary WPRE sequence is provided by SEQ ID NO: 52.
- Transgenes A nucleic acid cassette of the invention comprises a transgene. Typically, the transgene encodes a therapeutic protein. A therapeutic protein is one which has potential utility in the treatment or prevention of a disease or condition, such as those describe herein.
- a nucleic acid cassette of the invention may comprise a nucleic acid encoding a protein which has a therapeutic effect on a disease or condition to be treated.
- a nucleic acid cassette of the invention may comprise a nucleic acid encoding a therapeutic protein which is a functional or wild ⁇ type form of a protein which is present in a patient to be treated in a dysfunctional form (whether the dysfunction is inherent or acquired).
- the phrase "inherent dysfunction” refers to a protein which is innately dysfunctional due to genetic factors and the phrase “acquired dysfunction” refers to a protein which is dysfunctional due to environmental or other factors after birth.
- a nucleic acid cassette of the invention may comprise a nucleic acid encoding a protein which is a functional or wild ⁇ type form of a protein which is present in a patient, but which that has become dysfunctional due to a genetic disease, such as a genetic respiratory disease.
- the nucleic acid cassettes of the present invention are useful in the treatment of diseases via their use in expressing therapeutic proteins in lung parenchyma as described herein (e.g. ATI, ATII cells).
- the nucleic acid cassettes of the present invention are useful in the treatment of diseases via their use in expressing therapeutic proteins in ATII cells, ATI cells and/or club cells.
- the transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally Surfactant Protein B (SFTPB), Surfactant Protein C (SFTPC), alpha ⁇ 1 ⁇ antitrypsin (AAT), Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, von Willebrand Factor, Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF), decorin, an anti ⁇ inflammatory protein (e.g.
- the therapeutic protein is not an antibody, particularly not a monoclonal antibody.
- the transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally SFTPB, SP ⁇ C, AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, von Willebrand Factor, GM ⁇ CSF, an anti ⁇ inflammatory protein (e.g.
- the nucleic acid cassettes of the invention are particularly efficient at driving the expression, secretion and/or membrane insertion of proteins (e.g. therapeutic proteins as described herein) by the lung parenchyma. This is particularly the case when such cassettes are comprised within F/HN pseudotyped viral vectors of the invention (as described herein), which are efficient at targeting cells in the lung parenchyma.
- the nucleic acid cassettes of the invention and vectors comprising said cassettes
- nucleic acid cassettes of the invention are typically delivered to lung parenchyma as described herein. Accordingly, the nucleic acid cassettes of the invention (and vectors comprising said cassettes) are particularly suited for treatment of diseases or disorders of the lung parenchyma. Typically, the nucleic acid cassettes of the invention (and vectors comprising said cassettes) may be used for the treatment of a genetic respiratory disease.
- a nucleic acid cassette of the invention (or vector comprising said cassette) may comprise a nucleic acid encoding a polypeptide or protein that is therapeutic for the treatment of such diseases, particularly a disease or disorder of the lung parenchyma.
- a nucleic acid cassette of the invention may comprise a nucleic acid sequence encoding a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally SFTPB, SFTPC, AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, von Willebrand Factor, GM ⁇ CSF, an anti ⁇ inflammatory protein (e.g.
- IL ⁇ 10, TGG ⁇ , or TNF ⁇ alpha or monoclonal antibody, an anti ⁇ inflammatory decoy and a monoclonal antibody against an infectious agent; or (b) ABCA3, TRIM72, CSF2RA, CSF2RB or DCN.
- Other preferred examples of therapeutic proteins that may be encoded by a nucleic acid sequence comprised in a nucleic acid cassette of the invention (or vector comprising said cassette) include genes related to or associated with other surfactant deficiencies.
- the therapeutic protein encoded by a nucleic acid cassette of the invention (or vector comprising said cassette) may be SFTPB. Examples of an SFTPB therapeutic transgene are provided by SEQ ID NOs: 53, 54 and 56.
- SFTPB transgene An exemplary codon ⁇ optimised SFTPB transgene is provided by SEQ ID NO: 55.
- the therapeutic protein encoded by said SFTPB transgene may be exemplified by the polypeptide of SEQ ID NO: 57. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 53 to 57.
- the transgene may encode ABCA3. Examples of a ABACA3 transgene are provided by SEQ ID NOs: 58 and 59.
- An exemplary codon ⁇ optimised ABACA3 transgene is provided by SEQ ID NO: 60.
- the polypeptide encoded by said ABACA3 transgene may be exemplified by the polypeptide of SEQ ID NO: 61. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 58 to 61.
- the therapeutic protein encoded by a nucleic acid cassette of the invention may be SFTPC.
- An example of an SFTPC therapeutic transgene is provided by SEQ ID NO: 62.
- the therapeutic protein encoded by said SFTPC transgene may be exemplified by the polypeptide of SEQ ID NO: 63.
- variants thereof are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 62 or 63.
- the therapeutic protein encoded by a nucleic acid cassette of the invention may be an AAT.
- An example of an AAT therapeutic transgene (SERPINA1) is provided by SEQ ID NO: 70.
- SEQ ID NO: 70 is a codon ⁇ optimized CpG depleted AAT transgene (SERPINA1) previously designed by the present inventors to enhance translation in human cells. Such optimisation has been shown to enhance gene expression by up to 15 ⁇ fold.
- variants of same sequence which possess the same technical effect of enhancing translation compared with the unmodified (wild ⁇ type) AAT gene sequence are also encompassed by the present invention.
- the therapeutic protein encoded by said AAT transgene may be exemplified by the polypeptide of SEQ ID NO: 71. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 70 or 71.
- the therapeutic protein encoded by a nucleic acid cassette of the invention may be an FVIII.
- FVIII therapeutic transgene examples are provided by SEQ ID NOs: 72 and 73.
- the polypeptide encoded by the FVIII transgene may be exemplified by the polypeptide of SEQ ID NO: 74 and 75. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 72 to 75.
- the therapeutic protein encoded by a nucleic acid cassette of the invention may be GM ⁇ CSF.
- a GM ⁇ CSF transgene may comprise or consist of SEQ ID NO: 64 (human).
- the polypeptide encoded by the GM ⁇ CSF transgene may be exemplified by the polypeptide of SEQ ID NO: 65 (human). Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 64and 65.
- the transgene may encode decorin.
- An example of a DCN transgene is provided by SEQ ID NO: 66.
- the polypeptide encoded by said DCN transgene may be exemplified by the polypeptide of SEQ ID NO: 67.
- Variants thereof are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to SEQ ID NO: 66or 67.
- the transgene may encode TRIM72.
- An example of a TRIM72 transgene is provided by SEQ ID NO: 68.
- the polypeptide encoded by said TRIM72 transgene may be exemplified by the polypeptide of SEQ ID NO: 69.
- Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to SEQ ID NO: 68or 69.
- the therapeutic protein encoded by a nucleic acid cassette of the invention may be encoded by any one of SFTPB, SFTPC, Factor V, Factor VII, Factor IX, Factor X and/or Factor XI, von Willebrand Factor, GM ⁇ CSF, ABCA3, TRIM72 or DCN, or other known related gene.
- the transgene may preferably be SFTPB, SFTPC, ABCA3 or GM ⁇ CSF.
- the therapeutic protein may be a monoclonal antibody (mAb) against an infectious agent (bacterial, fungal or viral, e.g.
- the therapeutic protein may be anti ⁇ TNF alpha.
- the therapeutic protein may be one implicated in an inflammatory, immune or metabolic condition.
- a nucleic acid cassette of the invention (or a vector comprising said cassette) may be delivered to one or more cell/cell type of the lung parenchyma to allow production of proteins to be secreted into circulatory system.
- the therapeutic protein may be any one of Factor VII, Factor VIII, Factor IX, Factor X, Factor XI and/or von Willebrand’s factor.
- nucleic acid cassette of the invention may be used in the treatment of diseases, particularly cardiovascular diseases and blood disorders, preferably blood clotting deficiencies such as haemophilia.
- the therapeutic protein may be an mAb against an infectious agent or a protein implicated in an inflammatory, immune or metabolic condition, such as, lysosomal storage disease.
- the nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the promoter and the nucleic acid encoding the therapeutic protein.
- nucleic acid cassette when the nucleic acid cassette is comprised in a viral vector, there may be no intron between the promoter and the transgene in the vector genome (pDNA1) plasmid used to make said viral vector, as described herein.
- said nucleic acid cassette of the invention (or a vector comprising said cassette) may have an intron positioned between the promoter and the transgene, particularly if the cassette (or vector comprising said cassette) is non ⁇ viral.
- the nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an SFTPB transgene, including those described herein.
- nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an SFTPC transgene, including those described herein.
- said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an AAT transgene (SERPINA1), including those described herein.
- nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FVIII transgene, including those described herein.
- said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FVII transgene, including those described herein.
- nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FIX transgene, including those described herein.
- said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FX transgene, including those described herein.
- nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FXI transgene, including those described herein.
- said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and a von Willebrand Factor transgene, including those described herein.
- nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and a Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF) transgene, including those described herein.
- GM ⁇ CSF Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor
- said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the lentiviral e.g.
- SIV vector comprises a SFTPB promoter fragment and an DCN transgene, including those described herein.
- said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the lentiviral (e.g. SIV) vector comprises a SFTPB promoter fragment and a TRIM72 transgene, including those described herein.
- said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the lentiviral e.g.
- SIV vector comprises a SFTPB promoter fragment and a ABACA3 transgene, including those described herein.
- said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
- the nucleic acid cassette of the invention (or a vector comprising said cassette) comprises a nucleic acid encoding a therapeutic protein (said nucleic acid is referred to interchangeably herein as a transgene).
- the nucleic acid sequence encodes a gene product, e.g., a protein, particularly a therapeutic protein.
- the nucleic acid cassette of the invention (or a vector comprising said cassette) comprises a nucleic acid sequence encoding an SFTPB, SFTPC, AAT, GM ⁇ CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 and said nucleic acid sequence comprises (or consists of) a nucleic acid sequence having at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100%) sequence identity to the SFTPB, SFTPC, AAT, GM ⁇ CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 nucleic acid sequence respectively, examples of which are described herein.
- the nucleic acid sequence encoding SFTPB, SFTPC, AAT, GM ⁇ CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 may preferably comprise (or consist of) a nucleic acid sequence having at least 95% (such as at least 95, 96, 97, 98, 99 or 100%) sequence identity to the SFTPB, SFTPC, AAT, GM ⁇ CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 nucleic acid sequence respectively, examples of which are described herein.
- the amino acid sequence of the (therapeutic) protein encoded by the transgene may be a functional variant having at least 95% (such as at least 95, 96, 97, 98, 99 or 100%) sequence identity to the functional protein.
- SFTPB promoter fragment and/or enhancer of the invention may be linked to a transgene by a linker.
- Said linker is typically a short DNA sequence, as defined herein, and may comprise or consist of one or more restriction enzyme site, examples of which are also described herein.
- a SFTPB promoter fragment of the invention may be joined to a transgene by a linker which comprises or consists of a NheI restriction site and/or a BgIII restriction site, preferably a NheI restriction site.
- Signal peptides The transgene encoding for a (therapeutic) protein may further comprise a nucleic acid sequence encoding for a signal peptide.
- Said signal peptide may be the endogenous signal peptide of the (therapeutic) protein, or a signal peptide exogenous to said (therapeutic) protein.
- the transgene further comprises a nucleic acid encoding for an exogenous signal peptide
- the transgene preferably exclude a nucleic acid sequence encoding for the endogenous signal peptide.
- the exogenous signal peptide is typically the sole signal peptide linked with (and hence driving secretion and/or membrane insertion) of the therapeutic protein. All disclosure herein relates to both transgenes and therapeutic proteins including and excluding endogenous signal peptides unless explicitly stated.
- sequence identity of variants, and/or lengths of fragments may be based on the sequence with or without a signal peptide.
- Any signal peptide and therapeutic combination may be used, provided that this combination is effective in increasing the expression, secretion and/or membrane insertion of a (therapeutic) protein as defined herein.
- Selection of a signal peptide may depend on the specific (therapeutic) protein and/or the specific lung parenchymal cell type by which the (therapeutic) protein is to be expressed/secreted/inserted into the cell membrane.
- lentiviral vectors such as the lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus
- HN haemagglutinin ⁇ neuraminidase
- F fusion
- the invention therefore provides lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered in combination with a surfactant.
- HN haemagglutinin ⁇ neuraminidase
- F fusion
- the invention therefore provides lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered simultaneously or sequentially with a surfactant.
- the lentiviral (e.g. SIV) vector and surfactant are administered in combination.
- Administered "in combination,” encompasses both simultaneous (also referred to as concurrent) administration/delivery and sequential (also referred to as separate) administration/delivery.
- the delivery of the lentiviral (e.g. SIV) vector may still be occurring when the delivery of the surfactant begins, or the delivery of surfactant may still be occurring when the delivery of the lentiviral (e.g. SIV) vector begins, so that there is overlap in terms of administration.
- Simultaneous delivery may encompass delivery of the lentiviral (e.g. SIV) vector and surfactant within weeks to months or even years of each other, typically so that the lentiviral (e.g. SIV) vector delivery overlaps with the delivery of the surfactant.
- the delivery of the lentiviral (e.g. SIV) vector may still be occurring when the delivery of the surfactant begins, or the delivery of surfactant may still be occurring when the delivery of the lentiviral (e.g. SIV) vector begins, so that there is overlap in terms of administration.
- Simultaneous delivery may encompass delivery of the lentiviral (e.g. SIV) vector and surfactant within weeks to months or even
- SIV vector may end before the delivery of the surfactant begins, or the delivery of the surfactant may end before delivery of the lentiviral (e.g. SIV) vector begins.
- Sequential administration may involve the lentiviral (e.g. SIV) vector and surfactant being administered within 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 6 hours, 12 hours or 24 hours or longer of each other.
- the lentiviral vector may be administered before the surfactant.
- the surfactant may be administered before the lentiviral vector.
- the lentiviral vector and surfactant may be administered simultaneously, optionally wherein the lentiviral vector and surfactant are mixed prior to administration.
- the treatment is more effective because of combined administration.
- treatment with the lentiviral e.g.
- SIV vector may be more effective, e.g., an equivalent effect is seen with less of the lentiviral (e.g. SIV) vector, or the lentiviral (e.g. SIV) vector reduces symptoms to a greater extent, than would be seen if the lentiviral (e.g. SIV) vector were administered in the absence of the surfactant.
- treatment with the surfactant may be more effective, e.g., an equivalent effect is seen with less of the surfactant, or the surfactant reduces symptoms to a greater extent, than would be seen if the surfactant were administered in the absence of the lentiviral (e.g. SIV) vector.
- a combination therapy of the invention may increase transgene expression by at least 1.2 fold, at least 1.3 fold, at least 1.4 fold, at least 1.5 fold, at least 2 fold, at least 2.5 fold or more compared with treatment with the lentiviral (e.g. SIV) vector alone (i.e. compared with the increase in transgene expression achieved when treating with the lentiviral (e.g. SIV) alone).
- the surfactant may aid the distribution of the lentiviral (e.g. SIV) vector within the lungs, enabling it to penetrate more deeply into the respiratory tree and thus facilitating transduction of the lung parenchyma.
- appropriate dosage of the lentiviral (e.g. SIV) vector and/or the surfactant will depend on the specific agent, and can also vary from patient to patient. Any surfactant may be used in a combination therapy according to the present invention. It will be appreciated that it is within the routine practice of one of ordinary skill in the art to select such a surfactant.
- a disease to be treated with such a lentiviral vector and a surfactant may be a genetic disease.
- the disease to be treated may be a respiratory disease, particularly a genetic respiratory disease; or a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder.
- Non ⁇ limiting examples of diseases which may be treated according to this aspect of the invention include Surfactant Protein B (SP ⁇ B) Deficiency; Surfactant Protein C (SP ⁇ C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1 ⁇ antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID ⁇ 19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia.
- SP ⁇ B Surfactant Protein B
- SP ⁇ C Surfactant Protein C
- ABCA3 deficiency Pulmonary surfactant metabolism dysfunction 2
- SMDP3 Pulmonary surfactant metabolism dysfunction 3
- a lentiviral vector for use in combination with a surfactant may comprise HN and F proteins from a Sendai virus, as described herein.
- a lentiviral vector for use in combination with a surfactant may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector.
- said lentiviral vector may be a SIV vector.
- the transgene may encode any suitable therapeutic protein as described herein.
- said transgene may be selected from: (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP ⁇ B), Surfactant Protein C (SP ⁇ C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF), decorin, an anti ⁇ inflammatory protein (e.g.
- a lentiviral vector for use in combination with a surfactant may comprise a SFTPB promoter fragment as defined herein.
- said lentiviral vector may comprise a nucleic acid cassette as defined herein.
- Therapeutic Indications The nucleic acid cassettes and vectors of the present invention enable cell ⁇ preferred or cell ⁇ specific expression of a transgene encoding a (therapeutic) protein.
- nucleic acid cassette or vector facilitating efficient transgene expression.
- the nucleic acid cassettes and vectors of the invention, and particularly the F/HN ⁇ pseudotyped retroviral/lentiviral (e.g. SIV) vectors of the invention are capable of: (i) transduction of one or more cell/cell type of the lung parenchyma without disruption of epithelial integrity; (ii) persistent gene expression; (iii) lack of chronic toxicity; and/or (iv) efficient repeat administration. Long term/persistent stable gene expression, preferably at a therapeutically ⁇ effective level, may be achieved using repeat doses of a nucleic acid cassette or vector of the present invention.
- the nucleic acid cassettes and vectors of the present invention and particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention can be used in gene therapy.
- the present invention provides a nucleic acid cassette or gene therapy vector as defined herein for use in a method of treating or preventing a disease.
- the disease to be treated may be chronic or acute.
- the nucleic acid cassettes and vectors (viral and non ⁇ viral) of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention may be used to deliver any transgene useful in gene therapy.
- the nucleic acid cassettes and vectors (viral and non ⁇ viral) of the invention are for use in gene therapy for the treatment of a disease or disorder of the lung parenchyma.
- efficient airway cell uptake properties of the nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention make them highly suitable for treating respiratory or lung diseases, particularly genetic respiratory diseases, particularly preferably those of/involving the lung parenchyma.
- SIV vectors of the invention can also be used in methods of gene therapy to promote secretion of therapeutic proteins.
- the invention provides secretion of therapeutic proteins into alveoli or lumen of the bronchioles.
- Administration of a nucleic acid cassettes and vectors of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention and its uptake by airway cells may be used to enable the use of the lungs as a “factory” to produce a therapeutic protein that is then secreted and enters the general circulation at therapeutic levels, where it can travel to cells/tissues of interest to elicit a therapeutic effect.
- nucleic acid cassettes and vectors of the present invention can also be treated by the nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention.
- Nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention can effectively treat a disease by providing a transgene for the correction of the disease. For example, resulting in the expression and secretion of SFTPB from cells of the lung parenchyma, to compensate for the pathologically low levels of SFTPB expression in patients with SFTPB deficiency.
- nucleic acid cassettes and vectors of the present invention may be used to treat alpha ⁇ 1 ⁇ antitrypsin (AAT) deficiency, typically by gene therapy with a AAT transgene (SERPINA1) as described herein.
- AAT alpha ⁇ 1 ⁇ antitrypsin
- SERPINA1 AAT transgene
- AAT is a secreted anti ⁇ protease that is produced mainly in the liver and then trafficked to the lung, with smaller amounts also being produced in the lung itself.
- the main function of AAT is to bind and neutralise/inhibit neutrophil elastase.
- Gene therapy with AAT according to the present invention is relevant to AAT deficient patient, as well as in other lung diseases such as CF or chronic obstructive pulmonary disease (COPD), and offers the opportunity to overcome some of the problems encountered by conventional enzyme replacement therapy (in which AAT isolated from human blood and administered intravenously every week), providing stable, long ⁇ lasting expression in the target tissue (lung/nasal epithelium), ease of administration and unlimited availability.
- Transduction with a nucleic acid cassettes and vectors of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention may lead to secretion of the recombinant protein into the lumen of the lung as well as into the circulation.
- AAT gene therapy may therefore also be beneficial in other disease indications, non ⁇ limiting examples of which include type 1 and type 2 diabetes, acute myocardial infarction, ischemic heart disease, rheumatoid arthritis, inflammatory bowel disease, transplant rejection, graft versus host (GvH) disease, multiple sclerosis, liver disease, cirrhosis, vasculitides and infections, such as bacterial and/or viral infections.
- AAT has numerous other anti ⁇ inflammatory and tissue ⁇ protective effects, for example in pre ⁇ clinical models of diabetes, graft versus host disease and inflammatory bowel disease.
- AAT in the lung and/or nose following transduction according to the present invention may, therefore, be more widely applicable, including to these indications.
- diseases that may be treated with gene therapy of a secreted protein according to the present invention include cardiovascular diseases and blood disorders, particularly blood clotting deficiencies such as haemophilia (A, B or C), von Willebrand disease and Factor VII deficiency.
- the disease to be treated is selected from a surfactant protein deficiency, such as Surfactant Protein B (SFTPB) Deficiency, Surfactant Protein C (SFTPC) deficiency, ABCA3 deficiency, Pulmonary surfactant metabolism dysfunction 2 (SMDP2) Pulmonary surfactant metabolism dysfunction 3 (SMDP3), or other surfactant deficiencies; Primary Ciliary Dyskinesia (PCD); Alpha 1 ⁇ antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP, hereditary and/or acquired); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID ⁇ 19; a pulmonary fibrotic disease (including idiopathic pulmonary fibrosis); a pulmonary allergic condition; a pulmonary bacterial infection; asthma; lung cancer; a dysplastic change in the lungs; and haemophilia.
- SFTPB Surfactant Protein B
- SFTPC Surfactant Protein C
- diseases or disorders to be treated include Primary Ciliary Dyskinesia (PCD), acute lung injury, and/or inflammatory, infectious, immune or metabolic conditions, such as lysosomal storage diseases or a pulmonary bacterial infection, or any other lung disease or disorder.
- PCD Primary Ciliary Dyskinesia
- the nucleic acid cassettes and vectors (viral and non ⁇ viral) of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention typically provide high expression levels of a transgene of interest, and the (therapeutic) protein encoded thereby, when administered to a patient.
- high expression and therapeutic expression are used interchangeably herein.
- Expression may be measured by any appropriate method (qualitative or quantitative, preferably quantitative), and concentrations given in any appropriate unit of measurement, for example ng/ml or ⁇ M. Expression/secretion/membrane insertion of a transgene, or the (therapeutic) protein of interest encoded thereby may be given in absolute terms.
- expression/secretion/membrane insertion of a therapeutic protein may be given in relative terms, for example relative to the expression/secretion/membrane insertion of the therapeutic protein encoded by a corresponding nucleic acid cassette or vector of the invention without the SFTB promoter fragment of the invention, relative to the expression/secretion/membrane insertion of the same transgene using the full ⁇ length SFTB promoter as described herein, or relative to the expression/secretion/membrane insertion of the corresponding endogenous (defective) gene.
- Expression may be measured in terms of mRNA or protein expression.
- the expression of the therapeutic protein of the invention may be quantified relative to the endogenous protein or gene in terms of protein concentration, mRNA copies per cell or any other appropriate unit.
- Secretion and/or membrane insertion of a therapeutic protein may be quantified relative to secretion/membrane insertion of the corresponding endogenous protein, or relative to the level of secretion/membrane insertion of the therapeutic protein introduced via an expression cassette with the same transgene but without the SFTB promoter fragment of the invention or comprising the full ⁇ length SFTB promoter as described herein.
- Expression levels of a nucleic acid encoding a (therapeutic) protein and/or the expression/secretion/membrane insertion of the encoded (therapeutic) protein of the invention may be measured ex vivo (e.g. in the conditioned media used to culture the cells or within the cells themselves) or in vivo (e.g. in the lung tissue, epithelial lining fluid and/or serum/plasma) as appropriate.
- a high and/or therapeutic expression level may therefore refer to the concentration in the lung, epithelial lining fluid and/or serum/plasma.
- SIV vectors of the present invention may be administered twice ⁇ daily, daily, twice ⁇ weekly, weekly, monthly, every two months, every three months, every four months, every six months, yearly, every two years, or more. Dosing may be continued for as long as required, for example, for at least six months, at least one year, two years, three years, four years, five years, ten years, fifteen years, twenty years, or more, up to for the lifetime of the patient to be treated.
- the invention also provides nucleic acid cassettes and vectors of the present invention, and particularly retroviral/lentiviral (e.g.
- SIV vectors of the invention as described herein for use in a method of gene therapy comprising the steps of: (a) transducing cells (e.g. lung parenchyma) ex vivo to produce modified cells expressing a transgene of interest; and (b) administering the resulting modified cells.
- the invention provides a method of treating a disease, the method comprising administering a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention to a subject. Any disease described herein may be treated according to the invention.
- the invention provides a method of treating a lung disease using a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention.
- the disease to be treated may be a chronic disease.
- the invention also provides a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention as described herein for use in a method of treating a disease. Any disease described herein may be treated according to the invention.
- the invention provides a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention for use in a method of treating a lung disease.
- the disease to be treated may be a chronic disease.
- the invention also provides the use of a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention as described herein in the manufacture of a medicament for use in a method of treating a disease. Any disease described herein may be treated according to the invention.
- the invention provides the use of a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention for the manufacture of a medicament for use in a method of treating a lung disease.
- the disease to be treated may be a chronic disease.
- the invention also provides a cell comprising a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector.
- Said cell may be a lung parenchyma cell as described herein.
- the invention further provides a method of expressing a transgene, typically a transgene encoding a (therapeutic) protein in a target cell, comprising delivering a nucleic acid cassette or a vector of the invention into the target cells. Said method may be carried out in vitro, ex vivo, or in vivo, preferably in vitro or ex vivo.
- the target cell may be any appropriate cell type, such as those described herein.
- the target cells may be prokaryotic or eukaryotic, preferably eukaryotic. Particularly preferred are mammalian cells, such as human, non ⁇ human primate, mouse, rat, dog, cat, horse, or cow cells.
- the cells may be primary cells or cell lines. Non ⁇ limiting examples of cells include ATII cells, ATI cells, club cells and/or HEK293T cells.
- the step of delivering the nucleic acid cassette or vector may comprise integrating said nucleic acid cassette or gene therapy vector into said target cell's genome. Any appropriate technique may be used to deliver the nucleic acid cassette or vector, examples of which are known in the art and within the routine practice of one of ordinary skill in the art.
- the method may further comprise a step of culturing cells expressing the transgene, and/or isolating or purifying the expressed (therapeutic) protein from said cells.
- Any and all disclosure herein in relation to nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention applies equally and without reservation to the therapeutic uses and methods described herein. Further, for the avoidance of doubt, any and all disclosure herein in relation to therapeutic uses and methods using nucleic acid cassettes or vectors of the present invention, applies equally and without reservation to the combination therapies of lentiviral (e.g. SIV) vectors and surfactants as described herein.
- Long term/persistent stable gene expression may be achieved using repeat doses of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention of the present invention.
- a single dose may be used to achieve the desired long ⁇ term expression.
- Formulation and administration also provides a composition comprising a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention, and optionally a pharmaceutically acceptable carrier, excipient, buffer or diluent.
- the nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral e.g.
- SIV vectors of the invention may be administered in any dosage appropriate for achieving the desired therapeutic effect.
- Appropriate dosages may be determined by a clinician or other medical practitioner using standard techniques and within the normal course of their work.
- suitable dosages of viral vectors of the invention include 1x10 8 transduction units (TU), 1x10 9 TU, 1x10 10 TU, 1x10 11 TU or more.
- Non ⁇ limiting examples of suitable dosages of non ⁇ viral vectors/delivery means of the invention include a maximum of 30 mL per dose, a maximum of 25 mL per dose, a maximum of 20 mL per dose, a maximum of 15 mL per dose, a maximum of 10 mL per dose, or less, preferably a maximum of 20 mL per dose.
- Non ⁇ limiting examples of pharmaceutically acceptable carriers that may be comprised in a composition of the invention include water, saline, and phosphate ⁇ buffered saline. In some embodiments, however, the composition is in lyophilized form, in which case it may include a stabilizer, such as bovine serum albumin (BSA).
- BSA bovine serum albumin
- compositions with a preservative such as thiomersal or sodium azide
- a preservative such as thiomersal or sodium azide
- the nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered by any appropriate route. It may be desired to direct the compositions of the present invention (as described above) to the respiratory system of a subject. Efficient transmission of a therapeutic/prophylactic composition or medicament to the site of a disease or disorder in the respiratory tract may be achieved by oral or intra ⁇ nasal administration, for example, as aerosols (e.g. nasal sprays), or by catheters.
- aerosols e.g. nasal sprays
- nucleic acid cassettes or vectors of the present invention and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention are stable in clinically relevant nebulisers, inhalers (including metered dose inhalers), catheters and aerosols, etc.
- Other routes of administration including but not limited to i.v. administration, intranasal administration and intraplural injection are also encompassed by the present invention. Suitable administration routes are known in the art.
- the nose is a preferred production site for a therapeutic protein using nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g.
- SIV vectors of the invention for at least one of the following reasons: (i) extracellular barriers such as inflammatory cells and sputum are less pronounced in the nose; (ii) ease of vector administration; (iii) smaller quantities of vector required; and (iv) ethical considerations.
- nasal administration of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may result in efficient (high ⁇ level) and long ⁇ lasting expression of the therapeutic protein of interest. Accordingly, nasal administration of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be preferred.
- Formulations for intra ⁇ nasal administration may be in the form of nasal droplets or a nasal spray.
- An intra ⁇ nasal formulation may comprise droplets having approximate diameters in the range of 100 ⁇ 5000 ⁇ m, such as 500 ⁇ 4000 ⁇ m, 1000 ⁇ 3000 ⁇ m or 100 ⁇ 1000 ⁇ m.
- the droplets may be in the range of about 0.001 ⁇ 100 ⁇ l, such as 0.1 ⁇ 50 ⁇ l or 1.0 ⁇ 25 ⁇ l, or such as 0.001 ⁇ 1 ⁇ l.
- the aerosol formulation may take the form of a powder, suspension or solution. The size of aerosol particles is relevant to the delivery capability of an aerosol. Smaller particles may travel further down the respiratory airway towards the alveoli than would larger particles.
- the aerosol particles have a diameter distribution to facilitate delivery along the entire length of the bronchi, bronchioles, and alveoli.
- the particle size distribution may be selected to target a particular section of the respiratory airway, for example the alveoli.
- the particles may have diameters in the approximate range of 0.1 ⁇ 50 ⁇ m, preferably 1 ⁇ 25 ⁇ m, more preferably 1 ⁇ 5 ⁇ m.
- Aerosol particles may be for delivery using a nebulizer (e.g. via the mouth) or nasal spray.
- An aerosol formulation may optionally contain a propellant and/or surfactant. The formulation of pharmaceutical aerosols is routine to those skilled in the art, see for example, Sciarra, J.
- compositions comprising nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention, in particular where intranasal delivery is to be used, may comprise a humectant.
- Suitable humectants include, for instance, sorbitol, mineral oil, vegetable oil and glycerol; soothing agents; membrane conditioners; sweeteners; and combinations thereof.
- the compositions may comprise a surfactant.
- Suitable surfactants include non ⁇ ionic, anionic and cationic surfactants. Examples of surfactants that may be used include, for example, polyoxyethylene derivatives of fatty acid partial esters of sorbitol anhydrides, such as for example, Tween 80, Polyoxyl 40 Stearate, Polyoxy ethylene 50 Stearate, fusieates, bile salts and Octoxynol.
- nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be performed after an initial administration.
- the administration may, for instance, be at least a week, two weeks, a month, two months, six months, a year or more after the initial administration.
- nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered at least once a week, once a fortnight, once a month, every two months, every six months, annually or at longer intervals.
- administration is every six months, more preferably annually.
- nucleic acid cassettes or vectors of the present invention and particularly retroviral/lentiviral (e.g. SIV) vectors may, for instance, be administered at intervals dictated by when the effects of the previous administration are decreasing.
- retroviral/lentiviral vectors e.g. SIV
- any and all disclosure herein in relation formulations of nucleic acid cassettes or vectors of the present invention applies equally and without reservation to the formulations of lentiviral (e.g. SIV) vectors for combination therapies of lentiviral (e.g. SIV) vectors and surfactants as described herein.
- SEQUENCE HOMOLOGY Any of a variety of sequence alignment methods can be used to determine percent identity, including, without limitation, global methods, local methods and hybrid methods, such as, e.g., segment approach methods. Protocols to determine percent identity are routine procedures within the scope of one skilled in the art. Global methods align sequences from the beginning to the end of the molecule and determine the best alignment by adding up scores of individual residue pairs and by imposing gap penalties. Non ⁇ limiting methods include, e.g., CLUSTAL W, see, e.g., Julie D.
- Non ⁇ limiting methods include, e.g., Match ⁇ box, see, e.g., Eric Depiereux and Ernest Feytmans, Match ⁇ Box: A Fundamentally New Algorithm for the Simultaneous Alignment of Several Protein Sequences, 8(5) CABIOS 501 ⁇ 509 (1992); Gibbs sampling, see, e.g., C. E.
- % sequence identity between two or more nucleic acid or amino acid sequences is a function of the number of identical positions shared by the sequences. Thus, % identity may be calculated as the number of identical nucleotides / amino acids divided by the total number of nucleotides / amino acids, multiplied by 100. Calculations of % sequence identity may also take into account the number of gaps, and the length of each gap that needs to be introduced to optimize alignment of two or more sequences.
- a limited number of non ⁇ conservative amino acids, amino acids that are not encoded by the genetic code, and unnatural amino acids may be substituted for polypeptide amino acid residues.
- the polypeptides of the present invention can also comprise non ⁇ naturally occurring amino acid residues.
- Non ⁇ naturally occurring amino acids include, without limitation, trans ⁇ 3 ⁇ methylproline, 2,4 ⁇ methano ⁇ proline, cis ⁇ 4 ⁇ hydroxyproline, trans ⁇ 4 ⁇ hydroxy ⁇ proline, N ⁇ methylglycine, allo ⁇ threonine, methyl ⁇ threonine, hydroxy ⁇ ethylcysteine, hydroxyethylhomo ⁇ cysteine, nitro ⁇ glutamine, homoglutamine, pipecolic acid, tert ⁇ leucine, norvaline, 2 ⁇ azaphenylalanine, 3 ⁇ azaphenyl ⁇ alanine, 4 ⁇ azaphenyl ⁇ alanine, and 4 ⁇ fluorophenylalanine.
- coli cells are cultured in the absence of a natural amino acid that is to be replaced (e.g., phenylalanine) and in the presence of the desired non ⁇ naturally occurring amino acid(s) (e.g., 2 ⁇ azaphenylalanine, 3 ⁇ azaphenylalanine, 4 ⁇ azaphenylalanine, or 4 ⁇ fluorophenylalanine).
- a natural amino acid that is to be replaced e.g., phenylalanine
- non ⁇ naturally occurring amino acid(s) e.g., 2 ⁇ azaphenylalanine, 3 ⁇ azaphenylalanine, 4 ⁇ azaphenylalanine, or 4 ⁇ fluorophenylalanine.
- the non ⁇ naturally occurring amino acid is incorporated into the polypeptide in place of its natural counterpart. See, Koide et al., Biochem. 33:7470 ⁇ 6, 1994.
- Naturally occurring amino acid residues can be converted to non ⁇ naturally occurring species by in vitro chemical modification.
- Chemical modification can be combined with site ⁇ directed mutagenesis to further expand the range of substitutions (Wynn and Richards, Protein Sci. 2:395 ⁇ 403, 1993).
- a limited number of non ⁇ conservative amino acids, amino acids that are not encoded by the genetic code, non ⁇ naturally occurring amino acids, and unnatural amino acids may be substituted for amino acid residues of polypeptides of the present invention.
- Essential amino acids in the polypeptides of the present invention can be identified according to procedures known in the art, such as site ⁇ directed mutagenesis or alanine ⁇ scanning mutagenesis (Cunningham and Wells, Science 244: 1081 ⁇ 5, 1989).
- Sites of biological interaction can also be determined by physical analysis of structure, as determined by such techniques as nuclear magnetic resonance, crystallography, electron diffraction or photoaffinity labeling, in conjunction with mutation of putative contact site amino acids. See, for example, de Vos et al., Science 255:306 ⁇ 12, 1992; Smith et al., J. Mol. Biol. 224:899 ⁇ 904, 1992; Wlodaver et al., FEBS Lett. 309:59 ⁇ 64, 1992.
- the identities of essential amino acids can also be inferred from analysis of homologies with related components (e.g. the translocation or protease components) of the polypeptides of the present invention.
- phage display e.g., Lowman et al., Biochem. 30:10832 ⁇ 7, 1991; Ladner et al., U.S. Patent No. 5,223,409; Huse, WIPO Publication WO 92/06204
- region ⁇ directed mutagenesis e.g., region ⁇ directed mutagenesis
- Exon 1 of the SFTPB gene (as described by NG_016967.1) is dash ⁇ underlined (corresponding to bases 5543 ⁇ 5565 of NG_016967.1).
- the first A of this exon is the starting nucleotide of the mRNA generated by the SFTPB promoter.
- the bold and italicised ATG (corresponding to bases 5559 ⁇ 5561 of NG_016967.1) encodes the first methionine of pre ⁇ pro ⁇ SFTPB.
- 3’ of the double ⁇ underlined exon 1 is a partial portion of intron 1 (corresponding to bases 5626 ⁇ 5816 of NG_016967.1).
- the first base (base 81 of SEQ ID NO: 2) and last base (base 710 of SEQ ID NO: 2) of the core SFTP promoter fragment of the invention are double ⁇ underlined.
- the wavy ⁇ underlined sequence is the predicted TATA box of the SFTPB promoter.
- the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 76).
- SEQ ID NO: 42 VEGFA ⁇ 2 Enhancer & core SFTPB Promoter (1316bp)* gaggcacaaa gcgatcccca tcactgctcc acaatcattc attagctaac aagacagagc 60 agctcataaa aaaaagccg ttaaaaaat tccggggaaaaaagcag gaggtgatgc 120 aagccctggt taacaaggc tgagggttgg ggggaggcat gagagggtgt gagtggaata 180 acccaagcct gataagccac aaagcagcgc ctgctgctcccccccccc attag
- the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 77).
- SEQ ID NO: 43 CpG ⁇ Free CMV Enhancer & core SFTPB Promoter (938bp)* gttacataac ttatggtaaa tggcctgcct ggctgactgc ccaatgaccc ctgcccaatg 60 atgtcaataa tgatgtatgt tcccatgtaa tgccaatagg gactttccat tgatgtcaat 120 gggtggagta tttatggtaa ctgcccactt ggcagtacat caagtgtatc atatgccaag 180 tatgcccct attgatgtca atgatggtaa atggcctgcc tggcat
- the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 78).
- SEQ ID NO: 44 ELF3 Enhancer & core SFTPB Promoter (1290bp)* tagggatggg ccgaggctgg cactgatgct agacttccgt gcacagggca agtatggaca 60 agccccaagt ggctttgtga ggcccacaca gtgaagcttg ggaaatggga agtggggctg 120 cgcccagatt ctggtatcta tgacaactaa ggccgctgca catcctcatg gctctcccag 180 agacctcagg tgaggccctt ctgtgtgtcct caagcaccca
- the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 79).
- SEQ ID NO: 45 SV40 Enhancer & mSPB Promoter (873bp)* cgatggagcg gagaatgggc ggaactgggc ggagttaggg gcgggatggg cggagttagg 60 ggcgggacta tggttgctga ctaattgaga tgcatgctttt gcatacttct gcctgctggg 120 gagcctgggg actttccaca cctggttgct gactaattga gatgcatgct tgcatactt 180 ctgcctgctg gggagcctgg ggactttcca caccctaact gacacacatt ccacagc
- the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 80).
- SEQ ID NO: 46 Alv ⁇ 01 (CMV forward enhancer + mSFPB promoter) (955bp) AGATCTGTTACATAACTTATGGTAAATGGCCTGCCTGGCTGACTGCCCAATGACCCCTGCCCAATGATGTCAAT AATGATGTATGTTCCCATGTAATGCCAATAGGGACTTTCCATTGATGTCAATGGGTGGAGTATTTATGGTAACT GCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTATGCCCCCTATTGATGTCAATGATGGTAAATGGCC TGCCTGGCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTATGTATTAGTCATTGC TATTACCATGGATCTTATAGGAGCCACTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGG GT
- the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 81).
- the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 82).
- SEQ ID NO: 48 Alv ⁇ 03 (VEGFA ⁇ 2 forward enhancer + mSFPB promoter) (1333bp) AGATCTGAGGCACAAAGCGATCCCCATCACTGCTCCACAATCATTCATTAGCTAACAAGACAGAGCAGCTCAT AAAAAAAAAGCCGTTAAAAAAATTCCGGGGAAATGGAAAGCAGGAGGTGATGCAAGCCCTGGTTAACAAAG GCTGAGGGTTGGGGGGAGGCATGAGAGGGTGTGAGTGGAATAACCCAAGCCTGATAAGCCACAAAGCAGC GCCTGCTGCTCCCTCCTCCCTCTGCCGCCTGAGTCAGAAGCCGGGATGTGTTCAAAATCAAGCAATGTAATT AGCAGGTTGCATCATCATGCCTCTCGATTATAAATTACAATGGCCACAAAGAGGGTGGATGGAGGAGCAGGGATT AGATGGGGTAACCCAA
- the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 83).
- SEQ ID NO: 49 ATTTGAGCTCTTCTTTCTGCTGAACCATCG Underlined sequence is an SacI restriction enzyme site
- SEQ ID NO: 50 SFTPB promoter reverse primer TCTTAGATCTGTCAGACAGCTCTGGGTTCC Underlined sequence is a BgIII restriction enzyme site
- SEQ ID NO: 51 Exemplary linker between SFTPB promoter fragment and enhancer agatct
- SEQ ID NO: 52 Exemplified WPRE component (mWPRE) 1 GGGCCCAATC AACCTCTGGA TTACAAAATT TGTGAAAGAT TGACTGGTAT TCTTAACTAT 61 GTTGCTCCTT TTACGCTATG TGGATACGCT GCTTTAATGC CTTTGTATCA TGCTATTGCT 121 TCCCGTATGG CTTTCATTTT CTCCTCCTTG TATAAATCCT
- Example 1 Design and in vitro validation of a cell ⁇ specific core SFPB (mSPB) promoter sequence
- mSPB cell ⁇ specific core SFPB
- SFTPB promoter The expression driven by the core SFTPB promoter was compared to the full length SFTPB promoter (the 972bp fragment of SEQ ID NO: 2, fSPB) using a human surfactant air ⁇ liquid interface (SALI) culture model; an in vitro cell culture model that robustly recapitulates human ATII cells in primary cell culture (Munis et al. (2021) Molecular Therapy: Methods & Clinical Development 20: 237 ⁇ 246). H441 cells, when grown under SALI culture conditions, successfully mimic key characteristics of primary ATII cells. Briefly, SALI cultures were established by culturing cells in 12 ⁇ well Transwell inserts.
- SALI human surfactant air ⁇ liquid interface
- the polarization medium comprised either RPMI ⁇ 1640 (H441s cells and co ⁇ culture) or F12 ⁇ K (A549s) supplemented with 2 mM l ⁇ glutamine, 50 U/mL penicillin, 50 mg/mL streptomycin, 1% insulin ⁇ transferrin ⁇ selenium (GIBCO), 4% FCS, and 1 ⁇ M dexamethasone (Sigma). Media were subsequently changed three times per week throughout the described experiments. Cells were grown under SALI conditions for 14 days after air ⁇ lift prior to experimentation unless otherwise stated.
- the negative control sample (Na ⁇ ve; non ⁇ transduced cells H441 cells) shows a background level of 0.4%.
- the % EGFP ⁇ positive cells observed with the widely ⁇ used, non ⁇ specific CMV [SEQ ID NO: 6], EF1aS [SEQ ID NO; 7], and PGK [SEQ ID NO: 8] promoters is relatively high (17 ⁇ 22%) in line with expectations.
- the % EGFP ⁇ positive cells observed for the hCEF promoter [SEQ ID NO: 5], which has been used previously in the lungs of patients, is 11%.
- mSPB SEQ ID NO: 1
- fSPB SEQ ID NO: 4
- mSPC SEQ ID NO: 9
- fSPC fSPC
- human HEK293T cells were transduced with recombinant SIV lentiviral vectors pseudotyped with VSV ⁇ G and expressing the EGFP transgene from these same promoters (CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), EF1aS (SEQ ID NO: 7), PGK (SEQ ID NO: 8), the core SFTPB fragment (mSPB) (SEQ ID NO: 1), full ⁇ length SPB (fSPB) (SEQ ID NO: 4), minimal SPC (mSPC) (SEQ ID NO: 9) and full ⁇ length SPC (fSPC) (SEQ ID NO: 10)) at a multiplicity of infection (moi) of 0.5 or 1.0.
- CMV SEQ ID NO: 6
- hCEF SEQ ID NO: 5
- EF1aS SEQ ID NO: 7
- PGK SEQ ID NO: 8
- the core SFTPB fragment mSPB
- fSPB full ⁇ length
- the % EGFP ⁇ positive cells is low for mSPB, fSPB, mSPC and fSPC promoter sequences in generic HEK293T cells, indicating that the activity of these may be cell ⁇ specific.
- the mSPB promoter drives very low level expression in HEK293T cells (see, Figure 2).
- the expression driven by the mSPB promoter is likely to be specific to lung parenchymal cells, particularly ATII cells.
- Example 2 The core SFPB (mSPB) promoter sequence drives cell ⁇ specific gene expression in vivo
- the specificity of the mSPB promoter was further investigated using an in vivo mouse model. As shown in Figure 3A, on day 0 mice were dosed by nasal instillation with SIV vector. 7 days after dosing, lung tissue was harvested, fixed, frozen and cryosectioned. The sections were then analysed for expression of EGFP.
- mSPB core SFTPB promoter fragment
- mSPB core SFTPB promoter fragment
- Example 3 Design and production of improved mSPB promoters
- the inventors then sought to further increase transgene expression by and/or activity of (i.e. the number of lung parenchyma cells in which the promoter is active) the mSPB promoter, whilst retaining specificity for the lung parenchyma.
- the inventors generated a panel of improved mSPB promoters, each comprising mSPB and an enhancer.
- Lung ⁇ specific enhancers were selected using the ATII cell gene expression database from LungGENS and a tissue expression database.
- tissue expression databases including https://research.cchmc.org/pbge/lunggens/default.html and https://tissues.jensenlab.org/Search were interrogated to identify genes with high mean levels of expression in ATII cells.
- Enhancer regions were selected (and transcription factor binding sites identified) using University of California at Santa Cruz (UCSC) genome browser https://genome.ucsc.edu.
- the identified enhancer sequences were cloned into a construct comprising mSPB and an operably linked EGFP2ALux transgene.
- enhancers were added to the constructs in the forward (F) and reverse (R) direction. Schematics of exemplary constructs are shown in Figure 6, with the mSPB promoter sequence of SEQ ID NO: 1 and enhancer sequences as per SEQ ID NOs: 13 to 20, 23 to 33 and 36 to 40.
- Similar reporter transgene expression constructs were also generated incorporating the commonly used enhancers hB ⁇ Actin (SEQ ID NOs: 21 and 22), SV40 (SEQ ID NOs: 34 and 35) and CMV (SEQ ID NOs: 11 and 12).
- the expression constructs were incorporated into recombinant lentiviral vectors.
- the various lentiviral vectors were prepared for analysis of EGFP or Lux reporter transgene expression from each enhancer/promoter combination.
- Human HEK293T cells, murine LA ⁇ 4 cells and human SALI cells were transfected with the mSPB promoter constructs to determine whether the addition or the enhancer (or the orientation) increased gene expression.
- mSPB did not significantly increase gene expression relative to the na ⁇ ve control. Differences in expression driven by mSPB compared with mSPB + the different enhancers was also not significant. In contrast, significant differences in gene expression are seen in murine LA ⁇ 4 cells (see, Figure 8) and human SALI cells (see, Figure 9).
- murine LA ⁇ 4 cells the addition of a CMV enhancer (forwards or reverse), ELF3 enhancer (forwards), SV40 enhancer (forwards), and two of the VEGFA enhancers – VEG1 (forwards or reverse) and VEG2 (forwards) to the mSPB promoter significantly increased expression relative to the mSPB promoter alone.
- Example 4 Further characterisation of improved mSPB promoters Using the results of Example 3, the following seven candidate enhancers were selected for further screening: SV40, CMV, SLC34A2 ⁇ 1, SLC34A2 ⁇ 2, VEGFA ⁇ 2, ELF3 ⁇ 1 and ELF3 ⁇ 2.
- hCEF SEQ ID NO: 5
- CMV SEQ ID NO: 6
- hCEF SEQ ID NO: 5
- CMV SEQ ID NO: 6
- Lentivirus was dosed intranasally (1x10 7 TU per mouse in 100uL TSSM Buffer). On days 2, 7 and 28 post ⁇ dosing, all mice were anaesthetised and imaged for luciferase expression.
- Figure 12 shows the expression of luciferase as a heat map. Expression driven by hCEF or CMV promoters is shown as a reference. The inventors selected promoter which drive expression in the lungs, but not the nose. Localised expression in the lungs is important, because the nose has cells that are similar to (airway epithelial cells ciliated/non ⁇ ciliated epithelial) to those in the lungs, so it is important to determine that the transgene will not be expressed in the nose.
- FIG. 13A The expression of luciferase is quantified in Figure 13A.
- the ratio of expression in the lungs and the nose is shown in Figure 13B.
- the level of luciferase signal (Figure 13A) was greatest with the CMVenh group (SEQ ID NO: 11) (**** compared to mSPB (SEQ ID NO: 1) although the specificity for lung expression (Figure 13B), determined by the ratio of signal in the lung and nose (L ⁇ to ⁇ N), was not significantly different from CMV (SEQ ID NO: 6)or hCEF (SEQ ID NO: 5) at day 28.
- transgene expression from these constructs delivered as a non ⁇ viral (plasmid) formulation were similar to those obtained following delivery with viral (recombinant lentiviral) vectors.
- the enhancer sequences CMVenh (SEQ ID NO: 11), Slc2 (SEQ ID NO: 31) and VEGFA ⁇ 2 (SEQ ID NO: 38) combined with the mSPB (SEQ ID NO: 1) promoter sequence were assigned the nomenclature Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 respectively, as shown in SEQ ID NOs: 46, 47 and 48 respectively.
- Mouse lungs were processed for cryosections and imaging. Cryosections (7 ⁇ M) were subject to immunohistochemistry using primary antibodies to detect colocalization of EGFP and Pro/Mature Surfactant Protein ⁇ B, the latter being a marker for ATII cells.
- this experiment confirmed that the mSPB promoter/enhancer combinations tested drove transgene expression primarily in cells of the parenchyma, making them attractive candidates for targeted gene therapy of diseases resulting from or associated with deficiency of expression and/or expression of defective proteins in lung parenchymal cells, such as surfactant deficiencies.
- Repetition of this experiment using a different ATII cell marker, Surfactant Protein ⁇ C yielded the same pattern of expression, with Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 driving EGFP expression which colocalises with the ATII cell ⁇ specific marker SP ⁇ C (data not shown). Therefore, Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 have been reproducibly shown to drive expression in the lung parenchyma.
- Example 6 Improved mSPB promoters drive long ⁇ term expression in vivo
- Lentiviral vectors comprising mSPB, Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 were administered to intranasally. The mice were then monitored for luciferase expression.
- FIG. 16 shows the expression of luciferase as a heat map. Expression driven by hCEF or CMV promoters is shown as a reference. The expression of luciferase in the lungs over the time course of the experiment is quantified in Figure 16B. The area under the curve is shown in Figure 16C and the ratio of expression in the lungs and the nose is shown in Figure 16D.
- the results indicate that the Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 promoter drives high levels of expression in the lungs over at least a six ⁇ month period. Further, mSPB, Alv ⁇ 02 and Alv ⁇ 03 drive high levels of expression in the lungs, but not the nose.
- Example 7 – Improved mSPB promoters drive targeted expression in cells of the lung parenchyma Further analysis was conducted to confirm that Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 result in transgene expression in the target cells, particularly that expression is focussed in the cells of the lung parenchyma, rather than airway cells. This was investigated using immunohistochemistry.
- the hCEF promoter mainly drove expression in the cells of the airway, whereas as shown in Figure 17B, the mSPB promoter drove expression primarily in the parenchymal cells.
- the % of EGFP expression in the parenchymal cells for hCEF, mSPB, Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 was quantified.
- Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03, particularly Alv ⁇ 02 and Alv ⁇ 03 were able to drive EGFP expression in the parenchyma compared with the hCEF promoter.
- Example 8 Expression of SP ⁇ B restores transepithelial electrical resistance (TEER) in SFTPB knock out lung cells
- TEER transepithelial electrical resistance
- a Surfactant Air Liquid Interface (SALI) model was used with the human H441 lung cell line and the H441 SP ⁇ B KO cell line where the SFTPB gene had been ablated (Munis et al. (2021) Mo. Ther. Methods Clin. Dev 20:20:237 ⁇ 246, herein incorporated by reference).
- TEER transepithelial electrical resistance
- Example 9 Pulmonary surfactant has no effect on HEK293/T cell transduction with rSIV.F/HN vectors
- synthetic surfactant such as BLES or Curosurf
- rSIV.F/HN vector encoding EGFP under CMV promoter control was mixed 1:1 with TSSM (vehicle control) or with BLES or Curosurf and incubated at room temperature for 30mins. The mixtures were then diluted in OptiMEM ⁇ I (supplemented with polybrene for final 8 ⁇ g/mL working concentration) and used to transduce HEK293T cells.
- Example 10 Murine lung transduction with rSIV.F/HN vectors is at least as effective in the presence of a pulmonary surfactant Having surprisingly shown that pulmonary surfactants do not compromise the integrity of rSIV.F/HN lentiviral vectors in an in vitro setting, the effect of pulmonary surfactants on rSIV.FHN transduction of the murine lung was then investigated.
- rSIV.F/HN vector expressing firefly luciferase from the hCEF promoter was mixed 1:1 with TSSM (vehicle control) or BLES or Curosurf (to generate 2E8 TU in 100uL per dose).
- TSSM vehicle control
- BLES Curosurf
- TSSM diluent served as vehicle control.
- Mice were subjected to in vivo bioluminescent imaging to measure Firefly luciferase expression in the murine lungs.
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Genetics & Genomics (AREA)
- Engineering & Computer Science (AREA)
- Chemical & Material Sciences (AREA)
- Biotechnology (AREA)
- Biomedical Technology (AREA)
- General Engineering & Computer Science (AREA)
- Organic Chemistry (AREA)
- Molecular Biology (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Health & Medical Sciences (AREA)
- Biochemistry (AREA)
- Public Health (AREA)
- Physics & Mathematics (AREA)
- Biophysics (AREA)
- Veterinary Medicine (AREA)
- Plant Pathology (AREA)
- Animal Behavior & Ethology (AREA)
- Microbiology (AREA)
- Epidemiology (AREA)
- Pharmacology & Pharmacy (AREA)
- Medicinal Chemistry (AREA)
- Medicines That Contain Protein Lipid Enzymes And Other Medicines (AREA)
- Medicines Containing Material From Animals Or Micro-Organisms (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
- Medicinal Preparation (AREA)
- Pharmaceuticals Containing Other Organic And Inorganic Compounds (AREA)
Abstract
The present invention relates to nucleic acid cassettes for gene therapy, particularly to promoter and promoter/enhancer combinations for improved expression of transgenes in a lung parenchyma‐specific/preferred manner. The invention further relates nucleic acid cassettes comprising said promoters and promoter/enhancer combinations, viral and non‐viral vectors comprising such nucleic acid cassettes, and the use of such nucleic acid cassettes and vectors to increase expression of therapeutic proteins by lung parenchyma cells.
Description
SYNTHETIC PROMOTERS FIELD OF THE INVENTION The present invention relates to nucleic acid cassettes for gene therapy, particularly to promoter and promoter/enhancer combinations for improved expression of transgenes in a lung parenchyma‐specific/preferred manner. The invention further relates nucleic acid cassettes comprising said promoters and promoter/enhancer combinations, viral and non‐viral vectors comprising such nucleic acid cassettes, and the use of such nucleic acid cassettes and vectors to increase expression of therapeutic proteins by lung parenchyma cells. BACKGROUND TO THE INVENTION Surfactant protein B (SFTPB) deficiency is a severe monogenic interstitial lung disorder that leads to loss of life in infants as a result of alveolar collapse and respiratory distress syndrome. The only curative treatment is thought to be lung transplantation; however, the lack of suitable donor organs makes this a non‐viable option in most circumstances. The use of nucleic acids as medicine, or gene therapy, is a promising new treatment modality, both for SFTPB deficiency and other genetic diseases, including genetic diseases of the respiratory tract. The reason many gene therapies currently in use or under development are not effective at curing diseases is because it is difficult to make sufficient protein to reach the therapeutic threshold needed to treat or cure the disease. As such, generating sufficient gene expression in a given target cell is a major barrier to the success of many gene therapies. One approach to reaching the large doses needed for gene therapy to be successful is to administer massive amounts of the gene therapy to the patient, over 1 trillion viruses per kg of body mass. Producing so much virus is expensive, contributing to the $ 1,000,000 USD cost of gene therapies, and giving so much virus to a person can trigger immune responses that threaten the health of the patient and the efficacy of the therapy. To circumvent these problems, research to‐date has focused on gain of function mutations resulting in more potent proteins. Such an approach has been used previously in the gene therapies for haemophilia B (the Padua mutation in Factor IX) and lipoprotein lipase deficiency (the S447X variant of lipoprotein lipase). However, such gain‐of‐function mutations are not available for most gene therapies. Furthermore, and even with gain‐of‐function mutations, high doses of the gene therapy vector were still necessary to make the treatment effective. Therefore, such gain‐of‐function mutations along do not adequately address the exiting problems associated with producing sufficient quantities of vector, or the unwanted and clinically dangerous side effects associated with the large doses required.
The present inventors have previously developed a lentiviral vector, which has been pseudotyped with hemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, comprising a promoter and a transgene. Typically, the backbone of the vector is from a simian immunodeficiency virus (SIV), such as SIV1 or African green monkey SIV (SIV‐AGM). Preferably the backbone of a viral vector of the invention is from SIV‐AGM. The HN and F proteins function, respectively, to attach to sialic acids and mediate cell fusion for vector entry to target cells. The present inventors discovered that this specifically F/HN‐pseudotyped lentiviral vector can efficiently transduce airway epithelium, resulting in transgene expression sustained for periods beyond the proposed lifespan of airway epithelial cells. Importantly, the present inventors also found that re‐administration does not result in a loss of efficacy. These features make the vectors of the present invention attractive candidates for treating diseases via their use in expressing therapeutic proteins: (i) within the cells of the respiratory tract; (ii) secreted into the lumen of the respiratory tract; and (iii) secreted into the circulatory system. However, even using this state‐of‐the‐art platform technology, the levels of transgene expressed are at the lower predicted threshold required for clinical efficacy. There are other approaches which can also potentially increase the expression of therapeutic proteins by gene therapy vectors. For example, the use of exogenous signal peptides can be used to increase expression and secretion of therapeutic proteins by airway cells. By using the exogenous signal peptides it is possible to produce more protein for every copy of a gene therapy vector or transgene that is put into a cell, increasing the dose of therapeutic protein without increasing the amount of gene therapy vector given to a patient. However, the use of exogenous signal peptides is not appropriate for all therapeutic proteins or all conditions. For example, not all therapeutic proteins are secreted, and for some there may be clinical reasons why manipulating the signal peptides is undesirable. Generally, gene therapy vectors comprising promoters providing high and sustained gene expression in a variety of cell types are preferred, especially in a therapeutic context. For this reason, the inventors previously used a hCEF promoter to drive strong and persistent expression in mouse lung with non‐viral formulations. However for some conditions, expression in a specific tissue or cell type is required to (i) achieve desired therapeutic target, (ii) avoid gene expression‐related toxicities, and (iii) circumvent immune responses to the therapeutic agent stemming from gene expression in undesired cell types. In particular, it would be desirable to express surfactants in a tissue‐ or cell‐ specific manner, and with similar levels of gene expression to that of ubiquitously used strong promoters, for the treatment of genetic diseases, particularly genetic respiratory diseases, such as SFTPB deficiency.
There is therefore an unmet clinical need for new technologies to improve the tissue‐specific and/or cell‐specific expression of transgenes from a gene therapy vector. It is an object of the invention to address one or more of these problems. In particular, it is an object of the invention to provide new promoters, nucleic acid cassettes and gene therapy vectors which enable high expression of transgenes in the lung parenchyma. Such promoters, cassettes and vectors may be of particular use in the treatment of genetic diseases, particularly genetic respiratory diseases, such as surfactant protein deficiency. SUMMARY OF THE INVENTION At present, there remains a pressing need for technology that enables high levels of cell‐ or tissue‐specific expression of transgenes for gene therapy, including from the inventors’ own lentiviral platform. In particular, there is a need for novel promoters that drive cell‐specific expression in the lung parenchyma. The present inventors have developed a panel of novel promoter sequences comprising a functional fragment of the human SFTPB gene promoter. As demonstrated herein, the inventors have surprisingly shown that the SFTPB promoter fragment of the invention drives increased transgene expression compared with the full‐length SFTPB promoter. Furthermore, the inventors have also surprisingly demonstrated that the SFTPB promoter fragment of the invention can be combined with particular enhancers, including some enhancers which are not associated with cell‐specific expression in the lung parenchyma, to further improve transgene expression. Accordingly, the present invention provides an SFTPB promoter fragment which comprises or consists of at least part of exon 1 of the SFTPB gene, but does not comprise the SFTPB gene start codon. Typically said promoter is less than 800 bases in length, preferably less than 700 bases in length. Said promoter may comprise or consist of: (a) SEQ ID NO: 1 or a sequence with at least 80% identity to SEQ ID NO: 1; or (b) bases 81‐710 of SEQ ID NO: 2, or a sequence with at least 80% identity to bases 81‐710 of SEQ ID NO: 2; wherein optionally (i) said SFTPB promoter fragment further comprises up to 20 bases at the 5’ end, which may optionally correspond to up to 20 bases 5’ to base 81 of SEQ ID NO: 2; and/or (ii) said SFTPB promoter fragment further comprises up to 4 bases at the 3’ end, which may optionally correspond to up to 4 bases 3’ to base 710 of SEQ ID NO: 2. The SFTPB promoter fragment may comprise or consist of SEQ ID NO: 1, or a sequence with at least 90% identity to SEQ ID NO: 1. The SFTPB promoter fragment of the invention may further comprise a 5’ enhancer. Said enhancer may be in (i) the forward, or (ii) the reverse, orientation; preferably wherein the enhancer is in the forward orientation. Said enhancer may be selected from a SLC34A2 enhancer, a VEGFA
enhancer, a CMV enhancer, an SV40 enhancer, an ELF3 enhancer, an actin enhancer, an LMO7 enhancer, a SFTPC enhancer or a SFTPB enhancer. Preferably the enhancer may be selected from a SLC34A2 enhancer, a VEGFA enhancer or a CMV enhancer. The enhancer may be selected from (a) an SLC34A2 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 29‐33, preferably SEQ ID NO: 31; (b) a VEGFA enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38; (c) a CMV enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11; (d) a SV40 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or 35; (e) an ELF3 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 13‐20; (f) an actin enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 21 or 22; (g) an LMO7 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 23—25; (h) an SFTPC enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 27 or 28; or (i) an SFTPB enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 26. The SFTPB promoter fragment of the invention may comprise or consist of a nucleic acid sequence of any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 49, or a sequence with at least 90% identity to any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 48. Preferably the SFTPB promoter fragment of the invention comprises or consists of a nucleic acid sequence of any one of SEQ ID NOs: 46, 47 or 48, or a sequence with at least 90% identity to any one of SEQ ID NOs: 46, 47 or 48. The invention also provides a nucleic acid cassette comprising: (a) an SFTPB promoter fragment of the invention; and (b) a transgene. Said transgene may encode a therapeutic protein selected from: (a) secreted therapeutic protein selected from: Surfactant Protein B (SP‐B), Surfactant Protein C (SP‐C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte‐Macrophage Colony‐Stimulating Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody, an anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP‐binding cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB. The SFTPB promoter fragment may increase expression of the transgene by lung parenchymal cells, optionally compared with the full‐length SFTPB promoter. Expression of the transgene may be increased by at least 2‐fold, preferably at least 5‐fold compared with the full‐length SFTPB promoter. The lung parenchymal cells may comprise one or more cell type selected from: alveolar type I epithelial (ATI) cells, alveolar type II epithelial cells (ATII), and/or club cells, preferably ATII and/or ATI cells.
Expression of the transgene by a nucleic acid cassette or promoter of the invention may be specific to lung parenchymal cells; and/or the ratio of lung expression: nose expression of the transgene by the SFTPB promoter fragment is at least 2:1. The invention further provides a gene therapy vector, comprising a nucleic acid cassette of the invention. Said gene therapy vector may be a non‐viral vector, wherein optionally: (a) the non‐ viral vector is a plasmid; and/or (b) the non‐viral vector is comprised in a cationic liposome, which preferably comprises GL67A. Said gene therapy vector may be a viral vector, optionally selected from: a lentiviral vector; an AAV vector; and an adenoviral vector. Said lentiviral vector may be pseudotyped with haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, optionally from a Sendai virus. Said lentiviral vector may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector. Preferably said lentiviral vector is a SIV vector. The invention further provides a method of expressing a therapeutic protein in a target cell, comprising delivering a nucleic acid cassette of the invention or a gene therapy vector of the invention into the target cells. Said delivering may comprise integrating said nucleic acid cassette or gene therapy vector into said target cell's genome. The invention also provides a gene therapy vector of the invention for use in a method of treating a disease. Said disease may be a genetic disease. The disease may be: (a) a respiratory disease, particularly a genetic respiratory disease; or (b) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder. The disease may be selected from Surfactant Protein B (SP‐B) Deficiency; Surfactant Protein C (SP‐C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1‐antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID‐19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia. The invention further provides a cell comprising a nucleic acid cassette of the invention or a gene therapy vector of the invention. The invention also provides a composition comprising a nucleic acid cassette of the invention or a gene therapy vector of the invention and a pharmaceutically acceptable carrier, diluent or excipient.
In addition, the inventors have shown for the first time that lentiviral vectors, such as the lentiviral vector pseudotyped with haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus as exemplified herein, can be co‐administered with a synthetic surfactant, and achieve at least as efficient transduction into target cells, and potentially even enhanced transduction, compared with transduction of the lentiviral vector in vehicle alone. Whilst it is known in the art that AAV vectors can be administered with surfactant, it is surprising that this is possible with lentiviral vectors due to fundamental structural differences between AAV and lentiviral vectors. In particular, lentiviral vectors are surrounded by a lipid envelope, whereas AAV are not. It would therefore be expected that the lentiviral vector envelope would be disrupted by the hydrophobic portions of a surfactant, having a negative effect on the lentiviral structure. Accordingly, the invention also provides lentiviral vector pseudotyped with haemagglutinin‐ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered simultaneously or sequentially with a surfactant. Optionally said disease is: (a) a genetic disease; (b) a respiratory disease, particularly a genetic respiratory disease; (c) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder; and/or (d) selected from Surfactant Protein B (SP‐B) Deficiency; Surfactant Protein C (SP‐C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1‐antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID‐19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia. In such a lentiviral vector, the respiratory paramyxovirus may be a Sendai virus; and/or the lentiviral vector may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector, preferably a SIV vector The transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP‐B), Surfactant Protein C (SP‐C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte‐Macrophage Colony‐Stimulating Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody, an anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP‐binding cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.
In such a lentiviral vector, the promoter may comprise a SFTPB promoter fragment as defined herein; and/or the lentiviral vector may comprise a nucleic acid cassette as defined herein. The lentiviral vector may be administered before the surfactant. The surfactant may be administered before the lentiviral vector. The lentiviral vector and surfactant may be administered simultaneously, optionally wherein the lentiviral vector and surfactant are mixed prior to administration. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1: Flow cytometry plots illustrating EGFP expression levels in human SALI cultures transduced with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing the EGFP transgene at a dose of 1x106 Transducing Units (TU) from a range of promoters (CMV (SEQ ID NO: 6), hCEF ( (SEQ ID NO: 5), EF1aS ( (SEQ ID NO: 7), PGK ( (SEQ ID NO: 8), mSPB ( (SEQ ID NO: 1), fSPB (SEQ ID NO: 4), mSPC (SEQ ID NO: 9), fSPC (SEQ ID NO: 10)). Figure 2: Graphs quantifying EGFP expression in human HEK293T cells transduced with recombinant SIV lentiviral vectors pseudotyped with VSV‐G and expressing the EGFP transgene from a range of promoters (CMV (SEQ ID NO: 6), hCEF ( (SEQ ID NO: 5), EF1aS ( (SEQ ID NO: 7), PGK ( (SEQ ID NO: 8), mSPB ( (SEQ ID NO: 1), fSPB (SEQ ID NO: 4), mSPC (SEQ ID NO: 9), fSPC(SEQ ID NO: 10)) at a multiplicity of infection (moi) of 0.5 or 1.0. The % EGFP‐positive cells and the mean fluorescence intensity (MFI) at both moi are plotted. The negative control sample (Mock; non‐transduced, treated with buffer‐ only) represents the background levels. Figure 3: Experimental schematics for experiments to test EGFP expression in vivo using lentiviral vectors expressing EGFP from different promoters. A Schematic showing timing of dosing, on Day 0 mice were dosed with lentiviral vectors via nasal instillation and 7 days post‐dosing the mice were culled and lung tissue harvested for cryosections. B Schematic showing treatment groups. Female BALB/c mice (4‐6 weeks; n=3 per group) were dosed on Day 0 with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing the EGFP transgene (1x106 TU in 100µL TSSM Buffer) expressing EGFP from either the mSPB (SEQ ID NO: 1) or fSPB (SEQ ID NO: 4) promoter sequence (mSPB and fSPB groups, respectively). Female BALB/c mice (n=2) were treated with 100 µL TSSM buffer only (naïve group) as a negative control.
Figure 4: Representative immunohistochemistry images showing expression of EGFP in lung tissue samples taken from mice treated with lentiviral vectors expressing EGFP from different promoters. Images of native EFGP fluorescence (n=6 per mouse) were analysed in lung cryosections of mice dosed as shown in Figure 3. EGFP fluorescence in lung sections was analysed for the mSPB (SEQ ID NO: 1) and fSPB (SEQ ID NO: 4) groups and compared with mice in the naïve group imaged in parallel. The observed fluorescence appeared ‘punctate’ and scattered through the lung parenchyma (Arrows = foci corresponding to EGFP expression), which is indicative of expression in ATII cells. Very little EGFP fluorescence was visible in the airway epithelia with these promoters, which contrasts with previously observed EGFP expression from promiscuous promoters. Figure 5: Panel of ATII specific genes (and enhancers) generated by interrogation of LungGENS and a tissue expression database. Figure 6: Schematics of exemplary mSPB/enhancer constructs. Different lengths (indicated in base pairs (bp)) of the newly identified enhancer sequences (boxes) were sub‐cloned in front of the mSPB (SEQ ID NO: 1) promoter sequence (arrow boxes) in both forward (f) and reverse (r) orientations to generate candidate synthetic promoters expressing the EGFP2ALux reporter transgene. Similar reporter transgene expression constructs were also generated incorporating the commonly used enhancers hB‐Actin, SV40 and CMV. Some f and r permutations were not constructed. Figure 7: Graph showing expression of EGFP in HEK293T cells by different mSPB/enhancer constructs. Recombinant SIV lentiviral vectors pseudotyped with VSV‐G and expressing EGFP2ALux transgene from each of the enhancer/promoter combinations, were produced at small scale. These were used to transduce human HEK293T cells in a 24‐well plate (seeded @1x105 cells/well) at a multiplicity of infection (moi) of approximately 0.1. Luciferase levels (Relative Light Units (RLU)) were determined in n=4 independent biological replicates. Naïve (non‐transduced) cells were used as a control to determine background RLU level. There were no statistically significant differences observed between naïve and mSPB (SEQ ID NO: 1), or between mSPB (SEQ ID NO: 1) and all other enhancer constructs (Kruskal‐Wallis with Dunn’s multiple comparison test). Figure 8: Graph showing expression of EGFP in murine LA‐4 cells by different mSPB/enhancer constructs. P values are given where significant expression was observed. Recombinant SIV lentiviral vectors pseudotyped with VSV‐G and expressing EGFP2ALux transgene from each of the enhancer/promoter combinations, were produced at small scale. These were used to transduce
murine LA‐4 cells in a 24‐well plate (1x105 cells/well) at a multiplicity of infection (moi) of approximately 0.1. Luciferase levels (Relative Light Units (RLU)) were determined in n=4 independent biological replicates. Naïve (non‐transduced) cells were used as a control to determine background RLU level. P values are given where statistically significant differences were observed (Kruskal‐Wallis with Dunn’s multiple comparison test). Figure 9: Graph showing expression of EGFP in human SALI cells by different mSPB/enhancer constructs. P values are given where significant expression was observed. Recombinant SIV lentiviral vectors pseudotyped with VSV‐G and expressing EGFP2ALux transgene from each of the enhancer/promoter combinations, were produced at small scale. These were used to transduce human SALI cultures in (1x105 Transducing Units/well) at a multiplicity of infection (moi) of approximately 0.1. Luciferase levels (Relative Light Units (RLU)) were determined in n=4 independent biological replicates. Naïve (non‐transduced) cells were used as a control to determine background RLU level. P values are given where statistically significant differences were observed (Kruskal‐Wallis with Dunn’s multiple comparison test). Figure 10: Graph showing expression of EGFP in human SALI cells by lentiviral vectors with different mSPB/enhancers in a secondary screen. P values are given where significant expression was observed. Recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing EGFP2ALux reporter transgene, were used to transduce human SALI cultures (n=6 replicates) in a repeat secondary screening experiment, which included the following promoter sequences: CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), and mSPB ( (SEQ ID NO: 1), as well as with selected enhancer/promoter combinations SV40 (f)[SEQ ID NO: 34], CMVenh (f) [SEQ ID NO: 11], Elf1 (f)[SEQ ID NO: 13], Elf2 (f)[SEQ ID NO: 15], Slc1 (f)[SEQ ID NO: 29], Slc2 (f)[SEQ ID NO: 31], and Veg2 (f)[SEQ ID NO: 38]. Luciferase expression levels (Relative Light Units (RLU)) were determined and compared with naïve (non‐transduced) control cells. Vector Copy Number (VCN) was also determined in the transduced cell samples (n=3). Results from six independent biological replicates are shown; P values are given where statistically significant differences were observed (Kruskal‐Wallis with Dunn’s multiple comparison test). Figure 11: Graph showing expression of EGFP in HEK293T cells by lentiviral vectors with different mSPB/enhancers in a secondary screen. P values are given where significant expression was observed. Recombinant SIV lentiviral vectors pseudotyped with the F/HN and expressing EGFP2ALux reporter transgene, were used to transduce human HEK293T cell cultures (n=6 replicates) in a repeat secondary screening experiment, which included the following selected enhancer/promoter sequences: CMV
(SEQ ID NO: 6), hCEF (SEQ ID NO: 5), and mSPB ( (SEQ ID NO: 1), as well as with selected enhancer/promoter combinations SV40 (f)[SEQ ID NO: 34], CMVenh (f) [SEQ ID NO: 11], Elf1 (f)[SEQ ID NO: 13], Elf2 (f)[SEQ ID NO: 15], Slc1 (f)[SEQ ID NO: 29], Slc2 (f)[SEQ ID NO: 31], and Veg2 (f)[SEQ ID NO: 38]. Luciferase expression levels (Relative Light Units (RLU)) were determined and compared with naïve (non‐transduced) control cells. Vector Copy Number (VCN) was also determined in the transduced cell samples (n=3). Results from six independent biological replicates are shown; n.s: Results were not statistically significantly different (Kruskal‐Wallis with Dunn’s multiple comparison test). Figure 12: Heat map of luciferase expression in mice treated with lentiviral vectors with different mSPB/enhancers. Female BALB/c mice (6 weeks; n=5 per group) were dosed with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing EGFP2ALux reporter transgene from the following promoter sequences: CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), and mSPB ( (SEQ ID NO: 1), as well as with selected enhancer/promoter combinations SV40 (f)[SEQ ID NO: 34], CMVenh (f) [SEQ ID NO: 11], Elf1 (f)[SEQ ID NO: 13], Elf2 (f)[SEQ ID NO: 15], Slc1 (f)[SEQ ID NO: 29], Slc2 (f)[SEQ ID NO: 31], and Veg2 (f)[SEQ ID NO: 38]. Lentivirus was dosed intranasally (1x107 TU per mouse in 100uL TSSM Buffer). On days 2, 7 and 28 post‐dosing, all mice were anaesthetised and imaged for luciferase expression. Observed areas of luciferase positive signal correspond with luciferase transgene expression in the nose and/or lungs. A naïve group (n=3) of (non‐transduced) animals were included as a negative control and showed no luciferase signal (background levels). Luciferase signal was observed in the nose and lung areas at all timepoints with the non‐specific CMV[SEQ ID NO: 6] and hCEF[ SEQ ID NO: 5] lung promoter in line with expectations. Luciferase signal in the mSPB[SEQ IDN O: 1] group was overall lower and restricted to the lung, also as expected. Figure 13: Graphs showing in vivo luciferase expression by lentiviral vectors with different mSPB/enhancers (A) and the ratio of luciferase expression in the lungs and the nose (B). As described in Figure 12, female BALB/c mice (6 weeks; n=5 per group) were dosed with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing EGFP2ALux reporter transgene from the following enhancer/promoter sequences: CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), and mSPB ( (SEQ ID NO: 1), as well as the selected enhancer/promoter combinations SV40 (f)[SEQ ID NO: 34], CMVenh (f) [SEQ ID NO: 11], Elf1 (f)[SEQ ID NO: 13], Elf2 (f)[SEQ ID NO: 15], Slc1 (f)[SEQ ID NO: 29], Slc2 (f)[SEQ ID NO: 31], and Veg2 (f)[SEQ ID NO: 38]. Lentivirus was dosed intranasally (1x107 TU per mouse in 100uL TSSM Buffer). On days 2, 7 and 28 post‐dosing, mice were anaesthetised and imaged for luciferase expression. A) The day 28 luciferase signal (Radiance (photons/s/cm2/sr) and B) the ratio of in vivo
day 28 luciferase signal in the mouse lungs and nose, are plotted for each enhancer/mSPB combination tested. Figure 14: Graphs showing luciferase expression by plasmids with different mSPB/enhancers in human SALI cells (A), in vivo (B), and in HEK293T cells (C). The enhancer/mSPB promoter constructs expressing Lux reporter transgene were evaluated in the context of a non‐viral formulation, to deliver plasmids containing the CMV[SEQ ID NO: 6], hCEF[SEQ ID NO: 5] and mSPB[SEQ ID NO: 1] promoter sequences as well as the selected enhancer/mSPB promoter constructs: slc2[SEQ ID NO: 31], veg2[SEQ ID NO: 38], sv40[SEQ ID NO: 34], and CMVenh[SEQ ID NO: 11]. (A) Human SALI cultures (n=8) were transfected with plasmid DNA (2ug plasmid per culture) complexed with linear polyethyleneimine (PEIPro; 100µL per culture) and luciferase activity in Relative Light Units (RLU) determined. A naïve group (n=8; non‐transfected) was included as a negative control. Luciferase activity in the veg2[SEQ ID NO: 38] and CMVenh[SEQ ID NO: 11] groups were significantly different from mSPB [SEQ ID NO: 1] (** and ****, respectively; Kruskal‐Wallis). NB: a CMV[SEQ ID NO: 6] group was not included in this experiment. (B) Female BalbC mice (n=6 per group) were each dosed intranasally (100µL per mouse) with 60µg of plasmid complexed with 77.4µg 25KDa branched polyethyleneimine (PEI). After 72hours the lungs were harvested and processed to measure luciferase activity (RLU/mg protein) in lung homogenates. A naïve group (n=3; non‐transfected) was included as a negative control. (C) Human HEK293T cells were seeded (3x105 cells per well) and after 24 hours each well (n=4) was transfected with plasmid DNA (2µg plasmid per well) complexed with linear polyethyleneimine (PEIPro). After 48 hours cells were lysed and assayed for luciferase activity measured in duplicate and expressed as RLU per mg of protein. Figure 15: Representative immunohistochemistry images showing EGFP expression in lung parenchyma sections taken from mice treated with lentiviral vectors expressing EGFP from different enhancer/mSPB constructs. Representative images from (A) CMVenh (Alv‐01, SEQ ID NO: 46), (B) Slc2 (Alv‐02, SEQ ID NO: 47) and (C) Vegf2 (Alv‐03, SEQ ID NO: 48) groups are shown. On a background of (DAPI stained blue) cell nuclei, white arrows indicate examples of co‐localisation of (yellow) signal from EGFP transgene expression (green) and ATII cell‐specific marker SP‐B (red). Magnification: scale bar shown. Figure 16: To examine the level and duration of expression from these promoters, female BALB/c mice (n=10) were dosed with rSIV.F/HN encoding firefly luciferase under the control of CMV, hCEF, mSPB, Alv1, Alv2 or Alv3 promoter, or formulation buffer (TSSM) as a negative control (2.5e8 TU per
mouse via intranasal administration). (A)7, 14, 28 days, and 3 and 6 months after transduction, mice were administered D‐luciferin and imaged for luciferase activity in the lung and nasal cavity. Signal in regions of interest (ROI) capturing the chest and nasal areas were measured in photons/second/cm2/sr. Average radiance in the lung (B) in addition to area under the curve (AUC) analysis (C) were plotted. (D) Specificity for expression in the lung parenchyma was determined by calculating the ratio of signal in the lung (indicative of alveolar and airway cell transduction) to signal in the nasal cavity (indicative of airway cell transduction) 28 days after dosing. Significant differences were determined by Kruskall‐Wallis (H(6)=55.53, p<0.0001) with Dunn’s post hoc multiple comparison test comparing vector groups using the hCEF promoter. Data are shown as individual values and mean±SEM. A calculated p value of <0.05 was deemed significant (p<0.01 ,**; p<0.0001 ,****). Figure 17: The hCEF promoter drives expression in multiple lung cell types, including cells lining the airway epithelia, whereas the novel lung promoters mSPB, Alv‐1, Alv‐2 and Alv‐3 were designed to drive targeted expression in ATII cells in the parenchyma. To investigate the expression profile, BALB/c mice (n=5‐10 per group) were dosed with SIV.F/HN expressing EGFP from the hCEF, mSPB or Alv‐1, Alv‐2 or Alv‐3 promoters. Mice were culled on day 14 post‐dosing and whole lungs processed for cryosectioning. Sections were DAPI stained and imaged (using Axioscanner) to visualise native EGFP‐ positive cells. Representative images of EGFP positive cells in the airway (Aw) and parenchyma (P) are shown after administration of (A) rSIV.F/HN hCEF EGFP or (B) rSIV.F/HN mSP‐B EGFP. Images of lung sections were further analysed (using Visiopharm software) which required manual indication of airways and parenchyma. The percentage of EGFP‐positive cells from the total lung and from the airways was used to (C) estimate the percentage of EGFP‐positive cells observed in the parenchyma from each promoter. Collectively, the data presented indicate that the novel lung promoters (especially Alv‐2 and Alv‐3) show expression mainly in the parenchyma compared with the hCEF promoter. Figure 18: The effect of mSPB, Alv‐1, Alv‐2 and Alv‐3 driven SFP‐B expression on transepithelial electrical resistance (TEER) in a Surfactant Air Liquid Interface (SALI) model. (A & B) TEER values from SALI cultures generated from the H441 SP‐B KO cells were lower than the parental H441 cell line at 14 days post airlift constituting a phenotypic defect. (C) At 5 days post transduction, rSIV.F/HN expressing EGFP control vector did not increase the TEER in H441 SP‐B KO cells SALI cultures, however transduction with any/all of the vectors expressing SP‐B showed (D) correction of the TEER towards normal, calculated as a percentage of the mock‐transduced parental H441 cells. ns: statistically non‐ significant.
Figure 19: Graph showing the % transduction of HEK293T cells with rSIV.F/HN vector encoding EGFP under CMV promoter control was mixed with synthetic surfactants Beractat or Proactant alfa, or with TSSM (vehicle control). The presence of either synthetic surfactant had no significant effect on the % transduction of the HEK293T cells with the rSIV.F/HN vector. Figure 20: The effect of pulmonary surfactants on rSIV.FHN transduction of the murine lung. rSIV.F/HN vector expressing firefly luciferase from the hCEF promoter (rSIV.F/HN hCEF Flux) was mixed 1:1 with TSSM (vehicle control) or BLES or Curosurf (to generate 2E8 TU in 100uL per dose). (A) Time course of luciferase expression in the lungs of each mouse on days 7, 14, 28, 112 days post‐dosing. (B) Graph of luciferase activity in the lungs plotted as area under the curve (AUC) analysis (log10‐ transformed). (C) Graph of the fold difference in luciferase expression relative to mice dosed with control TSSM:vector. DETAILED DESCRIPTION OF THE INVENTION Definitions Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Singleton, et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY, 20 ED., John Wiley and Sons, New York (1994), and Hale & Marham, THE HARPER COLLINS DICTIONARY OF BIOLOGY, Harper Perennial, NY (1991) provide the skilled person with a general dictionary of many of the terms used in this disclosure. The meaning and scope of the terms should be clear; however, in the event of any latent ambiguity, definitions provided herein take precedent over any dictionary or extrinsic definition. It should be understood that this invention is not limited to the particular methodology, protocols, and reagents, etc., described herein and as such can vary. This disclosure is not limited by the exemplary methods and materials disclosed herein, and any methods and materials similar or equivalent to those described herein can be used in the practice or testing of embodiments of this disclosure. The terminology used herein is for the purpose of describing particular embodiments only, and is not intended to limit the scope of the present invention, which is defined solely by the claims. The description of embodiments of the disclosure is not intended to be exhaustive or to limit the disclosure to the precise form disclosed. While specific embodiments of, and examples for, the disclosure are described herein for illustrative purposes, various equivalent modifications are possible
within the scope of the disclosure, as those skilled in the relevant art will recognize. For example, while method steps or functions are presented in a given order, alternative embodiments may perform functions in a different order, or functions may be performed substantially concurrently. The teachings of the disclosure provided herein can be applied to other procedures or methods as appropriate. The various embodiments described herein can be combined to provide further embodiments. Aspects of the disclosure can be modified, if necessary, to employ the compositions, functions and concepts of the above references and application to provide yet further embodiments of the disclosure. Moreover, due to biological functional equivalency considerations, some changes can be made in protein structure without affecting the biological or chemical action in kind or amount. These and other changes can be made to the disclosure in light of the detailed description. All such modifications are intended to be included within the scope of the appended claims. The headings provided herein are not limitations of the various aspects or embodiments of this disclosure. As used herein, the term "capable of' when used with a verb, encompasses or means the action of the corresponding verb. For example, "capable of interacting" also means interacting, "capable of cleaving" also means cleaves, "capable of binding" also means binds and "capable of specifically targeting…" also means specifically targets. Numeric ranges are inclusive of the numbers defining the range. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within this disclosure. The upper and lower limits of these smaller ranges may independently be included or excluded in the range, and each range where either, neither or both limits are included in the smaller ranges is also encompassed within this disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in this disclosure. As used herein, the articles "a" and “an” may refer to one or to more than one (e.g. to at least one) of the grammatical object of the article. Further, unless otherwise required by context, singular terms shall include pluralities and plural terms shall include the singular. In this application, the use of "or" means "and/or" unless stated otherwise. Furthermore, the use of the term "including", as well as other forms, such as "includes" and "included", is not limiting. “About” may generally mean an acceptable degree of error for the quantity measured given the nature or precision of the measurements. Exemplary degrees of error are within 20 percent (%),
typically, within 10%, and more typically, within 5% of a given value or range of values. Preferably, the term “about” shall be understood herein as plus or minus (±) 5%, preferably ± 4%, ± 3%, ± 2%, ± 1%, ± 0.5%, ± 0.1%, of the numerical value of the number with which it is being used. The term "consisting of'' refers to compositions, methods, and respective components thereof as described herein, which are exclusive of any element not recited in that description of the invention. As used herein the term "consisting essentially of'' refers to those elements required for a given invention. The term permits the presence of elements that do not materially affect the basic and novel or functional characteristic(s) of that invention (i.e. inactive or non‐immunogenic ingredients). Embodiments described herein as “comprising” one or more features may also be considered as disclosure of the corresponding embodiments “consisting of” and/or “consisting essentially of” such features. Concentrations, amounts, volumes, percentages and other numerical values may be presented herein in a range format. It is also to be understood that such range format is used merely for convenience and brevity and should be interpreted flexibly to include not only the numerical values explicitly recited as the limits of the range but also to include all the individual numerical values or sub‐ranges encompassed within that range as if each numerical value and sub‐range is explicitly recited. A "vector" or "construct" (sometimes referred to as gene delivery or gene transfer "vehicle") refers to a macromolecule or complex of molecules comprising a polynucleotide to be delivered to a host cell, either in vitro or in vivo. A vector can be a linear or a circular molecule. A vector of the invention may be viral or non‐viral. All disclosure herein in relation vectors of the invention applies equally to viral and non‐viral vectors unless otherwise stated. All disclosure in relation to viral vectors of the invention applies equally and without reservation to lentiviral (e.g. SIV) vectors, particularly to lentiviral (e.g. SIV) vectors that are pseudotyped with hemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus (also referred to herein as SIV F/HN or SIV‐FHN). As used herein, the term "plasmid", refers to a common type of non‐viral vector. A plasmid is an extra‐chromosomal DNA molecule separate from the chromosomal DNA which is capable of replicating independently of the chromosomal DNA. Preferably a plasmid is circular and may be double‐stranded. The terms "nucleic acid cassette”, “nucleic acid construct", "expression cassette" and "nucleic acid expression cassette" are used interchangeably to mean a nucleic acid molecule that is capable of directing transcription. A nucleic acid cassette includes, at the least, a promoter or a structure
functionally equivalent to a promoter and a nucleic acid sequence to be transcribed. Thus, a nucleic acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter and a nucleic acid sequence encoding a protein of interest. In the present invention, a nucleic acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter, and a nucleic acid encoding a therapeutic protein. A nucleic acid cassette may include additional elements, such as an enhancer, and/or a transcription termination signal. As used herein, the terms “transduced” and “modified” are used interchangeably to describe cells which have been modified to express a transgene of interest. Typically the modification occurs through transduction of the cells. As used herein, the terms “titre” and “yield” are used interchangeably to mean the amount of viral (e.g. lentiviral, particularly SIV) vector produced by a method of the invention. Titre is the primary benchmark characterising manufacturing efficiency, with higher titres generally indicating that more vector is manufactured (e.g. using the same amount of reagents). Titre or yield may relate to the number of vector genomes that have integrated into the genome of a target cell (integration titre), which is a measure of “active” virus particles, i.e. the number of particles capable of transducing a cell. Transducing units (TU/mL also referred to as TTU/mL) is a biological readout of the number of host cells that get transduced under certain tissue culture/virus dilutions conditions, and is a measure of the number of “active” virus particles. The total number of (active+inactive) virus particles may also be determined using any appropriate means, such as by measuring either how much Gag is present in the test solution or how many copies of viral RNA are in the test solution. Assumptions are then made that a viral (e.g. lentivirus, particularly SIV) particle contains either 2000 Gag molecules or 2 viral RNA molecules. Once total particle number and a transducing titre/TU have been measured, a particle:infectivity ratio calculated. Amino acids are referred to herein using the name of the amino acid, the three‐letter abbreviation or the single letter abbreviation. Unless otherwise indicated, any nucleic acid sequences are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively. As used herein, the terms "protein" and "polypeptide" are used interchangeably herein to designate a series of amino acid residues, connected to each other by peptide bonds between the alpha‐amino and carboxyl groups of adjacent residues. The terms "protein", and "polypeptide" refer to a polymer of amino acids, including modified amino acids (e.g., phosphorylated, glycated, glycosylated, etc.) and amino acid analogues, regardless of its size or function. "Protein" and "polypeptide" are often used in reference to relatively large polypeptides, whereas the term "peptide"
is often used in reference to small polypeptides, but usage of these terms in the art overlaps. The terms "protein" and "polypeptide" are used interchangeably herein when referring to a gene product and fragments thereof. Thus, exemplary polypeptides or proteins include gene products, naturally occurring proteins, homologs, orthologs, paralogs, fragments and other equivalents, variants, fragments, and analogues of the foregoing. As used herein, the terms “polynucleotides”, "nucleic acid" and "nucleic acid sequence" refers to any molecule, preferably a polymeric molecule, incorporating units of ribonucleic acid, deoxyribonucleic acid or an analogue thereof. The nucleic acid can be either single‐stranded or double‐stranded. A single‐stranded nucleic acid can be one nucleic acid strand of a denatured double‐ stranded DNA Alternatively, it can be a single‐stranded nucleic acid not derived from any double‐ stranded DNA. In one aspect, the nucleic acid can be DNA. In another aspect, the nucleic acid can be RNA Suitable nucleic acid molecules are DNA, including genomic DNA or cDNA. Other suitable nucleic acid molecules are RNA, including siRNA, shRNA, and antisense oligonucleotides. The terms “transgene” and “gene” are also used interchangeably and both terms encompass fragments or variants thereof encoding the target protein. The transgenes of the present invention include nucleic acid sequences that have been removed from their naturally occurring environment, recombinant or cloned DNA isolates, and chemically synthesized analogues or analogues biologically synthesized by heterologous systems. Minor variations in the amino acid sequences of the invention are contemplated as being encompassed by the present invention, providing that the variations in the amino acid sequence(s) maintain at least 60%, at least 70%, more preferably at least 80%, at least 85%, at least 90%, at least 95%, and most preferably at least 97% or at least 99% sequence identity to the amino acid sequence of the invention or a fragment thereof as defined anywhere herein. The term homology is used herein to mean identity. As such, the sequence of a variant or analogue sequence of an amino acid sequence of the invention may differ on the basis of substitution (typically conservative substitution) deletion or insertion. Proteins comprising such variations are referred to herein as variants. Proteins of the invention may include variants in which amino acid residues from one species are substituted for the corresponding residue in another species, either at the conserved or non‐ conserved positions. Variants of protein molecules disclosed herein may be produced and used in the present invention. Following the lead of computational chemistry in applying multivariate data analysis techniques to the structure/property‐activity relationships [see for example, Wold, et al. Multivariate data analysis in chemistry. Chemometrics‐Mathematics and Statistics in Chemistry (Ed.: B. Kowalski); D. Reidel Publishing Company, Dordrecht, Holland, 1984 (ISBN 90‐277‐1846‐6] quantitative activity‐property relationships of proteins can be derived using well‐known mathematical
techniques, such as statistical regression, pattern recognition and classification [see for example Norman et al. Applied Regression Analysis. Wiley‐lnterscience; 3rd edition (April 1998) ISBN: 0471170828; Kandel, Abraham et al. Computer‐Assisted Reasoning in Cluster Analysis. Prentice Hall PTR, (May 11, 1995), ISBN: 0133418847; Krzanowski, Wojtek. Principles of Multivariate Analysis: A User's Perspective (Oxford Statistical Science Series, No 22 (Paper)). Oxford University Press; (December 2000), ISBN: 0198507089; Witten, Ian H. et al Data Mining: Practical Machine Learning Tools and Techniques with Java Implementations. Morgan Kaufmann; (October 11, 1999), ISBN:1558605525; Denison David G. T. (Editor) et al Bayesian Methods for Nonlinear Classification and Regression (Wiley Series in Probability and Statistics). John Wiley & Sons; (July 2002), ISBN: 0471490369; Ghose, Arup K. et al. Combinatorial Library Design and Evaluation Principles, Software, Tools, and Applications in Drug Discovery. ISBN: 0‐8247‐0487‐8]. The properties of proteins can be derived from empirical and theoretical models (for example, analysis of likely contact residues or calculated physicochemical property) of proteins sequence, functional and three‐dimensional structures and these properties can be considered individually and in combination. Amino acids are referred to herein using the name of the amino acid, the three‐letter abbreviation or the single letter abbreviation. The term “protein", as used herein, includes proteins, polypeptides, and peptides. As used herein, the term “amino acid sequence” is synonymous with the term “polypeptide” and/or the term “protein”. In some instances, the term “amino acid sequence” is synonymous with the term “peptide”. The terms "protein" and "polypeptide" are used interchangeably herein. In the present disclosure and claims, the conventional one‐letter and three‐ letter codes for amino acid residues may be used. The 3‐letter code for amino acids as defined in conformity with the IUPACIUB Joint Commission on Biochemical Nomenclature (JCBN). It is also understood that a polypeptide may be coded for by more than one nucleotide sequence due to the degeneracy of the genetic code. Amino acid residues at non‐conserved positions may be substituted with conservative or non‐ conservative residues. In particular, conservative amino acid replacements are contemplated. A “conservative amino acid substitution” is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain. Families of amino acid residues having similar side chains have been defined in the art, including basic side chains (e.g., lysine, arginine, or histidine), acidic side chains (e.g., aspartic acid or glutamic acid), uncharged polar side chains (e.g., glycine, asparagine, glutamine, serine, threonine, tyrosine, or cysteine), nonpolar side chains (e.g., alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, or tryptophan), beta‐branched side chains (e.g., threonine, valine, isoleucine) and aromatic side chains (e.g., tyrosine, phenylalanine, tryptophan, or histidine). Thus, if an amino acid in a polypeptide is replaced with another amino acid
from the same side chain family, the amino acid substitution is considered to be conservative. The inclusion of conservatively modified variants in a protein of the invention does not exclude other forms of variant, for example polymorphic variants, interspecies homologs, and alleles. “Non‐conservative amino acid substitutions” include those in which (i) a residue having an electropositive side chain (e.g., Arg, His or Lys) is substituted for, or by, an electronegative residue (e.g., Glu or Asp), (ii) a hydrophilic residue (e.g., Ser or Thr) is substituted for, or by, a hydrophobic residue (e.g., Ala, Leu, Ile, Phe or Val), (iii) a cysteine or proline is substituted for, or by, any other residue, or (iv) a residue having a bulky hydrophobic or aromatic side chain (e.g., Val, His, Ile or Trp) is substituted for, or by, one having a smaller side chain (e.g., Ala or Ser) or no side chain (e.g., Gly). “Insertions” or “deletions” are typically in the range of about 1, 2, or 3 amino acids. The variation allowed may be experimentally determined by systematically introducing insertions or deletions of amino acids in a protein using recombinant DNA techniques and assaying the resulting recombinant variants for activity. This does not require more than routine experiments for a skilled person. A “fragment” of a polypeptide comprises at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97% or more of the original polypeptide. For example, a fragment may comprise at least 5, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20 or more amino acids of the protein from which it is derived. A fragment may be continuous or discontinuous, preferably continuous. The polynucleotides of the present invention may be prepared by any means known in the art. For example, large amounts of the polynucleotides may be produced by replication in a suitable host cell. The natural or synthetic DNA fragments coding for a desired fragment will be incorporated into recombinant nucleic acid constructs, typically DNA constructs, capable of introduction into and replication in a prokaryotic or eukaryotic cell. Usually the DNA constructs will be suitable for autonomous replication in a unicellular host, such as yeast or bacteria, but may also be intended for introduction to and integration within the genome of a cultured insect, mammalian, plant or other eukaryotic cell lines. The polynucleotides of the present invention may also be produced by chemical synthesis, e.g. by the phosphoramidite method or the tri‐ester method, and may be performed on commercial automated oligonucleotide synthesizers. A double‐stranded fragment may be obtained from the single stranded product of chemical synthesis either by synthesizing the complementary strand and annealing the strand together under appropriate conditions or by adding the complementary strand using DNA polymerase with an appropriate primer sequence.
When applied to a nucleic acid sequence, the term “isolated” in the context of the present invention denotes that the polynucleotide sequence has been removed from its natural genetic milieu and is thus free of other extraneous or unwanted coding sequences (but may include naturally occurring 5' and 3' untranslated regions such as promoters and terminators), and is in a form suitable for use within genetically engineered protein production systems. Such isolated molecules are those that are separated from their natural environment. In view of the degeneracy of the genetic code, considerable sequence variation is possible among the polynucleotides of the present invention. Degenerate codons encompassing all possible codons for a given amino acid are set forth below: Amino Acid Codons Degenerate Codon Cys TGC TGT TGY Ser AGC AGT TCA TCC TCG TCT WSN Thr ACA ACC ACG ACT ACN Pro CCA CCC CCG CCT CCN Ala GCA GCC GCG GCT GCN Gly GGA GGC GGG GGT GGN Asn AAC AAT AAY Asp GAC GAT GAY Glu GAA GAG GAR Gln CAA CAG CAR His CAC CAT CAY Arg AGA AGG CGA CGC CGG CGT MGN Lys AAA AAG AAR Met ATG ATG Ile ATA ATC ATT ATH Leu CTA CTC CTG CTT TTA TTG YTN Val GTA GTC GTG GTT GTN Phe TTC TTT TTY Tyr TAC TAT TAY Trp TGG TGG Ter TAA TAG TGA TRR Asn/ Asp RAY Glu/ Gln SAR Any NNN
One of ordinary skill in the art will appreciate that flexibility exists when determining a degenerate codon, representative of all possible codons encoding each amino acid. For example, some polynucleotides encompassed by the degenerate sequence may encode variant amino acid sequences, but one of ordinary skill in the art can easily identify such variant sequences by reference to the amino acid sequences of the present invention. A “variant” nucleic acid sequence has substantial homology or substantial similarity to a reference nucleic acid sequence (or a fragment thereof). A nucleic acid sequence or fragment thereof is “substantially homologous” (or “substantially identical”) to a reference sequence if, when optimally aligned (with appropriate nucleotide insertions or deletions) with the other nucleic acid (or its complementary strand), there is nucleotide sequence identity in at least about 70%, 75%, 80%, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or more% of the nucleotide bases. Methods for homology determination of nucleic acid sequences are known in the art. Alternatively, a “variant” nucleic acid sequence is substantially homologous with (or substantially identical to) a reference sequence (or a fragment thereof) if the “variant” and the reference sequence they are capable of hybridizing under stringent (e.g. highly stringent) hybridization conditions. Nucleic acid sequence hybridization will be affected by such conditions as salt concentration (e.g. NaCl), temperature, or organic solvents, in addition to the base composition, length of the complementary strands, and the number of nucleotide base mismatches between the hybridizing nucleic acids, as will be readily appreciated by those skilled in the art. Stringent temperature conditions are preferably employed, and generally include temperatures in excess of 30°C, typically in excess of 37°C and preferably in excess of 45°C. Stringent salt conditions will ordinarily be less than 1000 mM, typically less than 500 mM, and preferably less than 200 mM. The pH is typically between 7.0 and 8.3. The combination of parameters is much more important than any single parameter. Methods of determining nucleic acid percentage sequence identity are known in the art. By way of example, when assessing nucleic acid sequence identity, a sequence having a defined number of contiguous nucleotides may be aligned with a nucleic acid sequence (having the same number of contiguous nucleotides) from the corresponding portion of a nucleic acid sequence of the present invention. Tools known in the art for determining nucleic acid percentage sequence identity include Nucleotide BLAST (as described below). One of ordinary skill in the art appreciates that different species exhibit “preferential codon usage”. As used herein, the term “preferential codon usage” refers to codons that are most frequently used in cells of a certain species, thus favouring one or a few representatives of the possible codons
encoding each amino acid. For example, the amino acid threonine (Thr) may be encoded by ACA, ACC, ACG, or ACT, but in mammalian host cells ACC is the most commonly used codon; in other species, different codons may be preferential. Preferential codons for a particular host cell species can be introduced into the polynucleotides of the present invention by a variety of methods known in the art. Introduction of preferential codon sequences into recombinant DNA can, for example, enhance production of the protein by making protein translation more efficient within a particular cell type or species. Thus, according to the invention, in addition to the gag‐pol genes any nucleic acid sequence may be codon‐optimised for expression in a host or target cell. In particular, the vector genome (or corresponding plasmid), the REV gene (or corresponding plasmid), the fusion protein (F) gene (or correspond plasmid) and/or the hemagglutinin‐neuraminidase (HN) gene (or corresponding plasmid, or any combination thereof may be codon‐optimised. A “fragment” of a polynucleotide of interest comprises a series of consecutive nucleotides from the sequence of said full‐length polynucleotide. By way of example, a “fragment” of a polynucleotide of interest may comprise (or consist of) at least 600 consecutive nucleotides from the sequence of said polynucleotide (e.g. at least 600, 650, 700, 750, 800 850, 900, or 950 consecutive nucleic acid residues of said polynucleotide). Typically, a fragment as defined herein retains the same function as the full‐length polynucleotide. The terms "decrease", "reduced", "reduction", or "inhibit" are all used herein to mean a decrease by a statistically significant amount. The terms "reduce," "reduction" or "decrease" or "inhibit" typically means a decrease by at least 10% as compared to a reference level (e.g. the absence of a given treatment) and can include, for example, a decrease by at least about 10%, at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, at least about 99% , or more. As used herein, "reduction" or "inhibition" encompasses a complete inhibition or reduction as compared to a reference level. "Complete inhibition" is a 100% inhibition (i.e. abrogation) as compared to a reference level. The terms "increased", "increase", "enhance", or "activate" are all used herein to mean an increase by a statically significant amount. The terms "increased", "increase", "enhance", or "activate" can mean an increase of at least 25%, at least 50% as compared to a reference level, for example an increase of at least about 50%, or at least about 75%, or at least about 80%, or at least about 90%, at least about 95%, or at least about 98%, or at least about 99%, or at least about 100%, or at least about 250% or more compared with a reference level, or at least about a 1.5‐fold, or at least about a 2‐fold, or at least about a 2.5‐fold, or at least about a 3‐fold, or at least about a 4‐fold, or at least about a 5‐
fold or at least about a 10‐fold increase, or any increase between 1.5‐fold and 10‐fold or greater as compared to a reference level. In the context of a yield or titre, an "increase" is an observable or statistically significant increase in such level. The terms "individual”, "subject”, and "patient”, are used interchangeably herein to refer to a mammalian subject for whom diagnosis, prognosis, disease monitoring, treatment, therapy, and/or therapy optimisation is desired. The mammal can be (without limitation) a human, non‐human primate, mouse, rat, dog, cat, horse, or cow. In a preferred embodiment, the individual, subject, or patient is a human. An “individual” may be an adult, juvenile or infant. An “individual” may be male or female. A "subject in need" of treatment for a particular condition can be an individual having that condition, diagnosed as having that condition, or at risk of developing that condition. A subject can be one who has been previously diagnosed with or identified as suffering from or having a condition in need of treatment or one or more complications or symptoms related to such a condition, and optionally, have already undergone treatment for a condition as defined herein or the one or more complications or symptoms related to said condition. Alternatively, a subject can also be one who has not been previously diagnosed as having a condition as defined herein or one or more or symptoms or complications related to said condition. For example, a subject can be one who exhibits one or more risk factors for a condition, or one or more or symptoms or complications related to said condition or a subject who does not exhibit risk factors. As used herein, the term “healthy individual” refers to an individual or group of individuals who are in a healthy state, e.g. individuals who have not shown any symptoms of the disease, have not been diagnosed with the disease and/or are not likely to develop the disease e.g. cystic fibrosis (CF) or any other disease described herein). Preferably said healthy individual(s) is not on medication affecting CF and has not been diagnosed with any other disease. The one or more healthy individuals may have a similar sex, age, and/or body mass index (BMI) as compared with the test individual. Application of standard statistical methods used in medicine permits determination of normal levels of expression in healthy individuals, and significant deviations from such normal levels. Herein the terms “control” and “reference population” are used interchangeably. The term “pharmaceutically acceptable” as used herein means approved by a regulatory agency of the Federal or a state government, or listed in the U.S. Pharmacopeia, European Pharmacopeia or other generally recognized pharmacopeia The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that such publications constitute prior art to the claims appended hereto.
Disclosure related to the various methods of the invention are intended to be applied equally to other methods, therapeutic uses or methods, the data storage medium or device, the computer program product, and vice versa. Surfactant Protein B (SFTPB) promoter fragments The full‐length SFTPB (surfactant protein B; SFTPB) promoter has previously been defined as a sequence that could be amplified from human genomic DNA with the PCR primers ATTTGAGCTCTTCTTTCTGCTGAACCATCG (sense, SEQ ID NO: 49) and TCTTAGATCTGTCAGACAGCTCTGGGTTCC (antisense, SEQ ID NO: 50), wherein the underlines sequences align with GenBank NCBI Reference Sequence: NG_016967.1 (version 1, accessed 18 November 2022) while the additional 5’ sequences provide SacI (sense) and BglII (antisense) restriction enzyme sites. The forward primer binds to the sense strand from bases 4845 to 4864 in NG_016967.1. The reverse primer binds to the reverse complement of bases 5797 to 5816 in NG_016967.1. Thus, the primer pair define a 972 bp genomic fragment. The 972 bp genomic fragment includes all of exon 1 of the SFTPB gene (bases 5543 to 5565 of NG_016967.1), wherein the A at base 5543 is reported as the starting nucleotide of the mRNA generated by the SFTPB promoter and the initiating ATG at bases 5559 to 5561 of NG_016967.1 encodes the first methionine of pre‐pro‐SFTPB. This 972 bp genomic fragment further includes a part of intron 1 from bases 5626 to 5816 of NG_016967.1. Preferably, references herein to a full‐length SFTPB promoter refer specifically to this 972 bp genomic fragment, which is present SEQ ID NO: 2. As described and exemplified herein, the present inventors have identified and isolated functional fragments of the SFTPB (surfactant protein B) gene promoter. The SFTPB promoter fragments generated by the inventors are able to drive cell‐specific expression of transgenes in the lung parenchyma. The fragments of the SFTPB gene promoter are shorter than the 972 bp genomic fragment previously identified, i.e. are shorter than the full‐length SFTPB promoter of SEQ ID NO: 2 previously reported in the art. Therefore, the SFTPB promoter fragments of the invention may be also referred to as a core SFTPB promoters. As discussed in more detail below and as exemplified herein, the SFTPB promoter fragments of the invention surprisingly increase transgene expression compared with the full‐length SFTPB promoter, and can do so in a lung parenchymal cell preferred/specific manner. A promoter of the invention is an SFTPB promoter fragment as described herein. Said SFTPB promoter fragment is functional, also as described herein. As used herein, the term “promoter of the invention” is used interchangeably with the terms “SFTPB promoter fragment” and “functional SFTPB
promoter fragment”. Thus, all disclosure herein to a (functional) SFTPB promoter fragment applies equally and without reservation to all promoters of the invention. An SFTPB promoter fragment of the invention comprises or consists of a core SFTPB promoter fragment, as described herein, or a variant thereof. An SFTPB promoter fragment of the invention may comprise or consist of a fragment of SEQ ID NO: 2, which lack all or part of SFTPB intron 1. SFTPB intron 1 begins at base 5626 of NG_016967.1 (corresponding to residue 782 of SEQ ID NO: 2) and corresponds to bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2). Thus, an SFTPB promoter fragment of the invention may comprise part (but not all) of the SFTPB intron 1. An SFTPB promoter fragment of the invention may not comprise one or more bases from bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2). By way of non‐limiting example, an SFTPB promoter fragment of the invention may comprise fewer than 180, fewer than 170, fewer than 160, fewer than 150, fewer than 140, fewer than 130, fewer than 120, fewer than 110, fewer than 100, fewer than 100, fewer than 90, fewer than 80, fewer than 70, fewer than 60, fewer than 50, fewer than 40, fewer than 30, fewer than 20, or fewer than 10 contiguous bases from 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2). Preferably, an SFTPB promoter fragment of the invention does not comprise bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2). Typically, the SFTPB promoter fragments of the invention do not comprise intron 1 of the SFTPB gene (or any portion thereof). Typically, an SFTPB promoter fragment of the invention comprises at least part of exon 1 of the SFTPB gene. SFTPB exon 1 begins at base 5543 of NG_016967.1 (corresponding to residue 699 of SEQ ID NO: 2) and corresponds to bases 5543 to 5565 of NG_016967.1 (corresponding to residues 699 to 781 of SEQ ID NO: 2). Thus, an SFTPB promoter fragment of the invention may comprise the contiguous nucleotide sequence of bases 5543 to 5565 of NG_016967.1 (corresponding to residues 699 to 781 of SEQ ID NO: 2). An SFTPB promoter fragment of the invention may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 bases, preferably 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 bases, 3’ to base 5542 of NG_016967.1 (corresponding to residue 698 of SEQ ID NO: 2), and wherein the additional bases correspond to a fragment of exon 1 of the SFTPB gene. By way of non‐limiting example, when an SFTPB promoter fragment of the invention comprises 1 base 3’ to base 5542 of NG_016967.1, the additional base corresponds to base 5543 of NG_016967.1 (corresponding to residue 699 of SEQ ID NO: 2), when an SFTPB promoter fragment of the invention comprises 2 bases 3’ to base 5542 of NG_016967.1, the additional bases correspond to bases 5543 to 5544 of NG_016967.1 (corresponding to residues 699 to 700 of SEQ ID NO: 2), when an SFTPB promoter
fragment of the invention comprises 3 bases 3’ to base 5542 of NG_016967.1, the additional bases correspond to bases 5543 to 5545 of NG_016967.1 (corresponding to residues 699 to 701 of SEQ ID NO: 2), when an SFTPB promoter fragment of the invention comprises 4 bases 3’ to base 5542 of NG_016967.1, the additional bases correspond to bases 5543 to 5546 of NG_016967.1 (corresponding to residues 699 to 702 of SEQ ID NO: 2), and so on. Preferably, an SFTPB promoter fragment of the invention comprises a portion of SFTPB corresponding to SEQ ID NO: 3. An SFTPB promoter fragment of the invention may not comprise the SFTPB gene start codon. Typically, an SFTPB promoter fragment of the invention may comprise of at least part of exon 1 of the SFTPB gene, but does not comprise the SFTPB gene start codon. Thus, an SFTPB promoter fragment of the invention typically does not comprise the initiating ATG at bases 5559 to 5561 of NG_016967.1 (corresponding to residues 715 to 717 of SEQ ID NO: 2). Preferably an SFTPB promoter fragment of the invention also does not comprise any of the SFTPB gene sequence 3’ of this start codon. An SFTPB promoter fragment of the invention may comprise fewer than 900 bases in length, fewer than 850 bases in length, fewer than 800 bases in length, fewer than 750 bases in length, fewer than 700 bases in length, fewer than 650 bases in length, fewer than 640 bases in length, fewer than 630 bases in length, fewer than 620 bases in length, fewer than 610 bases in length or fewer than 600 bases in length. Preferably, an SFTPB promoter fragment of the invention comprises fewer than 800 bases in length. A particularly preferred SFTPB promoter fragment of the invention comprises fewer than 700 bases in length. An SFTPB promoter fragment of the invention may consist of fewer than 900 bases in length, fewer than 850 bases in length, fewer than 800 bases in length, fewer than 750 bases in length, fewer than 700 bases in length, fewer than 650 bases in length, fewer than 640 bases in length, fewer than 630 bases in length, fewer than 620 bases in length, fewer than 610 bases in length or fewer than 600 bases in length. Preferably, an SFTPB promoter fragment of the invention consists of fewer than 800 bases in length. A particularly preferred SFTPB promoter fragment of the invention may consist of fewer than 700 bases in length. The exemplified SFTPB promoter fragment of the invention consists of 630 or 635 bases in length. An SFTPB promoter fragment of the invention may comprise from 600 bases to 800 bases in length, from 600 bases to 750 bases in length, from 600 bases to 700 bases in length, or from 600 bases to 650 bases in length. An SFTPB promoter fragment of the invention may consist of from 600 bases to 800 bases in length, from 600 bases to 750 bases in length, from 600 bases to 700 bases in length, or from 600 bases to 650 bases in length. An SFTPB promoter fragment of the invention may comprise from 630 bases to 800 bases in length, from 635 bases to 800 bases in length, from 630 bases to 750 bases in length, from 635 bases to 750 bases in length, from 630 bases to 700 bases in length, from 635 bases to 700 bases in length,
from 630 bases to 650 bases in length, or from 635 bases to 650 bases in length. An SFTPB promoter fragment of the invention may consist of from 630 bases to 800 bases in length, from 635 bases to 800 bases in length, from 630 bases to 750 bases in length, from 635 bases to 750 bases in length, from 630 bases to 700 bases in length, from 635 bases to 700 bases in length, from 630 bases to 650 bases in length, or from 635 bases to 650 bases in length. An SFTPB promoter fragment of the invention may be a fragment of a mammalian or avian, preferably a mammalian SFTPB promoter, i.e. an SFTPB promoter fragment of the invention may be a mammalian or avian SFTPB promoter fragment. By way of non‐limiting example, a mammalian SFTPB promoter fragment may be a human, non‐human primate, mouse, rat, dog, cat, horse, or cow SFTPB promoter fragment. Preferably, the SFTPB promoter fragment is a fragment of a human SFTPB promoter, i.e. preferably the SFTPB promoter fragment of the invention is a human SFTPB promoter fragment. By way of non‐limiting example, an SFTPB promoter fragment of the invention may comprise any functional fragment of SEQ ID NO: 2. Typically such an SFTPB promoter fragment of the invention is of a length as described herein. An SFTPB promoter fragment of the invention may comprise or consist of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8%sequence identity to a fragment of SEQ ID NO: 2, wherein the promoter retains the function of the SFTPB gene promoter, as defined herein. Typically such an SFTPB promoter fragment of the invention is of a length as described herein. Alternatively or in addition, such an SFTPB promoter fragment of the invention may comprise any additional feature (e.g. lack the SFTPB gene start codon and/or lack all or part of the SFTPB gene intron 1), as described herein. An SFTPB promoter fragment of the invention may comprise a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to bases 4925‐ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). For example, an SFTPB promoter fragment of the invention may comprise a sequence having at least 90% identity to bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of the invention may comprise a sequence having at least at least 95% identity to bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of the invention may comprise the sequence of bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of the invention may consist of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least
98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity or more to bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 90% identity to bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 95% identity to bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of the invention may consist of the sequence of bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). In some embodiments, an SFTPB promoter fragment of the invention derived from SEQ ID NO: 2, such as those described above, may comprise a substitution at residue 683 of SEQ ID NO: 2 (corresponding to base 5527 of NG_016967.1). Preferably, an SFTPB promoter fragment of the invention derived from SEQ ID NO: 2, such as those described above, may comprise substitution of an alanine residue by a cytosine at residue 683 of SEQ ID NO: 2 (corresponding to base 5527 of NG_016967.1), in other words, may comprise an A683C substitution. An exemplified SFTPB promoter fragment of the invention is SEQ ID NO: 1. Accordingly, the present invention provides an SFTPB promoter fragment which comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to SEQ ID NO: 1. For example, an SFTPB promoter fragment of the invention may comprise a sequence having at least 90% identity to SEQ ID NO: 1. An SFTPB promoter fragment of the invention may comprise a sequence having at least at least 95% identity to SEQ ID NO: 1. Preferably, an SFTPB promoter fragment of the invention may comprise the sequence of SEQ ID NO: 1. The present invention provides an SFTPB promoter fragment which consists of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to SEQ ID NO: 1. For example, an SFTPB promoter fragment of the invention may consist of a sequence having at least 90% identity to SEQ ID NO: 1. An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 95% identity to SEQ ID NO: 1. Preferably, an SFTPB promoter fragment of the invention may consist of the sequence of SEQ ID NO: 1. An SFTPB promoter fragment of the invention may further comprise up to 10 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bases), up to 15 bases, up to 20 bases, up to 25 bases or up to 25 bases at the 5’ end. Such additional bases may preferably be present when said SFTPB promoter fragment (i) comprises a sequence having at least 70% identity to SEQ ID NO: 1; and/or (ii) comprises a sequence
having at least 70% identity to bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2); as described herein. Where additional bases are present at the 5’ end of an SFTPB promoter fragment of the invention, said additional bases may be bases which correspond to the corresponding number of bases 5’ to base 4925 of NG_016967.1 (corresponding to residue 81 of SEQ ID NO: 2). By way of non‐limiting example, when an SFTPB promoter fragment of the invention comprises 10 bases 5’ to base 4925 of NG_016967.1, the additional base corresponds to bases 4915 to 4924 of NG_016967.1 (corresponding to residues 71 to 80 of SEQ ID NO: 2), or when an SFTPB promoter fragment of the invention comprises 20 bases 5’ to base 4925 of NG_016967.1, the additional bases correspond to bases 4905 to 4924 of NG_016967.1 (corresponding to residues 61 to 80 of SEQ ID NO: 2), and so on. Preferably, an SFTPB promoter fragment of the invention may further comprise up to 20 bases at the 5’ end, and particularly preferably, the up to 20 additional bases correspond to bases 5’ of residue 4925 of NG_016967.1 (corresponding to residue 81 of SEQ ID NO: 2). Alternatively or in addition, an SFTPB promoter fragment of the invention may further comprise up to 10 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bases) bases at the 3’ end. Such additional bases may preferably be present when said SFTPB promoter fragment (i) comprises a sequence having at least 70% identity to SEQ ID NO: 1; and/or (ii) comprises a sequence having at least 70% identity to bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2); as described herein. Where additional bases are present at the 3’ end of an SFTPB promoter fragment of the invention, said additional bases may be bases which correspond to the corresponding number of bases 3’ to base 5554 of NG_016967.1 (corresponding to residue 710 of SEQ ID NO: 2). By way of non‐limiting example, when an SFTPB promoter fragment of the invention comprises 10 bases 3’ to base 5554 of NG_016967.1, the additional base corresponds to bases 5555 to 5564 of NG_016967.1 (corresponding to residues 711 to 720 of SEQ ID NO: 2), when an SFTPB promoter fragment of the invention comprises 5 bases 3’ to base 5554 of NG_016967.1, the additional bases correspond to bases 5555 to 5559 of NG_016967.1 (corresponding to residues 711 to 715 of SEQ ID NO: 2), or when an SFTPB promoter fragment of the invention comprises 4 bases 3’ to base 5554 of NG_016967.1, the additional bases correspond to bases 5555 to 5558 of NG_016967.1 (corresponding to residues 711 to 714 of SEQ ID NO: 2), and so on. Preferably, an SFTPB promoter fragment of the invention may further comprise up to 4 bases at the 3’ end, and particularly preferably, the up to 4 additional bases correspond to bases 3’ of residue 5554 of NG_016967.1 (corresponding to residue 710 of SEQ ID NO: 2). An SFTPB promoter fragment of the invention may comprise additional sequences at the 5’ and/or 3’ end to facilitate molecular biology applications of said SFTPB promoter fragment. By way of
non‐limiting example, one or more restriction enzyme site may be added at the 5’ and/or 3’ end of the sequence. Such restriction enzyme sites may be used to facilitate cloning of the SFTPB promoter fragment into a non‐viral vector (e.g. plasmid) according to the invention, or into a manufacturing plasmid for use in the production of a viral vector according to the invention. When restriction enzyme sites are present at both the 5’ and 3’ end of the SFTPB promoter fragment, each restriction enzyme site may be selected independently, and thus the 5’ and 3’ sites may be the same or different. Examples of restriction enzyme sites that may be used include NheI and BgIII. For example, a SFTPB promoter fragment of the invention may have a 5’ BgIII restriction enzyme site (5’AGATCT3’) and/or a 3’ NheI restriction enzyme site (5’GCTAGC3’). Enhancer As exemplified herein, the inventors have also modified the SFTPB promoter fragments of the invention to further increase transgene expression levels and/or to increase promoter activity, particularly in the lung parenchyma. In particular, the inventors modified the SFTPB promoter fragment to include an enhancer sequence. Thus, the invention further provides an SFTPB promoter fragment which further comprises an enhancer. All disclosure herein to SFTPB promoter fragments of the invention applies equally and without reservation to SFTPB promoter fragment which further comprise an enhancer. An enhancer is a cis‐acting DNA sequence which can increase gene transcription. An enhancer of the invention may be from about 150 to about 900 bp in length, such as from about 200 to about 900 bp, from about 200 to about 800 bp, from about 300 to about 700 bp, from about 400 to about 800 bp, or from about 300 to about 600 bp in length. An enhancer of the invention may be linked to an SFTPB promoter fragments of the invention by a linker. Said linker is typically a short DNA sequence, which may be from about 1 to about 50 bp in length, such as from about 1 to about 20 bp, from about 1 to about 10 bp, from about 5 to about 20 bp in length, or from about 5 to about 10 bp in length. A linker may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp, particularly 6 bp, in length. A non‐limiting example of a linker sequence is given in SEQ ID NO: 51. Said linker may comprise or consist of one or more restriction enzyme site, non‐limiting examples of which are described herein. By way of example, a promoter/enhancer combination of the invention may be joined by a linker which comprises or consists of a NheI restriction site and/or a BgIII restriction site, preferably a BgIII restriction site. An enhancer of the invention may comprise additional sequences at the 5’ and/or 3’ end. By way of non‐limiting example, one or more restriction enzyme site may be added at the 5’ and/or 3’ end of the enhancer. When restriction enzyme sites are present at both the 5’ and 3’ end of the
enhancer, each restriction enzyme site may be selected independently, and thus the 5’ and 3’ sites may be the same or different. Said 5’ or 3’ restriction site may comprise all or part of a linker joining the enhancer to the SFTPB promoter fragment of the invention. The enhancer may be (i) 5’ to an SFTPB promoter fragment of the invention, or (ii) 3’ to an SFTPB promoter fragment of the invention. Preferably, the enhancer is 5’ to an SFTPB promoter fragment of the invention. The (5’ or 3’, preferably 5’) enhancer may be in (i) the forward, or (ii) the reverse orientation. Preferably, the enhancer is in the forward orientation. The inventors surprisingly found that the (cell‐specific) expression driven by a SFTPB promoter fragment further comprising an enhancer sequence as defined herein is higher than gene expression levels when using an SFTPB promoter fragment of the invention alone. Thus, an SFTPB promoter fragment further comprising an enhancer as provided herein has the potential to provide an even greater increase in transgene expression compared with the full‐length SFTPB promoter (i.e. transgene expression increases full‐length SFTPB promoter < SFTPB promoter fragment of the invention < SFTPB promoter of the invention further comprising an enhancer). Higher expression may be defined as greater than 105%, greater than 110%, greater than 120%, greater than 130%, greater than 140%, greater than 150%, greater than 175%, greater than 200%, greater than 225%, greater than 250%, greater than 300%, greater than 350%, greater than 400% or more of the transgene expression compared with a suitable control, such as the SFTPB promoter fragment alone, or the full‐ length SFTPB promoter. By way of non‐limiting example, expression of a transgene by an SFTPB promoter of the invention further comprising an enhancer may be greater than 105%, greater than 110%, greater than 120%, greater than 130%, greater than 140%, greater than 150%, greater than 175%, greater than 200%, greater than 225%, or greater than 250% of the transgene expression when using the SFTPB promoter fragment alone. Transgene expression may be quantified at the nucleic acid and/or protein level, and can be quantified by any suitable standard technique known to the person skilled in the art, for example, by real‐time reverse transcription polymerase chain reaction (RT‐qPCR), Western blotting and enzyme‐linked immunosorbent assay or ELISA. The inventors also surprisingly found that the (cell‐specific) expression driven by a SFTPB promoter fragment comprising an enhancer sequence as defined herein is comparable to gene expression levels when using ubiquitously used strong promoters (e.g. CMV, hCEF). Comparable expression may be defined as at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% of gene expression relative to expression of the same transgene using the same strong promoter (e.g. CMV or hCEF promoter [such as SEQ ID NOs: 6 and 5 respectively]). Expression of the transgene using the
SFTPB promoter fragment may be higher than expression using a strong promoter (e.g. CMV or hCEF) promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more relative to expression of the same transgene using the same strong promoter (e.g. CMV or hCEF promoter). Where the SFTPB promoter fragment comprises a SFTPB promoter fragment as defined above operably linked to an enhancer, said enhancer may preferably be of viral origin, a lung‐preferred enhancer, a lung‐parenchyma‐preferred enhancer, a pneumocyte‐preferred enhancer, an ATII and club cell‐preferred enhancer, or an ATII cell preferred enhancer, a lung‐specific enhancer, a lung‐ parenchyma‐specific enhancer, a pneumocyte‐specific enhancer, an ATII and club cell‐specific enhancer, or an ATII cell specific enhancer. The terms “preferred” and “specific” are defined herein. As described herein, the term “preferred” may alternatively or additionally be defined as higher expression in the lung relative to expression in the nose (particularly the nasal cavity), and thus give rise to an increased ratio of lung expression: nose expression, as described herein. Expression in the lung and nasal cavity can be determined using an in vivo luciferase reporter assay (e.g., wherein vectors comprising nucleic acid cassettes of the invention are administered to mice, and the relative bioluminescence of the lungs and nasal cavity is quantified). The enhancer may be selected from a hB‐actin enhancer, a SLC34A2 (Sodium‐dependent phosphate transport protein 2B) enhancer, a VEGFA (Vascular Endothelial Growth Factor A) enhancer, a CMV (Cytomegalovirus) enhancer, an SV40 (simian virus 40) enhancer, an ELF3 (E74 like ETS transcription factor 3) enhancer, an SFTPC (surfactant protein C) enhancer, an SFTPB enhancer, or a LMO7 (LIM domain 7) enhancer. Preferably, the enhancer is selected from a SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer. For example, in some preferred embodiments, the enhancer is a SLC34A2 enhancer. In other preferred embodiments, the enhancer is a VEGFA enhancer. In other preferred embodiments, the enhancer is a CMV enhancer. In other preferred embodiments, the enhancer is a SV40 enhancer. In other preferred embodiments, the enhancer is an ELF3 enhancer. Combinations of enhancers, typically those identified herein, and combinations of one or more preferred enhancer described herein, may be used according to the present invention. The CMV enhancer may be in the forwards orientation. Alternatively, the CMV enhancer is in the reverse orientation. The ELF3 enhancer may be in the forwards orientation. The SV40 enhancer may be in the forwards orientation. The SV40 enhancer may be in the reverse orientation. The VEGFA enhancer may be in the forwards orientation. Alternatively, the VEGFA enhancer may be in the reverse orientation.
The SLC34A2 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31. The SLC34A2 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SQE ID NO: 31. The SLC34A2 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31. The SLC34A2 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31. The SLC34A2 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31. The SLC34A2 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31. The VEGFA enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38. The VEGFA enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38. The VEGFA enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38. The VEGFA enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38. The VEGFA enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38. The VEGFA enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38. The CMV enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11. The CMV enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11. The CMV enhancer may comprise the nucleotide sequence of SEQ ID NO: 5. The CMV enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11. The CMV enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11. The CMV enhancer may consist of the nucleotide sequence of SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11. This exemplary CMV enhancer sequence is CpG‐free. CpG‐containing variants of this CMV enhancer (or other CpG‐comprising CMV enhancers) are also encompassed by the present invention.
The SV40 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 34 or 35. The SV40 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 34 or 35. the SV40 enhancer may comprise the nucleotide sequence of SEQ ID NO: 34 or 35. The SV40 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 34 or 35. The SV40 enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or 35. The SV40 enhancer may consist of the nucleotide sequence of SEQ ID NO: 34 or 35. The ELF3 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 13 to 20. The ELF3 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 13 to 20. The ELF3 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 13 to 20. The ELF3 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 13 to 20. The ELF3 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 13 to 20. The ELF3 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 13 to 20. The actin enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 21 or 22. The actin enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 21 or 22. The actin enhancer may comprise the nucleotide sequence of SEQ ID NO: 21 or 22. The actin enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 21 or 22. The actin enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 21 or 22. The actin enhancer may consist of the nucleotide sequence of SEQ ID NO: 21 or 22. The LMO7 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 23 to 25. The LMO7 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 23 to 25. The LMO7 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 23 to 25. The LMO7 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 23 to 25. The LMO7 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 23 to 25. The LMO7 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 23 to 25.
The SFTPC enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 27 or 28. The SFTPC enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 27 or 28. The SFTPC enhancer may comprise the nucleotide sequence of SEQ ID NO: 27 or 28. The SFTPC enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 27 or 28. The SFTPC enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 27 or 28. The SFTPC enhancer may consist of the nucleotide sequence of SEQ ID NO: 27 or 28. The SFTPB enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 26. The SFTPB enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 26. The SFTPB enhancer may comprise the nucleotide sequence of SEQ ID NO: 26. The SFTPB enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 26. The SFTPB enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 26. The SFTPB enhancer may consist of the nucleotide sequence of SEQ ID NO: 26. Any SFTPB promoter fragment of the invention may be combined with any enhancer of the invention. For the avoidance of doubt, and by way of non‐limiting example, it is envisaged that any preferred SFTPB promoter fragment of the invention may be combined with any preferred enhancer (e.g. an SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer). A preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76). Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76). Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76). Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT,
identified in the sequence information section herein) is omitted (SEQ ID NO: 76). A particularly preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82). Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82). Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82). Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82). A particularly preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 47. Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 47. Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 47. Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 47. A preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77). Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77). Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77). Said
preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77). A particularly preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83). Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83). Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83). Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83). A particularly preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 48. Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 48. Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 48. Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 48. A preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78). Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78). Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic
acid sequence of SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78). Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78). A particularly preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81). Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81). Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81). Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81). A particularly preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 46. Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 46. Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 46. Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 46. A preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80). Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the
BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80). Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80). Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80). A preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79). Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79). Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79). Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79). Thus, a preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%,a t least 98%, at least 99% identity or more to any one of SEQ ID NOs: 41 to 48. A preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to any one of SEQ ID NOs: 41 to 48. A preferred SFTPB promoter fragment further comprising an enhancer may comprise a nucleic acid sequence of any one of SEQ ID NOs: 41 to 48. A preferred SFTPB promoter fragment further comprising an enhancer may consist of a nucleic acid sequence of any one of SEQ ID NOs: 41 to 48. A particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%,a t least 98%, at least 99% identity or more to any one of SEQ ID NOs: 46 to 48. A particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity
to any one of SEQ ID NOs: 46 to 48. A particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise a nucleic acid sequence of any one of SEQ ID NOs: 46 to 48. A particularly preferred SFTPB promoter fragment further comprising an enhancer may consist of a nucleic acid sequence of any one of SEQ ID NOs: 46 to 48. Lung parenchyma and specific/preferred expression therein The respiratory system can be divided into airways and lung parenchyma. The airways consist of the bronchus, which bifurcates off the trachea and divides into bronchioles and then further into alveoli. The parenchyma is responsible for gas exchange and includes the alveoli, alveolar ducts, and terminal and respiratory bronchioles. The most prominent structure in the lung parenchyma is the alveolus. Two types of epithelial cell line the alveolus. Alveolar type I (ATI) cells exhibit a broad, flattened morphology and cover around 95% of the surface area, whilst the cuboidal alveolar type II cells (ATII cells) line the remainder of the alveolus. ATI cells provide a gas exchange interface with the underlying endothelium, whereas ATII cells serve as both progenitors of ATI cells and also play a critical role in maintaining the homeostasis of the alveolus. The latter role is fulfilled by the secretion of surfactant proteins from specialised organelles within ATII cells, so‐called ‘lamellar bodies’, into the alveolar space. Secretion of surfactant proteins maintain surface tension and prevents atelectasis at the end of expiration, whilst contributing to the varied functions of the ATII cells. ATII cells are the only epithelial cell of the lung which synthesise and release all four surfactant proteins A, B, C and D, with surfactant protein C being unique to the ATII cell. In addition to synthesising, storing and secreting surfactant components, ATII cells have the following functions: (1) the transepithelial movement of water and ions regulating the volume of the alveolar surface liquid (ASL) preventing alveoli flooding, (2) the expression of immunomodulatory proteins necessary for host defence and the regulation of innate immunity and (3) the regeneration of alveolar epithelium after injury. In addition, surfactant proteins A, B and D are also synthesised by club cells (previously named Clara Cells) founds in the terminal and respiratory bronchioles of humans. Club cells are non‐ciliated epithelial cells found mainly in bronchioles as well as basal cells found in large airways. They have been ascribed several protective roles, including airway repair after injury, secretion of anti‐inflammatory and immunomodulatory proteins, and detoxification. ATI dysfunction, ATII dysfunction and/or club cell dysfunction or dropout is associated with the pathogenesis of various parenchymal lung diseases. Accordingly, the lung parenchyma may be targeted for treating genetic diseases such as surfactant deficiencies and interstitial lung disease.
The promoters of the invention drive transgene expression in the lung parenchyma, typically one or more cell type of the lung parenchyma. In particular, the promoters of the invention drive transgene expression in (i) ATI cells, (ii) ATII cells, (iii) club cells, (iv) ATI cells and ATII cells, (v) ATI cells and club cells, (vi) ATII cells and club cells, or (vii) ATI cells, ATII cells and club cells. A SFTPB promoter fragment of the invention is functional. As used herein, the term “functional” may mean that an SFTPB promoter fragment of the invention retains the functionality of the full‐length SFTPB promoter. In other words, a promoter of the invention may be defined as being capable of expressing a gene of interest (e.g., a transgene) in a cell type which, in a healthy subject, would express SFTPB. For example, the an SFTPB promoter fragment may be used to express a transgene in ATII cells, club cells and/or ATI cells. Additionally, or alternatively, an SFTPB promoter fragment of the invention may be defined as functional as it is capable of preferentially or specifically expressing a gene of interest (e.g., a transgene) in a tissue‐type or cell‐type which, in a healthy subject, would express SP‐B. Thus, tan SFTPB promoter fragment of the invention may be a lung‐parenchyma preferred promoter. An SFTPB promoter fragment of the invention may be a lung‐parenchyma specific promoter. An SFTPB promoter fragment of the invention may be a pneumocyte‐preferred promoter, whereby a pneumocyte is defined as any of the specialized cells of the alveoli of the lungs. An SFTPB promoter fragment of the invention may be a pneumocyte‐specific promoter. An SFTPB promoter fragment of the invention may be an ATII cell‐preferred promoter, a club cell‐preferred promoter and/or an ATI cell‐preferred promoter. An SFTPB promoter fragment of the invention may be an ATII cell‐specific promoter, a club cell‐specific expression and/or an ATI cell‐specific promoter. An SFTPB promoter fragment of the invention may preferably be an ATII cell‐preferred promoter. An SFTPB promoter fragment of the invention may preferably be an ATII cell‐specific promoter. Tissue or cell preferred expression may be defined as expression that is higher in said tissue or cell than other tissue or cell types. For example, lung‐parenchyma preferred expression may be defined as expression that is significantly higher in the lung‐parenchyma (or one or more cell type therein, as described above) than expression in one or more of: the brain, the eye, the endocrine tissues, the proximal digestive tract, the gastrointestinal tract, liver and gall bladder, pancreas, the kidney and/or urinary bladder, male tissues (i.e., the testis, epididymis, prostate and/or seminal vesicle) and female tissues (i.e., the vagina, breast, cervix, endometrium, fallopian tube, ovary and placenta), muscle tissues, connective and soft tissues, skin, bone marrow and lymphoid tissue. Lung‐ parenchyma preferred expression may be defined as expression that is at least about 5 times greater, at least about 10 times greater, at least about 20 times greater, at least about 50 times greater or at
least about 100 or more times greater in the lung parenchyma than one or more of the above reference tissue types, especially the gastrointestinal tract and/or the brain. Tissue or cell specific expression may be defined as expression that is at least about 10 times greater, at least about 20 times greater, at least about 50 times greater or at least about 100 or more times greater in said tissue or cell than any other tissue or cell types. Preferential and/or specific expression may be assessed at the level of RNA and/or protein expression, preferably protein expression. Expression can be measured by any suitable standard technique known to the person skilled in the art. For example, RNA expression levels can be measured by quantitative real‐time PCR. Protein expression can be measured by western blotting or immunohistochemistry. Advantageously, restricting the expression of transgenes to cells expressing endogenous SFTPB is expected to reduce the effects of off‐target gene expression, overexpression (e.g., toxicity/ER stress/UPR, etc) and/or reduce immune responses. As such, the SFTPB promoters of the invention, as a consequence of their preferential and/or specific expression in the lung parenchyma, or one or more cell type thereof, have potential clinical benefits as a result of these advantageous properties. By way of non‐limiting example, compared with a hCEF promoter of SEQ ID NO: 5 and/or a CMV promoter of SEQ ID NO: 6, an SFTPB promoter fragment of the invention may reduce the effects of off‐target gene expression, overexpression, and/or reduce immune responses. Additionally, or alternatively, compared to a hCEF promoter of SEQ ID NO: 5and/or a CMV promoter of SEQ ID NO: 6, an SFTPB promoter fragment of the invention may preferentially or specifically express a transgene in the lung parenchyma, pneumocytes, such as ATII cells, ATI cells, and/or club cells. As described and exemplified herein, an SFTPB promoter fragment of the invention may increase transgene expression compared with a full‐length SFTPB promoter, such as that described herein. Typically, an SFTPB promoter fragment of the invention increases transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)). An SFTPB promoter fragment of the invention increases transgene expression in lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2‐fold, at least about 2.5‐fold, at least about 3‐ fold, at least about 4‐fold, at least about 5‐fold, at least about 7.5‐fold or more. The increase in transgene expression by an SFTPB promoter fragment of the invention may be quantified compared with a suitable control, preferably compared with a full‐length SFTPB promoter, such as that described herein. Thus, an SFTPB promoter fragment of the invention increases transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with expression of the same transgene in the lung parenchyma (or one
or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with a full‐length SFTPB promoter, such as that described herein. An SFTPB promoter fragment of the invention increases transgene expression in lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2‐fold, at least about 2.5‐fold, at least about 3‐fold, at least about 4‐fold, at least about 5‐fold, at least about 7.5‐fold or more compared with expression of the same transgene in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with a full‐length SFTPB promoter, such as that described herein. In some preferred embodiments, the SFTPB promoter fragment of the invention increases transgene expression in ATII cells, and optionally one or more additional lung parenchymal cell type as described herein. Again, this increase in expression is preferably compared with expression of the same transgene in the same cell type(s) by a full‐length SFTPB promoter, such as that described herein. Alternatively or additionally, as described and exemplified herein, an SFTPB promoter fragment of the invention may preferentially drive transgene expression in the lung (particularly the lung parenchyma or one or more cell type thereof) compared with transgene expression in the nose (or cells thereof). Thus, an SFTPB promoter fragment of the invention may preferentially increase transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with transgene expression in the nose (or cells thereof). An SFTPB promoter fragment of the invention may preferentially increase transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2‐fold, at least about 2.5‐ fold, at least about 3‐fold, at least about 4‐fold, at least about 5‐fold, at least about 7.5‐fold or more compared with transgene expression in the nose (or cells thereof). The ratio of lung expression: nose expression by an SFTPB promoter fragment of the invention may be at least about 2:1, such as at least about 2.5:1, at least about 3:1, at least about 4:1, at least about 5:1 or more. Any disclosure herein in relation to promoters (i.e. SFTPB promoter fragments) of the invention applies equally and without reservation to other aspects of the invention comprising, or relating to, said promoters. Thus, any disclosure herein in relation to promoters (i.e. SFTPB promoter fragments) of the invention applies equally and without reservation to promoter/enhancer combinations, viral and non‐viral vectors of the invention. By way of non‐limiting example, as for promoters (i.e. SFTPB promoter fragments) of the invention, promoter/enhancer combinations, viral and non‐viral vectors of the invention drive transgene expression in the lung parenchyma, typically one or more cell type of the lung parenchyma. In particular, the promoters (i.e. SFTPB promoter fragments) and promoter/enhancer combinations, viral and non‐viral vectors of the invention drive
transgene expression in (i) ATI cells, (ii) ATII cells, (iii) club cells, (iv) ATI cells and ATII cells, (v) ATI cells and club cells, (vi) ATII cells and club cells, or (vii) ATI cells, ATII cells and club cells. Nucleic Acid Cassettes The invention also provides a nucleic acid cassette. In particular, the present invention provides a nucleic acid cassette comprising (a) an SFTPB promoter fragment; and (b) a transgene. As defined herein, a transgene may be defined as anucleic acid sequence encoding a therapeutic protein.Thus, the terms “nucleic acid sequence encoding a therapeutic protein” and the term “transgene” may be used interchangeably. In a nucleic acid of the invention, an SFTPB promoter fragment may increase expression of the transgene by the lung parenchyma (e.g. ATII cells), as defined herein. The increase in expression of a transgene by an SFTPB promoter fragment of the invention may be as defined herein, including disclosure of increasing transgene expression using SFTPB promoter fragments and/or SFTPB promoter fragments combined with an enhancer, as described above. Accordingly, any disclosure herein in relation to increasing transgene expression using an SFTPB promoter fragment and/or SFTPB promoter fragment combined with an enhancer of the invention applies equally and without reservation to nucleic acid cassettes of the invention. For example, the increase in expression of a transgene by an SFTPB promoter fragment of the invention is an increase of at least about 50%, at least about 60%, at least about 70%, at least about 80% or more. Preferably, the increase in expression of a transgene by an SFTPB promoter fragment of the invention is an increase of at least about 50%. In a nucleic acid of the invention, an SFTPB promoter fragment of the invention may increase expression of the transgene by the lung parenchyma (e.g. ATII cells) relative to a suitable control For example, the SFTPB promoter fragment of the invention may increase expression of the transgene by the lung parenchyma (e.g. ATII cells) relative to expression using the full‐length SFTPB gene promoter (e.g., the 972 bp genomic fragment defined above). Alternatively or in addition, expression of the transgene by the lung parenchyma (e.g. ATII cells) using the SFTPB promoter fragment in a nucleic acid of the invention, may be comparable to the expression of the transgene using a ubiquitously used strong promoter (e.g. CMV or hCEF). For example, the expression of the transgene using the SFTPB promoter fragment may be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% of the expression using a CMV promoter, as described herein. Alternatively, expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention, may be higher than expression using a CMV promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more
relative to expression of the same transgene using a CMV promoter. Alternatively or in addition, expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% of the expression using a hCEF promoter. Alternatively, the expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may be higher than expression using a hCEF promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more relative to expression of the same transgene using a CMV promoter. A nucleic acid cassette or vector of the invention enables long‐term transgene expression, resulting in long‐term expression of a transgene, which offers clinical benefits for the expression of therapeutic proteins in patients. As described herein, the phrases “long‐term expression”, “sustained expression”, “long‐lasting expression” and “persistent expression” are used interchangeably. Long‐ term expression according to the present invention means expression of a transgene, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more. Preferably long‐ term expression means expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more. The long‐term expression is typically accompanied by long‐term secretion or long‐term membrane insertion of the (therapeutic) protein encoded by the transgene, depending on whether the (therapeutic) protein is a secreted protein (e.g. SFTPB) or a membrane protein. Long‐term secretion according to the present invention means secretion of a (therapeutic) protein encoded by a transgene of the invention, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more. Preferably long‐term secretion means secretion for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more. Long‐term membrane insertion according to the present invention means that a (therapeutic) protein encoded by a transgene of the invention is inserted into and present in the cell membrane, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more. Preferably long‐term expression means membrane insertion for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
In particular, a nucleic acid cassette or vector of the invention may drive (increased) long‐ lasting expression of a transgene, and/or (preferably and) expression and secretion/membrane insertion of a (therapeutic) protein encoded by said transgene in one or more cell type of the lung parenchyma, as described herein, in vivo in a patient. Preferably, a nucleic acid cassette or vector of the invention drives expression of a transgene, and/or (preferably and) expression and secretion/membrane insertion of a (therapeutic) protein encoded by a transgene of the invention in one or more cell / cell type of the lung parenchyma, as described herein, for at least 45 days, more preferably at least 90 days. The nucleic acid of the nucleic acid cassette may be as defined herein. The nucleic acid cassette comprise DNA or RNA. Preferably the nucleic acid cassette is DNA. A nucleic acid cassette of the invention may optionally be codon optimised for expression in a particular cell type, for example, eukaryotic cells (e.g. mammalian cells, yeast cells, insect cells or plants cells) or prokaryotic cells (e.g. E.coli). The term “codon optimised” refers to the replacement of at least one codon within a base polynucleotide sequence with a codon that is preferentially used by the host organism in which the polynucleotide is to be expressed. Typically, the most frequently used codons in the host organism are used in the codon‐optimised polynucleotide sequence. Methods of codon optimisation are well known in the art. It will be understood by a skilled person that numerous different polynucleotides can encode the same polypeptide as a result of the degeneracy of the genetic code. It is also understood that skilled persons may, using routine techniques, make nucleotide substitutions that do not affect the polypeptide sequence encoded by the nucleic acid molecules to reflect the codon usage of any particular host organism in which the polypeptides are to be expressed. Therefore, unless otherwise specified, a nucleic acid cassette that comprises or consists of the SFTPB promoter fragment and transgene (e.g., encoding a therapeutic protein) of the invention includes all polynucleotide sequences that are degenerate versions of each other and that encode the same amino acid sequence. A nucleic acid cassette of the invention preferably comprises an SFTPB promoter fragment comprising an SFTPB promoter fragment operably linked to an enhancer sequence, as described herein. A nucleic acid cassette of the invention typically comprises a SFTPB promoter fragment comprising a SFTPB promoter fragment and an enhancer sequence, wherein the SFTPB promoter fragment is operably linked to a nucleic acid sequence comprising or consisting of a transgene (e.g., encoding a therapeutic protein). By operably linked, it is meant that the SFTPB promoter fragment is configured to express the transgene (e.g., encoding the therapeutic protein). The transgene encoding the therapeutic protein may also be linked to a suitable terminator sequence. Suitable terminator sequences are well known in the art.
The promoter included in the nucleic acid cassettes and vectors of the invention may be specifically selected and/or modified to further refine regulation of expression of the therapeutic gene. Again, suitable promoters and standard techniques for their modification are known in the art. As a non‐limiting example, an SFTPB promoter fragment of the invention may be modified to reduce the number of CpG dinucleotides, or to render the SFTPB promoter fragment CpG‐free. A number of (CpG‐free) promoters, and methods for the generation of CpG‐free promoters which are suitable for use in the present invention are described in Pringle et al. (J. Mol. Med. Berl. 2012, 90(12): 1487‐96), which is herein incorporated by reference in its entirety. Preferably, the nucleic acid cassettes and vectors of the invention comprise an SFTPB promoter fragment having low or no CpG dinucleotide content. Low CpG dinucleotide content may be defined as 10 CpG dinucleotides or less, preferably 5 CpG dinucleotides or less, such as 5, 4, 3, 2 or 1 CpG dinucleotides. An SFTPB promoter fragment may have some or all CG dinucleotides replaced with any one of AG, TG or GT. The absence (or reduction) of CpG dinucleotides further improves the performance of some nucleic acid cassettes and vectors of the invention, particularly lentiviral (e.g. SIV) vectors of the invention and in particular in situations where it is not desired to induce an immune response against an expressed antigen or an inflammatory response against the delivered expression construct. The elimination or reduction of CpG dinucleotides reduces the occurrence of flu‐like symptoms and inflammation which may result from administration of constructs, particularly when administered to the airways. The nucleic acid cassettes and vectors of the invention may be modified to allow shut down of gene expression. Standard techniques for modifying the vector in this way are known in the art. As a non‐limiting example, Tet‐responsive promoters are widely used. The nucleic acid cassette of the invention (or a vector comprising said cassette) may have an intron positioned between the promoter and the transgene. Non‐limiting examples of suitable introns are found for example, in UK Application No. 2213936.4, which is herein incorporated by reference in its entirety. Wherein the nucleic acid of the invention is present in a non‐viral vector (e.g. plasmid), the presence of at least one intron between the SFTPB promoter fragment and the transgene may be preferred, for example an intron as described in UK Application No. 2213936.4. The nucleic acid cassettes and vectors of the invention may include at least one part of a vector, in particular, regulatory elements. By way of non‐limiting example, the promoter within a nucleic acid cassette of the invention may be used to express more than one polypeptide, including one or more therapeutic protein. Thus, the nucleic acid cassette may comprise a nucleic acid sequence which, when transcribed, gives rise to multiple polypeptides, for instance a transcript may contain multiple open reading frames (ORFs) and also one or more Internal Ribosome Entry Sites (IRES) to allow translation of ORFs after the first ORF. A transcript may be polycistronic, i.e. it may be translated
to give a polypeptide which is subsequently cleaved to give a plurality of polypeptides. Alternatively, a nucleic acid cassette of the invention may comprise multiple promoters, including multiple SFTPB promoter fragments of the invention, or multiple copies of any specific an SFTPB promoter fragment of the invention, and hence give rise to a plurality of transcripts and hence a plurality of polypeptides, including a plurality of therapeutic proteins. Nucleic acid cassettes may, for instance, express one, two, three, four or more polypeptides via a promoter or promoters, including one or more SFTPB promoter fragment of the invention. A nucleic acid cassette may comprise one or more translation initiation sequence (TIS). Translation initiation plays an important role in mRNA translation, canonically a methionyl tRNA unique for initiation (Met‐tRNAi) identifies the AUG start codon and triggers the downstream translation process. Non‐canonical start codons (e.g. CUG for valyl‐tRNA)/TIS may also be used. The nucleic acid cassettes of the present invention may comprise at least one termination signal. A “termination signal" or "terminator" is comprised of the DNA sequences involved in specific termination of an RNA transcript by an RNA polymerase. Thus, a termination signal that ends the production of an RNA transcript is contemplated according to the present invention. A terminator may be necessary in vivo to achieve desirable message levels. In eukaryotic systems, a terminator region may also comprise specific DNA sequences that permit site‐specific cleavage of the new transcript so as to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch of about 200 A residues (polyA) to the 3’ end of the transcript. RNA molecules modified with this polyA tail appear to more stable and are translated more efficiently. Thus, when the nucleic acid cassette is for expression in eukaryotes, a terminator typically comprises a signal for the cleavage of the RNA, and it is preferred that the terminator signal promotes polyadenylation of the message. The terminator and/or polyadenylation site elements can serve to enhance message levels and to minimize read through from the cassette into other sequences. Terminators contemplated for use in the invention include any known terminator of transcription described herein or known to one of ordinary skill in the art, including but not limited to, for example, the termination sequences of genes, such as for example the bovine growth hormone terminator or viral termination sequences, such as for example the SV40 terminator. In certain embodiments, the termination signal may be a lack of transcribable or translatable sequence, such as due to a sequence truncation. The invention also provides gene therapy vectors comprising a nucleic acid cassette of the invention. Any and all disclosure herein in relation to nucleic acid cassettes of the invention applies equally and without reservation to gene therapy vectors of the invention.
The nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in a given host cell. Any appropriate host cell may be used, such as mammalian, bacterial, insect, yeast, and/or plant host cells. In addition, cell‐free expression systems may be used. Such expression systems and host cells are standard in the art. Typically the nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in the lung. The nucleic acid cassettes and vectors of the invention may be capable of expressing the transgene in the lung parenchyma. The nucleic acid cassettes and vectors of the invention may be capable of expressing the transgene in one or more cell type selected from ATII cells, ATI cells, club cells, and/or bronchioalveolar stem cells. The nucleic acid cassettes and vectors of the invention are typically capable of expressing the transgene in one or more ATII cells, ATI cells, club cells, bronchioalveolar stem cells in the terminal bronchioles. Preferably, the nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in ATII cells. Thus, the nucleic acid cassettes and vectors of the invention are capable of expressing the (therapeutic) protein encoded by the transgene in one or more cell/cell type of the lung parenchyma, particularly ATII cells, ATI cells, club cells and/or bronchioalveolar stem cells in the terminal bronchioles, particularly in ATII cells. The nucleic acid cassettes of the invention may be made using any suitable process known in the art. Thus, the nucleic acid cassettes may be made using chemical synthesis techniques. Alternatively, the nucleic acid cassettes of the invention may be made using molecular biology techniques. Non‐Viral Vectors The present invention also provides a vector comprising a nucleic acid cassette of the invention. The vector(s) may be present in the form of a therapeutic composition or formulation. The vector may be a non‐viral vector. The non‐viral vector(s) may be a DNA vector, such as a DNA plasmid. The vector(s) may be an RNA vector, such as a mRNA vector or a self‐amplifying RNA vector. The non‐viral vector may be an exosomes or microvesicle (MV). The non‐viral (e.g. DNA and/or RNA) vector(s) of the invention may be capable of expression in eukaryotic and/or prokaryotic cells. Typically, the non‐viral (e.g. DNA and/or RNA) vector(s) are capable of expression in a cell of a subject, for example, a cell of a mammalian or avian subject to be immunised. Typically the nucleic acid cassettes and vectors of the invention are capable of expressing a transgene in airway cells, preferably lung parenchymal cells (as described herein).
A non‐viral vector of the present invention may be a phage vector, such as an AAV/phage hybrid vector as described in Hajitou et al., Cell 2006; 125(2) pp. 385‐398; herein incorporated by reference. Vector(s) of the present invention (e.g. non‐viral DNA or RNA vectors) may be designed in silico, and then synthesised by conventional polynucleotide synthesis techniques. Non‐viral plasmids cannot replicate in the subject to be treated, as they lack the viral genetic material which hijacks the body's normal production machinery. However they are capable of replicating in appropriate host cells, such as yeasts or bacteria including E. coli, and particularly airway cells as defined herein. The term "plasmid" as used herein refers to a construction comprised of genetic material designed to direct transformation of a targeted cell. The plasmid contains a plasmid backbone. A "plasmid backbone" as used herein contains multiple genetic elements positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transfected cells. The plasmid backbone can contain one or more unique restriction sites within the backbone. The plasmid may be capable of autonomous replication in a defined host or organism such that the cloned sequence is reproduced. The plasmid can confer some well‐defined phenotype on the host organism which is either selectable or readily detected. The plasmid or plasmid backbone may have a linear or circular configuration. The components of a plasmid can contain, but is not limited to, a DNA molecule incorporating: (1) the plasmid backbone; (2) a sequence comprising or consisting of an SFTPB promoter fragment; (3) a transgene sequence encoding a (therapeutic) protein; and optionally (4) additional regulatory elements for transcription, translation, RNA stability and replication. The purpose of the plasmid in human gene therapy for the efficient delivery of nucleic acid sequences to, and expression of therapeutic proteins in, a cell or tissue. In particular, the purpose of the plasmid is to achieve high copy number, avoid potential causes of plasmid instability and provide a means for plasmid selection. As for expression, the nucleic acid cassette contains the necessary elements for expression of the nucleic acid within the cassette. Expression includes the efficient transcription of an inserted gene, nucleic acid sequence, or nucleic acid cassette with the plasmid. A DNA plasmid may be CpG‐free, or be optimised to reduce CpG dinucleotides as described herein. A DNA plasmid of the invention may be codon‐optimised as described herein. Methods of preparing plasmid DNA are well known in the art. Typically, they are capable of autonomous replication in an appropriate host or producer cell. The term "exosome" as used herein refers to an extracellular vesicle formed by exocytosis from a cell of origin. An exosome typically comprises a nucleic cassette of the invention. Exosomes
may be used to transform a targeted cell. The nucleic acid within an exosome may contain one or more genetic elements, which may be positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transformed cells. The term "microvesicle" (MV) as used herein refers to an extracellular vesicle formed typically between about 30 to about 1,000 nm in diameter. An MV typically comprises a nucleic cassette of the invention. Exosomes may be used to transform a targeted cell. The nucleic acid within an MV may contain one or more genetic elements, which may be positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transformed cells. Host cells containing (e.g. transformed, transfected, or electroporated with) the plasmid may be prokaryotic or eukaryotic in nature, either stably or transiently transformed, transfected, or electroporated with the plasmid. Suitable host cells include bacterial, yeast, fungal, invertebrate, and mammalian cells. Preferably the host cell is bacterial; more preferably E. coli. Host cells can then be used in methods for the large scale production of the plasmid. The cells are grown in a suitable culture medium under favourable conditions, and the desired plasmid isolated from the cells, or from the medium in which the cells are grown, by any purification technique well known to those skilled in the art; e.g. see Sambrook et al, supra. Any appropriate delivery means can be used to deliver a non‐viral vector (e.g. plasmid) of the invention to a target cell or patient. Suitable delivery means are known in the art and within the routine skill of one of ordinary skill in the art. Non‐limiting examples include the use of cationic lipids, polymers (e.g. polyethyleneimine and poly‐L‐lysine) and electroporation. Preferably cationic lipids may be used to deliver non‐viral (e.g. plasmid) vectors of the invention to target cells or to a patient. Non‐limiting examples of cationic lipids suitable for use according to the invention are GL67A and lipofectamine. The cationic lipid mixture GL67A is a mixture of three components ‐ GL67 (Cholest‐5‐en‐3‐ol (3β)‐,3‐[(3‐aminopropyl)[4‐[(3‐ aminopropyl)amino]butyl]carbamate], (CAS Number: 179075‐30‐0)), DOPE (1,2‐dioleoyl‐sn‐glycero‐3‐phosphoethanolamine) and DMPE‐PEG5000 (1,2‐Dimyristoyl‐sn‐ Glycero‐3‐Phosphoethanolamine‐N‐[methoxy (Polyethylene glycol)5000]). These components are formulated at a 1:2:0.05 molar ratio to form GL67A. The composition of GL67A and methods for its production are disclosed in WO2013/061091, as are methods for preparing mixtures of GL67A with exemplary non‐viral vectors. The contents of WO2013/061091 are herein incorporated by reference in their entirety.
Lipofectamine consists of a 3:1 mixture of DOSPA (2,3‐dioleoyloxy‐N‐ [2(sperminecarboxamido)ethyl]‐N,N‐dimethyl‐1‐propaniminium trifluoroacetate) and DOPE. Viral vectors The present invention also provides a vector comprising a nucleic acid cassette of the invention. The vector(s) may be present in the form of a therapeutic composition or formulation. The vector may be a viral vector. Any appropriate viral vector may be used to deliver a nucleic acid cassette of the invention. By way of non‐limiting example, a viral vector of the invention may be a lentiviral vector, an adeno‐ associated virus (AAV) vector, an adenoviral vector, a poxvirus vector, a herpes simplex virus (HSV) vector. Derivatives of these viral vectors, such as lentivirus‐derived particles are also encompassed within the invention. Non‐limiting examples of adenoviral vectors include human serotypes such as AdHu5, simian serotypes such as ChAd63, ChAdOX1 or ChAdOX2, and other forms. Non‐limiting examples of poxvirus vectors include a modified vaccinia Ankara (MVA)). ChAdOX1 and ChAdOX2 are disclosed in WO2012/172277 (herein incorporated by reference in its entirety). ChAdOX2 is a BAC‐derived and E4 modified AdC68‐based viral vector. Viral vectors are usually non‐replicating or replication impaired vectors, which means that the viral vector cannot replicate to any significant extent in normal cells (e.g. normal human cells), as measured by conventional means – e.g. via measuring DNA synthesis and/or viral titre. Non‐replicating or replication impaired vectors may have become so naturally (i.e. they have been isolated as such from nature) or artificially (e.g. by breeding in vitro or by genetic manipulation). There will generally be at least one cell‐type in which the replication‐impaired viral vector can be grown – for example, modified vaccinia Ankara (MVA) can be grown in CEF cells. Typically, the viral vector is incapable of causing a significant infection in an animal subject, typically in a mammalian subject such as a human or other primate. Viral vector(s) of the present invention may be designed in silico, and then synthesised by conventional polynucleotide synthesis techniques. Preferably the invention relates to retroviral vectors, particularly lentiviral vectors. The term “lentivirus” refers to a family of retroviruses. Retroviral/lentiviral vectors of the invention, can integrate into the genome of transduced cells and lead to long‐lasting expression. Examples of retroviruses suitable for use in the present invention include gammaretroviruses such as murine leukaemia virus (MLV) and feline leukaemia virus (FLV). Examples of lentiviruses suitable for use in the present invention include Simian immunodeficiency virus (SIV), Human immunodeficiency virus (HIV),
Feline immunodeficiency virus (FIV), Equine infectious anaemia virus (EIAV), and Visna/maedi virus. A particularly preferred lentiviral vector is an SIV vector (including all strains and subtypes), such as a SIV‐AGM (originally isolated from African green monkeys, Cercopithecus aethiops). The retroviral/lentiviral (e.g. SIV) vectors of the present invention are typically pseudotyped with hemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, or with G glycoprotein from Vesicular Stomatitis Virus (G‐VSV). Preferably the lentiviral (e.g. SIV) vectors of the present invention are pseudotyped with HN and F from a respiratory paramyxovirus. Particularly preferably the respiratory paramyxovirus is a Sendai virus (murine parainfluenza virus type 1). The F protein may be a truncated F protein, typically one in which the cytoplasmic domain is truncated. Preferably the truncated F protein is Fct4, in which 38 amino acids have been truncated from the C‐terminus of the F protein, with 4 amino acids of the F protein cytoplasmic domain being retained. The F protein may be a truncated F protein, typically one in which the cytoplasmic domain is truncated. Preferably the truncated F protein is Fct4, in which 38 amino acids have been truncated from the C‐terminus of the F protein, with 4 amino acids of the F protein cytoplasmic domain being retained. A retroviral/lentiviral (e.g. SIV) vector for use according to the invention may be integrase‐ competent (IC). Alternatively, the lentiviral (e.g. SIV) vector may be integrase‐deficient (ID). Viral vectors of the invention, particularly retroviral/lentiviral (e.g. SIV) vectors as described herein may transduce one or more cells types as described herein to achieve long term transgene expression. The HN protein may be a truncated and/or chimeric HN protein, typically one in which the cytoplasmic domain is truncated or substituted. Preferably, the HN protein is a chimeric HN protein in which (i) the cytoplasmic domain of the HN is replaced by the cytoplasmic domain of the transmembrane (TMP) protein; or (ii) the cytoplasmic domain of the TMP is added to the cytoplasmic domain of the HN protein. The HN protein may be as described in Kobayashi et al. (J. Virol. (2003) 77(4):2607‐2614), which is herein incorporated by reference in its entirety. Particularly preferred truncated and/or chimeric forms of the F and HN proteins are described in UK Patent Application No. 2212472.1, which is herein incorporated by reference in its entirety. The viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention enable high levels of transgene expression. Together with the increased levels of expression/secretion/membrane insertion of a therapeutic protein resulting from the use of an SFTPB promoter fragment of the invention, these viral vectors typically result in high levels (therapeutic levels) of expression of the transgene, and the (therapeutic) protein encoded by said transgene.
The transgene to be included in a viral vector of the invention, particularly a retroviral/lentiviral (e.g. SIV) vector of the invention may be modified to facilitate expression. For example, the transgene sequence may be in CpG‐depleted /low (or CpG‐fee) and/or codon‐optimised form to facilitate gene expression. Standard techniques for modifying the transgene sequence in this way are known in the art. The viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention exhibit enhanced expression of the transgene. Accordingly, the viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention are capable of producing long‐lasting, repeatable, high‐level transgene expression, particularly in lung parenchyma without inducing side effects (e.g., an undue immune response). The viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention enable long‐term transgene expression, resulting in long‐term expression (and secretion or membrane insertion) of a (therapeutic) protein by cells of the lung parenchyma as described herein. Long‐term expression according to the present invention means expression of a transgene gene and/or encoded (therapeutic) protein, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more. Preferably long‐term expression means expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more. Preferably, the invention relates to the use of F/HN lentiviral vectors comprising a nucleic acid cassette of the invention, particularly SIV F/HN vectors. The nucleic acid cassette comprised in a viral vector of the invention, particularly a retroviral/lentiviral (e.g. SIV) vector of the invention, may have no intron positioned between the promoter and the nucleic acid encoding the signal peptide and/or the nucleic acid encoding the therapeutic protein. Similarly, there may be no intron between the promoter and the nucleic acid encoding the signal peptide and/or the nucleic acid encoding the therapeutic protein in the vector genome (pDNA1) plasmid (for example, pGM326 or pGM830 as illustrated in Figures 2A and B and the corresponding sequences in UK Application No. 2102832.9, which is herein incorporated by reference in its entirety). The viral vectors of the invention may be made using any suitable process known in the art. In particular, retroviral/lentiviral (e.g. SIV) vectors of the invention may be made using the methods disclosed in UK Application No. 2102832.9, which is herein incorporated by reference in its entirety).
The viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention may comprise a central polypurine tract (cPPT) and/or the Woodchuck hepatitis virus posttranscriptional regulatory elements (WPRE). An exemplary WPRE sequence is provided by SEQ ID NO: 52. Transgenes A nucleic acid cassette of the invention comprises a transgene. Typically, the transgene encodes a therapeutic protein. A therapeutic protein is one which has potential utility in the treatment or prevention of a disease or condition, such as those describe herein. Thus, a nucleic acid cassette of the invention may comprise a nucleic acid encoding a protein which has a therapeutic effect on a disease or condition to be treated. A nucleic acid cassette of the invention may comprise a nucleic acid encoding a therapeutic protein which is a functional or wild‐type form of a protein which is present in a patient to be treated in a dysfunctional form (whether the dysfunction is inherent or acquired). As used herein, the phrase "inherent dysfunction" refers to a protein which is innately dysfunctional due to genetic factors and the phrase "acquired dysfunction" refers to a protein which is dysfunctional due to environmental or other factors after birth. Thus, a nucleic acid cassette of the invention may comprise a nucleic acid encoding a protein which is a functional or wild‐type form of a protein which is present in a patient, but which that has become dysfunctional due to a genetic disease, such as a genetic respiratory disease. The nucleic acid cassettes of the present invention are useful in the treatment of diseases via their use in expressing therapeutic proteins in lung parenchyma as described herein (e.g. ATI, ATII cells). Preferably, the nucleic acid cassettes of the present invention are useful in the treatment of diseases via their use in expressing therapeutic proteins in ATII cells, ATI cells and/or club cells. The transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally Surfactant Protein B (SFTPB), Surfactant Protein C (SFTPC), alpha‐1‐antitrypsin (AAT), Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, von Willebrand Factor, Granulocyte‐ Macrophage Colony‐Stimulating Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGGβ) or monoclonal antibody, an anti‐inflammatory decoy and a monoclonal antibody against an infectious agent; or (b) ATP‐binding cassette sub‐family member A (ABCA3), TRIM72, CSF2RA, CSF2RB and decorin. In some embodiments, the therapeutic protein is not an antibody, particularly not a monoclonal antibody. In such embodiments, the transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally SFTPB, SP‐C, AAT, Factor VIII, Factor VII, Factor IX,
Factor X, Factor XI, von Willebrand Factor, GM‐CSF, an anti‐inflammatory protein (e.g. IL‐10, or TGGβ) and an anti‐inflammatory decoy; or (b) ABCA3, TRIM72, CSF2RA, CSF2RB and decorin. The nucleic acid cassettes of the invention are particularly efficient at driving the expression, secretion and/or membrane insertion of proteins (e.g. therapeutic proteins as described herein) by the lung parenchyma. This is particularly the case when such cassettes are comprised within F/HN pseudotyped viral vectors of the invention (as described herein), which are efficient at targeting cells in the lung parenchyma. As such, for therapeutic applications the nucleic acid cassettes of the invention (and vectors comprising said cassettes) are typically delivered to cells of the respiratory tract, particularly the cells of the lung parenchyma. In other words, the nucleic acid cassettes of the invention (and vectors comprising said cassettes) are typically delivered to lung parenchyma as described herein. Accordingly, the nucleic acid cassettes of the invention (and vectors comprising said cassettes) are particularly suited for treatment of diseases or disorders of the lung parenchyma. Typically, the nucleic acid cassettes of the invention (and vectors comprising said cassettes) may be used for the treatment of a genetic respiratory disease. A nucleic acid cassette of the invention (or vector comprising said cassette) may comprise a nucleic acid encoding a polypeptide or protein that is therapeutic for the treatment of such diseases, particularly a disease or disorder of the lung parenchyma. The transgene and therapeutic protein of the invention are not limited, one of ordinary skill in the art will be able to identify transgenes an therapeutic proteins which may be usefully delivered according to the invention, particularly in the context of genetic diseases, particularly genetic respiratory diseases and diseases or disorders of the lung parenchyma as those described herein. Accordingly, a nucleic acid cassette of the invention (or vector comprising said cassette) may comprise a nucleic acid sequence encoding a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally SFTPB, SFTPC, AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, von Willebrand Factor, GM‐CSF, an anti‐inflammatory protein (e.g. IL‐10, TGGβ, or TNF‐alpha) or monoclonal antibody, an anti‐inflammatory decoy and a monoclonal antibody against an infectious agent; or (b) ABCA3, TRIM72, CSF2RA, CSF2RB or DCN. Other preferred examples of therapeutic proteins that may be encoded by a nucleic acid sequence comprised in a nucleic acid cassette of the invention (or vector comprising said cassette) include genes related to or associated with other surfactant deficiencies. The therapeutic protein encoded by a nucleic acid cassette of the invention (or vector comprising said cassette) may be SFTPB. Examples of an SFTPB therapeutic transgene are provided by SEQ ID NOs: 53, 54 and 56. An exemplary codon‐optimised SFTPB transgene is provided by SEQ ID NO:
55. The therapeutic protein encoded by said SFTPB transgene, may be exemplified by the polypeptide of SEQ ID NO: 57. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 53 to 57. The transgene may encode ABCA3. Examples of a ABACA3 transgene are provided by SEQ ID NOs: 58 and 59. An exemplary codon‐optimised ABACA3 transgene is provided by SEQ ID NO: 60. The polypeptide encoded by said ABACA3 transgene, may be exemplified by the polypeptide of SEQ ID NO: 61. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 58 to 61. The therapeutic protein encoded by a nucleic acid cassette of the invention (or vector comprising said cassette) may be SFTPC. An example of an SFTPC therapeutic transgene is provided by SEQ ID NO: 62. The therapeutic protein encoded by said SFTPC transgene, may be exemplified by the polypeptide of SEQ ID NO: 63. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 62 or 63. The therapeutic protein encoded by a nucleic acid cassette of the invention (or vector comprising said cassette) may be an AAT. An example of an AAT therapeutic transgene (SERPINA1) is provided by SEQ ID NO: 70. SEQ ID NO: 70 is a codon‐optimized CpG depleted AAT transgene (SERPINA1) previously designed by the present inventors to enhance translation in human cells. Such optimisation has been shown to enhance gene expression by up to 15‐fold. Variants of same sequence (as defined herein) which possess the same technical effect of enhancing translation compared with the unmodified (wild‐type) AAT gene sequence are also encompassed by the present invention. The therapeutic protein encoded by said AAT transgene, may be exemplified by the polypeptide of SEQ ID NO: 71. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 70 or 71. The therapeutic protein encoded by a nucleic acid cassette of the invention (or vector comprising said cassette) may be an FVIII. Examples of a FVIII therapeutic transgene are provided by SEQ ID NOs: 72 and 73. The polypeptide encoded by the FVIII transgene, may be exemplified by the polypeptide of SEQ ID NO: 74 and 75. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 72 to 75. The therapeutic protein encoded by a nucleic acid cassette of the invention (or vector comprising said cassette) may be GM‐CSF. A GM‐CSF transgene may comprise or consist of SEQ ID NO: 64 (human). The polypeptide encoded by the GM‐CSF transgene may be exemplified by the polypeptide of SEQ ID NO: 65 (human). Variants thereof (as described therein) are also included,
particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 64and 65. The transgene may encode decorin. An example of a DCN transgene is provided by SEQ ID NO: 66. The polypeptide encoded by said DCN transgene, may be exemplified by the polypeptide of SEQ ID NO: 67. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to SEQ ID NO: 66or 67. The transgene may encode TRIM72. An example of a TRIM72 transgene is provided by SEQ ID NO: 68. The polypeptide encoded by said TRIM72 transgene, may be exemplified by the polypeptide of SEQ ID NO: 69. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to SEQ ID NO: 68or 69. The therapeutic protein encoded by a nucleic acid cassette of the invention (or vector comprising said cassette) may be encoded by any one of SFTPB, SFTPC, Factor V, Factor VII, Factor IX, Factor X and/or Factor XI, von Willebrand Factor, GM‐CSF, ABCA3, TRIM72 or DCN, or other known related gene. As the lung parenchyma is preferably targeted for delivery of the nucleic acid cassettes of the invention (and vectors comprising said cassettes), the transgene may preferably be SFTPB, SFTPC, ABCA3 or GM‐CSF. The therapeutic protein may be a monoclonal antibody (mAb) against an infectious agent (bacterial, fungal or viral, e.g. the SARS‐Co‐V2 virus). The therapeutic protein may be anti‐TNF alpha. The therapeutic protein may be one implicated in an inflammatory, immune or metabolic condition. A nucleic acid cassette of the invention (or a vector comprising said cassette) may be delivered to one or more cell/cell type of the lung parenchyma to allow production of proteins to be secreted into circulatory system. In such embodiments, the therapeutic protein may be any one of Factor VII, Factor VIII, Factor IX, Factor X, Factor XI and/or von Willebrand’s factor. Such a nucleic acid cassette of the invention (or a vector comprising said cassette) may be used in the treatment of diseases, particularly cardiovascular diseases and blood disorders, preferably blood clotting deficiencies such as haemophilia. Again, the therapeutic protein may be an mAb against an infectious agent or a protein implicated in an inflammatory, immune or metabolic condition, such as, lysosomal storage disease. The nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the promoter and the nucleic acid encoding the therapeutic protein. Similarly, when the nucleic acid cassette is comprised in a viral vector, there may be no intron between the promoter and the transgene in the vector genome (pDNA1) plasmid used to make said viral vector, as described herein. Alternatively, said nucleic acid cassette of the invention (or a vector comprising
said cassette) may have an intron positioned between the promoter and the transgene, particularly if the cassette (or vector comprising said cassette) is non‐viral. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an SFTPB transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an SFTPC transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an AAT transgene (SERPINA1), including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an FVIII transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an FVII transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an FIX transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an FX transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an FXI transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene.
The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and a von Willebrand Factor transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and a Granulocyte‐Macrophage Colony‐Stimulating Factor (GM‐CSF) transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. In some preferred embodiments, the lentiviral (e.g. SIV) vector comprises a SFTPB promoter fragment and an DCN transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. In some preferred embodiments, the lentiviral (e.g. SIV) vector comprises a SFTPB promoter fragment and a TRIM72 transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. In some preferred embodiments, the lentiviral (e.g. SIV) vector comprises a SFTPB promoter fragment and a ABACA3 transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene. The nucleic acid cassette of the invention (or a vector comprising said cassette) comprises a nucleic acid encoding a therapeutic protein (said nucleic acid is referred to interchangeably herein as a transgene). The nucleic acid sequence encodes a gene product, e.g., a protein, particularly a therapeutic protein. For example, the nucleic acid cassette of the invention (or a vector comprising said cassette) comprises a nucleic acid sequence encoding an SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 and said nucleic acid sequence comprises (or consists of) a nucleic acid sequence having at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100%) sequence identity to the SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 nucleic acid sequence respectively, examples of which are described herein. The nucleic acid sequence encoding SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 may preferably comprise (or consist of) a nucleic acid sequence having at least 95% (such as at least 95, 96, 97, 98, 99
or 100%) sequence identity to the SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 nucleic acid sequence respectively, examples of which are described herein. The amino acid sequence of the (therapeutic) protein encoded by the transgene may be a functional variant having at least 95% (such as at least 95, 96, 97, 98, 99 or 100%) sequence identity to the functional protein. For example, an SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 polypeptide encoded by the respective SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 transgene may comprise (or consist of) an amino acid sequence having at least 95% (such as at least 95, 96, 97, 98, 99 or 100%) sequence identity to the functional SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 polypeptide sequence respectively. An SFTPB promoter fragment and/or enhancer of the invention may be linked to a transgene by a linker. Said linker is typically a short DNA sequence, as defined herein, and may comprise or consist of one or more restriction enzyme site, examples of which are also described herein. By way of example, a SFTPB promoter fragment of the invention may be joined to a transgene by a linker which comprises or consists of a NheI restriction site and/or a BgIII restriction site, preferably a NheI restriction site. Signal peptides The transgene encoding for a (therapeutic) protein may further comprise a nucleic acid sequence encoding for a signal peptide. Said signal peptide may be the endogenous signal peptide of the (therapeutic) protein, or a signal peptide exogenous to said (therapeutic) protein. When the transgene further comprises a nucleic acid encoding for an exogenous signal peptide, the transgene preferably exclude a nucleic acid sequence encoding for the endogenous signal peptide. In such instances, the exogenous signal peptide is typically the sole signal peptide linked with (and hence driving secretion and/or membrane insertion) of the therapeutic protein. All disclosure herein relates to both transgenes and therapeutic proteins including and excluding endogenous signal peptides unless explicitly stated. By way of non‐limiting example, sequence identity of variants, and/or lengths of fragments may be based on the sequence with or without a signal peptide. Any signal peptide and therapeutic combination may be used, provided that this combination is effective in increasing the expression, secretion and/or membrane insertion of a (therapeutic) protein as defined herein. Selection of a signal peptide may depend on the specific (therapeutic)
protein and/or the specific lung parenchymal cell type by which the (therapeutic) protein is to be expressed/secreted/inserted into the cell membrane. Exogenous signal peptides which increase expression, secretion and/or membrane insertion of a (therapeutic) protein have been described by the inventors in International Patent Application No: WO2022/219333, which is herein incorporated by reference in its entirety. Co‐administration of lentiviral vectors and surfactants As exemplified herein, lentiviral vectors, such as the lentiviral vector pseudotyped with haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, can be administered in combination with a synthetic surfactant, and achieve at least as efficient transduction into target cells, and potentially even enhanced transduction, compared with transduction of the lentiviral vector in vehicle alone. This is surprising as lentiviral vectors are surrounded by a lipid envelope, which would be predicted to be disrupted by the hydrophobic portions of a surfactant, having a negative effect on the lentiviral structure. Whilst the combination of a lentiviral vector and a surfactant is exemplified herein with a specific lentiviral vector, SIV.F/HN with a luciferase transgene, for the avoidance of doubt this example provides proof of concept for other lentiviral vectors, particularly other SIV.F/HN vectors to be co‐administered with a surfactant in this way. This is because it is the nature of the lentiviral structure, i.e. the presence of an envelope that is determinative. Having shown that co‐administration is feasible with one lentiviral vector, a skilled person would understand that co‐administration could be carried out with any lentiviral vector and a surfactant. Similarly, the nature of any pseudotyping proteins will not affect the ability of a lentiviral vector to be co‐ administered with a surfactant. The invention therefore provides lentiviral vector pseudotyped with haemagglutinin‐ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered in combination with a surfactant. Thus, the invention therefore provides lentiviral vector pseudotyped with haemagglutinin‐ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered simultaneously or sequentially with a surfactant. The lentiviral (e.g. SIV) vector and surfactant are administered in combination. Administered "in combination," encompasses both simultaneous (also referred to as concurrent) administration/delivery and sequential (also referred to as separate) administration/delivery.
For , "simultaneous" or "concurrent delivery”, the delivery of the lentiviral (e.g. SIV) vector may still be occurring when the delivery of the surfactant begins, or the delivery of surfactant may still be occurring when the delivery of the lentiviral (e.g. SIV) vector begins, so that there is overlap in terms of administration. Simultaneous delivery may encompass delivery of the lentiviral (e.g. SIV) vector and surfactant within weeks to months or even years of each other, typically so that the lentiviral (e.g. SIV) vector delivery overlaps with the delivery of the surfactant. Alternatively, the delivery of the lentiviral (e.g. SIV) vector may end before the delivery of the surfactant begins, or the delivery of the surfactant may end before delivery of the lentiviral (e.g. SIV) vector begins. Sequential administration may involve the lentiviral (e.g. SIV) vector and surfactant being administered within 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 6 hours, 12 hours or 24 hours or longer of each other. The lentiviral vector may be administered before the surfactant. The surfactant may be administered before the lentiviral vector. The lentiviral vector and surfactant may be administered simultaneously, optionally wherein the lentiviral vector and surfactant are mixed prior to administration. Typically the treatment is more effective because of combined administration. For example, treatment with the lentiviral (e.g. SIV) vector may be more effective, e.g., an equivalent effect is seen with less of the lentiviral (e.g. SIV) vector, or the lentiviral (e.g. SIV) vector reduces symptoms to a greater extent, than would be seen if the lentiviral (e.g. SIV) vector were administered in the absence of the surfactant. By way of further example, treatment with the surfactant may be more effective, e.g., an equivalent effect is seen with less of the surfactant, or the surfactant reduces symptoms to a greater extent, than would be seen if the surfactant were administered in the absence of the lentiviral (e.g. SIV) vector. Typically, delivery is such that the reduction in a symptom, or other parameter related to a disease to be treated is at least equivalent to what would be observed with the lentiviral (e.g. SIV) vector delivered in the absence of the surfactant, or the analogous situation is seen with the surfactant. Alternatively, a combination therapy of the invention may increase transgene expression by at least 1.2 fold, at least 1.3 fold, at least 1.4 fold, at least 1.5 fold, at least 2 fold, at least 2.5 fold or more compared with treatment with the lentiviral (e.g. SIV) vector alone (i.e. compared with the increase in transgene expression achieved when treating with the lentiviral (e.g. SIV) alone). Without being bound by theory, it is believed that the surfactant may aid the distribution of the lentiviral (e.g. SIV) vector within the lungs, enabling it to penetrate more deeply into the respiratory tree and thus facilitating transduction of the lung parenchyma.
It will be appreciated that appropriate dosage of the lentiviral (e.g. SIV) vector and/or the surfactant, will depend on the specific agent, and can also vary from patient to patient. Any surfactant may be used in a combination therapy according to the present invention. It will be appreciated that it is within the routine practice of one of ordinary skill in the art to select such a surfactant. By way of non‐limiting example, two animal‐derived surfactants in clinical use are Beractant (BLES) and Poractant alfa (Curosurf). Other synthetic and/or modified surfactants may also be used. A disease to be treated with such a lentiviral vector and a surfactant may be a genetic disease. Alternatively or in addition, the disease to be treated may be a respiratory disease, particularly a genetic respiratory disease; or a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder. Non‐limiting examples of diseases which may be treated according to this aspect of the invention include Surfactant Protein B (SP‐B) Deficiency; Surfactant Protein C (SP‐C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1‐antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID‐19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia. In some embodiments, include Surfactant Protein B (SP‐B) Deficiency may be a preferred indication according to the invention. A lentiviral vector for use in combination with a surfactant may comprise HN and F proteins from a Sendai virus, as described herein. Alternatively or in addition, a lentiviral vector for use in combination with a surfactant may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector. Preferably said lentiviral vector may be a SIV vector. The transgene may encode any suitable therapeutic protein as described herein. By way of non‐limiting example, said transgene may be selected from: (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP‐B), Surfactant Protein C (SP‐C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte‐Macrophage Colony‐Stimulating Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody, an anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP‐binding cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.
In some preferred embodiments, a lentiviral vector for use in combination with a surfactant may comprise a SFTPB promoter fragment as defined herein. Alternatively or in addition, said lentiviral vector may comprise a nucleic acid cassette as defined herein. Therapeutic Indications The nucleic acid cassettes and vectors of the present invention enable cell‐preferred or cell‐ specific expression of a transgene encoding a (therapeutic) protein. This may be further increased by the nucleic acid cassette or vector facilitating efficient transgene expression. The nucleic acid cassettes and vectors of the invention, and particularly the F/HN‐pseudotyped retroviral/lentiviral (e.g. SIV) vectors of the invention are capable of: (i) transduction of one or more cell/cell type of the lung parenchyma without disruption of epithelial integrity; (ii) persistent gene expression; (iii) lack of chronic toxicity; and/or (iv) efficient repeat administration. Long term/persistent stable gene expression, preferably at a therapeutically‐effective level, may be achieved using repeat doses of a nucleic acid cassette or vector of the present invention. Alternatively, a single dose may be used to achieve the desired long‐term expression. Thus, advantageously, the nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention can be used in gene therapy. Accordingly, the present invention provides a nucleic acid cassette or gene therapy vector as defined herein for use in a method of treating or preventing a disease. The disease to be treated may be chronic or acute. The nucleic acid cassettes and vectors (viral and non‐viral) of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention may be used to deliver any transgene useful in gene therapy. Typically, the nucleic acid cassettes and vectors (viral and non‐viral) of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention are for use in gene therapy for the treatment of a disease or disorder of the lung parenchyma. By way of example, efficient airway cell uptake properties of the nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention make them highly suitable for treating respiratory or lung diseases, particularly genetic respiratory diseases, particularly preferably those of/involving the lung parenchyma. The nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention can also be used in methods of gene therapy to promote secretion of therapeutic proteins. By way of further example, the invention provides secretion of therapeutic proteins into alveoli or lumen of the bronchioles. Administration of a nucleic acid cassettes and vectors of the present invention, and particularly a retroviral/lentiviral (e.g. SIV)
vector of the invention and its uptake by airway cells may be used to enable the use of the lungs as a “factory” to produce a therapeutic protein that is then secreted and enters the general circulation at therapeutic levels, where it can travel to cells/tissues of interest to elicit a therapeutic effect. Thus, other diseases which are not respiratory tract diseases, such as cardiovascular diseases, particularly genetic cardiovascular diseases or blood disorders, particularly blood clotting deficiencies, can also be treated by the nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention. Nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention can effectively treat a disease by providing a transgene for the correction of the disease. For example, resulting in the expression and secretion of SFTPB from cells of the lung parenchyma, to compensate for the pathologically low levels of SFTPB expression in patients with SFTPB deficiency. By way of further example, nucleic acid cassettes and vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be used to treat alpha‐1‐antitrypsin (AAT) deficiency, typically by gene therapy with a AAT transgene (SERPINA1) as described herein. AAT is a secreted anti‐protease that is produced mainly in the liver and then trafficked to the lung, with smaller amounts also being produced in the lung itself. The main function of AAT is to bind and neutralise/inhibit neutrophil elastase. Gene therapy with AAT according to the present invention is relevant to AAT deficient patient, as well as in other lung diseases such as CF or chronic obstructive pulmonary disease (COPD), and offers the opportunity to overcome some of the problems encountered by conventional enzyme replacement therapy (in which AAT isolated from human blood and administered intravenously every week), providing stable, long‐lasting expression in the target tissue (lung/nasal epithelium), ease of administration and unlimited availability. Transduction with a nucleic acid cassettes and vectors of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention may lead to secretion of the recombinant protein into the lumen of the lung as well as into the circulation. One benefit of this is that the therapeutic protein reaches the interstitium. AAT gene therapy may therefore also be beneficial in other disease indications, non‐limiting examples of which include type 1 and type 2 diabetes, acute myocardial infarction, ischemic heart disease, rheumatoid arthritis, inflammatory bowel disease, transplant rejection, graft versus host (GvH) disease, multiple sclerosis, liver disease, cirrhosis, vasculitides and infections, such as bacterial and/or viral infections. AAT has numerous other anti‐inflammatory and tissue‐protective effects, for example in pre‐ clinical models of diabetes, graft versus host disease and inflammatory bowel disease. The production
of AAT in the lung and/or nose following transduction according to the present invention may, therefore, be more widely applicable, including to these indications. Other examples of diseases that may be treated with gene therapy of a secreted protein according to the present invention include cardiovascular diseases and blood disorders, particularly blood clotting deficiencies such as haemophilia (A, B or C), von Willebrand disease and Factor VII deficiency. In some preferred embodiments, the disease to be treated is selected from a surfactant protein deficiency, such as Surfactant Protein B (SFTPB) Deficiency, Surfactant Protein C (SFTPC) deficiency, ABCA3 deficiency, Pulmonary surfactant metabolism dysfunction 2 (SMDP2) Pulmonary surfactant metabolism dysfunction 3 (SMDP3), or other surfactant deficiencies; Primary Ciliary Dyskinesia (PCD); Alpha 1‐antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP, hereditary and/or acquired); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID‐19; a pulmonary fibrotic disease (including idiopathic pulmonary fibrosis); a pulmonary allergic condition; a pulmonary bacterial infection; asthma; lung cancer; a dysplastic change in the lungs; and haemophilia. Other examples of diseases or disorders to be treated include Primary Ciliary Dyskinesia (PCD), acute lung injury, and/or inflammatory, infectious, immune or metabolic conditions, such as lysosomal storage diseases or a pulmonary bacterial infection, or any other lung disease or disorder. The nucleic acid cassettes and vectors (viral and non‐viral) of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention, typically provide high expression levels of a transgene of interest, and the (therapeutic) protein encoded thereby, when administered to a patient. The terms high expression and therapeutic expression are used interchangeably herein. Expression may be measured by any appropriate method (qualitative or quantitative, preferably quantitative), and concentrations given in any appropriate unit of measurement, for example ng/ml or µM. Expression/secretion/membrane insertion of a transgene, or the (therapeutic) protein of interest encoded thereby may be given in absolute terms. Alternatively, expression/secretion/membrane insertion of a therapeutic protein may be given in relative terms, for example relative to the expression/secretion/membrane insertion of the therapeutic protein encoded by a corresponding nucleic acid cassette or vector of the invention without the SFTB promoter fragment of the invention, relative to the expression/secretion/membrane insertion of the same transgene using the full‐length SFTB promoter as described herein, or relative to the expression/secretion/membrane insertion of the corresponding endogenous (defective) gene.
Expression may be measured in terms of mRNA or protein expression. The expression of the therapeutic protein of the invention may be quantified relative to the endogenous protein or gene in terms of protein concentration, mRNA copies per cell or any other appropriate unit. Secretion and/or membrane insertion of a therapeutic protein may be quantified relative to secretion/membrane insertion of the corresponding endogenous protein, or relative to the level of secretion/membrane insertion of the therapeutic protein introduced via an expression cassette with the same transgene but without the SFTB promoter fragment of the invention or comprising the full‐length SFTB promoter as described herein. Expression levels of a nucleic acid encoding a (therapeutic) protein and/or the expression/secretion/membrane insertion of the encoded (therapeutic) protein of the invention may be measured ex vivo (e.g. in the conditioned media used to culture the cells or within the cells themselves) or in vivo (e.g. in the lung tissue, epithelial lining fluid and/or serum/plasma) as appropriate. A high and/or therapeutic expression level may therefore refer to the concentration in the lung, epithelial lining fluid and/or serum/plasma. Repeated doses of nucleic acid cassettes and vectors (viral and non‐viral) of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention may be administered twice‐daily, daily, twice‐weekly, weekly, monthly, every two months, every three months, every four months, every six months, yearly, every two years, or more. Dosing may be continued for as long as required, for example, for at least six months, at least one year, two years, three years, four years, five years, ten years, fifteen years, twenty years, or more, up to for the lifetime of the patient to be treated. The invention also provides nucleic acid cassettes and vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention as described herein for use in a method of gene therapy, wherein said method comprises the steps of: (a) transducing cells (e.g. lung parenchyma) ex vivo to produce modified cells expressing a transgene of interest; and (b) administering the resulting modified cells. The invention provides a method of treating a disease, the method comprising administering a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention to a subject. Any disease described herein may be treated according to the invention. In particular, the invention provides a method of treating a lung disease using a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention. The disease to be treated may be a chronic disease. The invention also provides a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention as described herein for use in a method of treating a disease. Any disease described herein may be treated according to the invention.
In particular, the invention provides a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention for use in a method of treating a lung disease. The disease to be treated may be a chronic disease. The invention also provides the use of a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention as described herein in the manufacture of a medicament for use in a method of treating a disease. Any disease described herein may be treated according to the invention. In particular, the invention provides the use of a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention for the manufacture of a medicament for use in a method of treating a lung disease. The disease to be treated may be a chronic disease. The invention also provides a cell comprising a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector. Said cell may be a lung parenchyma cell as described herein. The invention further provides a method of expressing a transgene, typically a transgene encoding a (therapeutic) protein in a target cell, comprising delivering a nucleic acid cassette or a vector of the invention into the target cells. Said method may be carried out in vitro, ex vivo, or in vivo, preferably in vitro or ex vivo. The target cell may be any appropriate cell type, such as those described herein. By way of non‐limiting example, the target cells may be prokaryotic or eukaryotic, preferably eukaryotic. Particularly preferred are mammalian cells, such as human, non‐human primate, mouse, rat, dog, cat, horse, or cow cells. The cells may be primary cells or cell lines. Non‐ limiting examples of cells include ATII cells, ATI cells, club cells and/or HEK293T cells. The step of delivering the nucleic acid cassette or vector may comprise integrating said nucleic acid cassette or gene therapy vector into said target cell's genome. Any appropriate technique may be used to deliver the nucleic acid cassette or vector, examples of which are known in the art and within the routine practice of one of ordinary skill in the art. The method may further comprise a step of culturing cells expressing the transgene, and/or isolating or purifying the expressed (therapeutic) protein from said cells. Any and all disclosure herein in relation to nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention applies equally and without reservation to the therapeutic uses and methods described herein. Further, for the avoidance of doubt, any and all disclosure herein in relation to therapeutic uses and methods using nucleic acid cassettes or vectors of the present invention, applies equally and without reservation to the combination therapies of lentiviral (e.g. SIV) vectors and surfactants as described herein.
Long term/persistent stable gene expression, preferably at a therapeutically‐effective level, may be achieved using repeat doses of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention of the present invention. Alternatively, a single dose may be used to achieve the desired long‐term expression. Formulation and administration The invention also provides a composition comprising a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention, and optionally a pharmaceutically acceptable carrier, excipient, buffer or diluent. The nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered in any dosage appropriate for achieving the desired therapeutic effect. Appropriate dosages may be determined by a clinician or other medical practitioner using standard techniques and within the normal course of their work. Non‐ limiting examples of suitable dosages of viral vectors of the invention include 1x108 transduction units (TU), 1x109 TU, 1x1010 TU, 1x1011 TU or more. Non‐limiting examples of suitable dosages of non‐viral vectors/delivery means of the invention include a maximum of 30 mL per dose, a maximum of 25 mL per dose, a maximum of 20 mL per dose, a maximum of 15 mL per dose, a maximum of 10 mL per dose, or less, preferably a maximum of 20 mL per dose. Non‐limiting examples of pharmaceutically acceptable carriers that may be comprised in a composition of the invention include water, saline, and phosphate‐buffered saline. In some embodiments, however, the composition is in lyophilized form, in which case it may include a stabilizer, such as bovine serum albumin (BSA). In some embodiments, it may be desirable to formulate the composition with a preservative, such as thiomersal or sodium azide, to facilitate long‐ term storage. The nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered by any appropriate route. It may be desired to direct the compositions of the present invention (as described above) to the respiratory system of a subject. Efficient transmission of a therapeutic/prophylactic composition or medicament to the site of a disease or disorder in the respiratory tract may be achieved by oral or intra‐nasal administration, for example, as aerosols (e.g. nasal sprays), or by catheters. Typically the nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention are stable in clinically relevant nebulisers, inhalers (including metered dose inhalers), catheters and aerosols, etc.
Other routes of administration, including but not limited to i.v. administration, intranasal administration and intraplural injection are also encompassed by the present invention. Suitable administration routes are known in the art. In some embodiments the nose is a preferred production site for a therapeutic protein using nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention for at least one of the following reasons: (i) extracellular barriers such as inflammatory cells and sputum are less pronounced in the nose; (ii) ease of vector administration; (iii) smaller quantities of vector required; and (iv) ethical considerations. Thus, nasal administration of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may result in efficient (high‐level) and long‐lasting expression of the therapeutic protein of interest. Accordingly, nasal administration of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be preferred. Formulations for intra‐nasal administration may be in the form of nasal droplets or a nasal spray. An intra‐nasal formulation may comprise droplets having approximate diameters in the range of 100‐5000 µm, such as 500‐4000 µm, 1000‐3000 µm or 100‐1000 µm. Alternatively, in terms of volume, the droplets may be in the range of about 0.001‐100 µl, such as 0.1‐50 µl or 1.0‐25 µl, or such as 0.001‐1 µl. The aerosol formulation may take the form of a powder, suspension or solution. The size of aerosol particles is relevant to the delivery capability of an aerosol. Smaller particles may travel further down the respiratory airway towards the alveoli than would larger particles. In one embodiment, the aerosol particles have a diameter distribution to facilitate delivery along the entire length of the bronchi, bronchioles, and alveoli. Alternatively, the particle size distribution may be selected to target a particular section of the respiratory airway, for example the alveoli. In the case of aerosol delivery of the medicament, the particles may have diameters in the approximate range of 0.1‐50 µm, preferably 1‐25 µm, more preferably 1‐5 µm. Aerosol particles may be for delivery using a nebulizer (e.g. via the mouth) or nasal spray. An aerosol formulation may optionally contain a propellant and/or surfactant. The formulation of pharmaceutical aerosols is routine to those skilled in the art, see for example, Sciarra, J. in Remington's Pharmaceutical Sciences (supra). The agents may be formulated as solution aerosols, dispersion or suspension aerosols of dry powders, emulsions or semisolid preparations. The aerosol may be delivered using any propellant system known to those skilled in the art. The aerosols may be applied to the upper respiratory tract, for example by nasal inhalation, or to the lower respiratory tract or to both. The part of the lung that the medicament is delivered to may
be determined by the disorder. Compositions comprising nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention, in particular where intranasal delivery is to be used, may comprise a humectant. This may help reduce or prevent drying of the mucus membrane and to prevent irritation of the membranes. Suitable humectants include, for instance, sorbitol, mineral oil, vegetable oil and glycerol; soothing agents; membrane conditioners; sweeteners; and combinations thereof. The compositions may comprise a surfactant. Suitable surfactants include non‐ionic, anionic and cationic surfactants. Examples of surfactants that may be used include, for example, polyoxyethylene derivatives of fatty acid partial esters of sorbitol anhydrides, such as for example, Tween 80, Polyoxyl 40 Stearate, Polyoxy ethylene 50 Stearate, fusieates, bile salts and Octoxynol. In some cases after an initial administration a subsequent administration of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be performed. The administration may, for instance, be at least a week, two weeks, a month, two months, six months, a year or more after the initial administration. In some instances, nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered at least once a week, once a fortnight, once a month, every two months, every six months, annually or at longer intervals. Preferably, administration is every six months, more preferably annually. The nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors may, for instance, be administered at intervals dictated by when the effects of the previous administration are decreasing. Further, for the avoidance of doubt, any and all disclosure herein in relation formulations of nucleic acid cassettes or vectors of the present invention, applies equally and without reservation to the formulations of lentiviral (e.g. SIV) vectors for combination therapies of lentiviral (e.g. SIV) vectors and surfactants as described herein. SEQUENCE HOMOLOGY Any of a variety of sequence alignment methods can be used to determine percent identity, including, without limitation, global methods, local methods and hybrid methods, such as, e.g., segment approach methods. Protocols to determine percent identity are routine procedures within the scope of one skilled in the art. Global methods align sequences from the beginning to the end of the molecule and determine the best alignment by adding up scores of individual residue pairs and by imposing gap penalties. Non‐limiting methods include, e.g., CLUSTAL W, see, e.g., Julie D. Thompson et al., CLUSTAL W: Improving the Sensitivity of Progressive Multiple Sequence Alignment Through Sequence Weighting, Position‐ Specific Gap Penalties and Weight Matrix Choice, 22(22) Nucleic Acids
Research 4673‐4680 (1994); and iterative refinement, see, e.g., Osamu Gotoh, Significant Improvement in Accuracy of Multiple Protein. Sequence Alignments by Iterative Refinement as Assessed by Reference to Structural Alignments, 264(4) J. MoI. Biol. 823‐838 (1996). Local methods align sequences by identifying one or more conserved motifs shared by all of the input sequences. Non‐limiting methods include, e.g., Match‐box, see, e.g., Eric Depiereux and Ernest Feytmans, Match‐ Box: A Fundamentally New Algorithm for the Simultaneous Alignment of Several Protein Sequences, 8(5) CABIOS 501 ‐509 (1992); Gibbs sampling, see, e.g., C. E. Lawrence et al., Detecting Subtle Sequence Signals: A Gibbs Sampling Strategy for Multiple Alignment, 262(5131 ) Science 208‐214 (1993); Align‐M, see, e.g., Ivo Van WaIIe et al., Align‐M ‐ A New Algorithm for Multiple Alignment of Highly Divergent Sequences, 20(9) Bioinformatics:1428‐1435 (2004). Thus, percent sequence identity is determined by conventional methods. See, for example, Altschul et al., Bull. Math. Bio. 48: 603‐16, 1986 and Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89:10915‐19, 1992. Briefly, two amino acid sequences are aligned to optimize the alignment scores using a gap opening penalty of 10, a gap extension penalty of 1, and the "blosum 62" scoring matrix of Henikoff and Henikoff (ibid.) as shown below (amino acids are indicated by the standard one‐letter codes). The "percent sequence identity" between two or more nucleic acid or amino acid sequences is a function of the number of identical positions shared by the sequences. Thus, % identity may be calculated as the number of identical nucleotides / amino acids divided by the total number of nucleotides / amino acids, multiplied by 100. Calculations of % sequence identity may also take into account the number of gaps, and the length of each gap that needs to be introduced to optimize alignment of two or more sequences. Sequence comparisons and the determination of percent identity between two or more sequences can be carried out using specific mathematical algorithms, such as BLAST, which will be familiar to a skilled person. ALIGNMENT SCORES FOR DETERMINING SEQUENCE IDENTITY A R N D C Q E G H I L K M F P S T W Y V A 4 R ‐1 5 N ‐2 0 6 D ‐2 ‐2 1 6 C 0 ‐3 ‐3 ‐3 9 Q ‐1 1 0 0 ‐3 5 E ‐1 0 0 2 ‐4 2 5
G 0 ‐2 0 ‐1 ‐3 ‐2 ‐2 6 H ‐2 0 1 ‐1 ‐3 0 0 ‐2 8 I ‐1 ‐3 ‐3 ‐3 ‐1 ‐3 ‐3 ‐4 ‐3 4 L ‐1 ‐2 ‐3 ‐4 ‐1 ‐2 ‐3 ‐4 ‐3 2 4 K ‐1 2 0 ‐1 ‐3 1 1 ‐2 ‐1 ‐3 ‐2 5 M ‐1 ‐1 ‐2 ‐3 ‐1 0 ‐2 ‐3 ‐2 1 2 ‐1 5 F ‐2 ‐3 ‐3 ‐3 ‐2 ‐3 ‐3 ‐3 ‐1 0 0 ‐3 0 6 P ‐1 ‐2 ‐2 ‐1 ‐3 ‐1 ‐1 ‐2 ‐2 ‐3 ‐3 ‐1 ‐2 ‐4 7 S 1 ‐1 1 0 ‐1 0 0 0 ‐1 ‐2 ‐2 0 ‐1 ‐2 ‐1 4 T 0 ‐1 0 ‐1 ‐1 ‐1 ‐1 ‐2 ‐2 ‐1 ‐1 ‐1 ‐1 ‐2 ‐1 1 5 W ‐3 ‐3 ‐4 ‐4 ‐2 ‐2 ‐3 ‐2 ‐2 ‐3 ‐2 ‐3 ‐1 1 ‐4 ‐3 ‐2 11 Y ‐2 ‐2 ‐2 ‐3 ‐2 ‐1 ‐2 ‐3 2 ‐1 ‐1 ‐2 ‐1 3 ‐3 ‐2 ‐2 2 7 V 0 ‐3 ‐3 ‐3 ‐1 ‐2 ‐2 ‐3 ‐3 3 1 ‐2 1 ‐1 ‐2 ‐2 0 ‐3 ‐1 4 The percent identity is then calculated as: Total number of identical matches __________________________________________ x 100 [length of the longer sequence plus the number of gaps introduced into the longer sequence in order to align the two sequences] Substantially homologous polypeptides are characterized as having one or more amino acid substitutions, deletions or additions. These changes are preferably of a minor nature, that is conservative amino acid substitutions (as described herein) and other substitutions that do not significantly affect the folding or activity of the polypeptide; small deletions, typically of one to about 30 amino acids; and small amino‐ or carboxyl‐terminal extensions, such as an amino‐terminal methionine residue, a small linker peptide of up to about 20‐25 residues, or an affinity tag. In addition to the 20 standard amino acids, non‐standard amino acids (such as 4‐ hydroxyproline, 6‐N‐methyl lysine, 2‐aminoisobutyric acid, isovaline and α ‐methyl serine) may be substituted for amino acid residues of the polypeptides of the present invention. A limited number of non‐conservative amino acids, amino acids that are not encoded by the genetic code, and unnatural amino acids may be substituted for polypeptide amino acid residues. The polypeptides of the present invention can also comprise non‐naturally occurring amino acid residues.
Non‐naturally occurring amino acids include, without limitation, trans‐3‐methylproline, 2,4‐ methano‐proline, cis‐4‐hydroxyproline, trans‐4‐hydroxy‐proline, N‐methylglycine, allo‐threonine, methyl‐threonine, hydroxy‐ethylcysteine, hydroxyethylhomo‐cysteine, nitro‐glutamine, homoglutamine, pipecolic acid, tert‐leucine, norvaline, 2‐azaphenylalanine, 3‐azaphenyl‐alanine, 4‐ azaphenyl‐alanine, and 4‐fluorophenylalanine. Several methods are known in the art for incorporating non‐naturally occurring amino acid residues into proteins. For example, an in vitro system can be employed wherein nonsense mutations are suppressed using chemically aminoacylated suppressor tRNAs. Methods for synthesizing amino acids and aminoacylating tRNA are known in the art. Transcription and translation of plasmids containing nonsense mutations is carried out in a cell free system comprising an E. coli S30 extract and commercially available enzymes and other reagents. Proteins are purified by chromatography. See, for example, Robertson et al., J. Am. Chem. Soc. 113:2722, 1991; Ellman et al., Methods Enzymol. 202:301, 1991; Chung et al., Science 259:806‐9, 1993; and Chung et al., Proc. Natl. Acad. Sci. USA 90:10145‐9, 1993). In a second method, translation is carried out in Xenopus oocytes by microinjection of mutated mRNA and chemically aminoacylated suppressor tRNAs (Turcatti et al., J. Biol. Chem. 271:19991‐8, 1996). Within a third method, E. coli cells are cultured in the absence of a natural amino acid that is to be replaced (e.g., phenylalanine) and in the presence of the desired non‐naturally occurring amino acid(s) (e.g., 2‐azaphenylalanine, 3‐ azaphenylalanine, 4‐azaphenylalanine, or 4‐fluorophenylalanine). The non‐naturally occurring amino acid is incorporated into the polypeptide in place of its natural counterpart. See, Koide et al., Biochem. 33:7470‐6, 1994. Naturally occurring amino acid residues can be converted to non‐naturally occurring species by in vitro chemical modification. Chemical modification can be combined with site‐directed mutagenesis to further expand the range of substitutions (Wynn and Richards, Protein Sci. 2:395‐403, 1993). A limited number of non‐conservative amino acids, amino acids that are not encoded by the genetic code, non‐naturally occurring amino acids, and unnatural amino acids may be substituted for amino acid residues of polypeptides of the present invention. Essential amino acids in the polypeptides of the present invention can be identified according to procedures known in the art, such as site‐directed mutagenesis or alanine‐scanning mutagenesis (Cunningham and Wells, Science 244: 1081‐5, 1989). Sites of biological interaction can also be determined by physical analysis of structure, as determined by such techniques as nuclear magnetic resonance, crystallography, electron diffraction or photoaffinity labeling, in conjunction with mutation of putative contact site amino acids. See, for example, de Vos et al., Science 255:306‐12, 1992; Smith et al., J. Mol. Biol. 224:899‐904, 1992; Wlodaver et al., FEBS Lett. 309:59‐64, 1992. The identities of
essential amino acids can also be inferred from analysis of homologies with related components (e.g. the translocation or protease components) of the polypeptides of the present invention. Multiple amino acid substitutions can be made and tested using known methods of mutagenesis and screening, such as those disclosed by Reidhaar‐Olson and Sauer (Science 241:53‐7, 1988) or Bowie and Sauer (Proc. Natl. Acad. Sci. USA 86:2152‐6, 1989). Briefly, these authors disclose methods for simultaneously randomizing two or more positions in a polypeptide, selecting for functional polypeptide, and then sequencing the mutagenized polypeptides to determine the spectrum of allowable substitutions at each position. Other methods that can be used include phage display (e.g., Lowman et al., Biochem. 30:10832‐7, 1991; Ladner et al., U.S. Patent No. 5,223,409; Huse, WIPO Publication WO 92/06204) and region‐directed mutagenesis (Derbyshire et al., Gene 46:145, 1986; Ner et al., DNA 7:127, 1988). Multiple amino acid substitutions can be made and tested using known methods of mutagenesis and screening, such as those disclosed by Reidhaar‐Olson and Sauer (Science 241:53‐7, 1988) or Bowie and Sauer (Proc. Natl. Acad. Sci. USA 86:2152‐6, 1989). Briefly, these authors disclose methods for simultaneously randomizing two or more positions in a polypeptide, selecting for functional polypeptide, and then sequencing the mutagenized polypeptides to determine the spectrum of allowable substitutions at each position. Other methods that can be used include phage display (e.g., Lowman et al., Biochem. 30:10832‐7, 1991; Ladner et al., U.S. Patent No. 5,223,409; Huse, WIPO Publication WO 92/06204) and region‐directed mutagenesis (Derbyshire et al., Gene 46:145, 1986; Ner et al., DNA 7:127, 1988). SEQUENCE INFORMATION Key to Sequences SEQ ID NO: 1 Core SFTPB promoter (mSPB) SEQ ID NO: 2 972bp SFTPB genomic promoter sequence SEQ ID NO: 3 5’ fragment portion of SFTPB exon 1 SEQ ID NO: 4 full length SFTPB promoter sequence (fSPB) SEQ ID NO: 5 Exemplary hCEF promoter SEQ ID NO: 6 Exemplary CMV promoter SEQ ID NO: 7 Exemplary EF1aS promoter SEQ ID NO: 8 Exemplary PGK promoter SEQ ID NO: 9 Exemplary core SFTPC promoter (mSPC)
SEQ ID NO: 10 full length SFTPC promoter sequence (fSPC) SEQ ID NO: 11 CpG‐free CMV enhancer forward SEQ ID NO: 12 CpG‐free CMV enhancer reverse SEQ ID NO: 13 ELF3‐1 enhancer forward SEQ ID NO: 14 ELF3‐1 enhancer reverse SEQ ID NO: 15 ELF3‐2 enhancer forward SEQ ID NO: 16 ELF3‐2 enhancer reverse SEQ ID NO: 17 ELF3‐3 enhancer forward SEQ ID NO: 18 ELF3‐3 enhancer reverse SEQ ID NO: 19 ELF3‐4 enhancer forward SEQ ID NO: 20 ELF3‐4 enhancer reverse SEQ ID NO: 21 actin enhancer forward SEQ ID NO: 22 actin enhancer reverse SEQ ID NO: 23 LMO7‐1 enhancer forward SEQ ID NO: 24 LMO7‐1 enhancer reverse SEQ ID NO: 25 LMO7‐2 enhancer reverse SEQ ID NO: 26 SFTPB enhancer forward SEQ ID NO: 27 SFTPC enhancer forward SEQ ID NO: 28 SFTPC enhancer reverse SEQ ID NO: 29 SLC34A2‐1 enhancer forward SEQ ID NO: 30 SLC34A2‐1 enhancer reverse SEQ ID NO: 31 SLC34A2‐2 enhancer forward SEQ ID NO: 32 SLC34A2‐2 enhancer reverse SEQ ID NO: 33 SLC34A2‐3 enhancer reverse SEQ ID NO: 34 SV40 enhancer forward SEQ ID NO: 35 SV40 enhancer reverse SEQ ID NO: 36 VEGFA‐1 enhancer forward SEQ ID NO: 37 VEGFA‐1 enhancer reverse SEQ ID NO: 38 VEGFA‐2 enhancer forward SEQ ID NO: 39 VEGFA‐3 enhancer forward SEQ ID NO: 40 VEGFA‐3 enhancer reverse SEQ ID NO: 41 SLC34A2 enhancer + core SFTPB promoter SEQ ID NO: 42 VEGFA enhancer + core SFTPB promoter SEQ ID NO: 43 CpG‐free CMV enhancer + core SFTPB promoter
SEQ ID NO: 44 ELF3 enhancer + core SFTPB promoter SEQ ID NO: 45 SV40 enhancer + core SFTPB promoter SEQ ID NO: 46 Alv‐01 (CMV forward enhancer + mSFPB promoter) SEQ ID NO: 47 Alv‐02 (SLC34A2‐2 forward enhancer + mSFPB promoter) SEQ ID NO: 48 Alv‐03 (VEGFA‐2 forward enhancer + mSFPB promoter) SEQ ID NO: 49 SFTPB promoter forward primer SEQ ID NO: 50 SFTPB promoter reverse primer SEQ ID NO: 51 Exemplary linker between SFTPB promoter fragment and enhancer SEQ ID NO: 52 Exemplary WPRE component (mWPRE) SEQ ID NO: 53 Exemplary SFTPB transgene SEQ ID NO: 54 human surfactant protein B (hSP‐B) transgene SEQ ID NO: 55 codon‐optimised hSP‐B transgene SEQ ID NO: 56 Homo sapiens surfactant protein B (SFTPB), RefSeqGene on chromosome 2. (NCBI Reference Sequence: NG_016967.1) (18425bp) SEQ ID NO: 57 Exemplary SFTPB polypeptide SEQ ID NO: 58 Exemplified ABACA3 (ABCA3) transgene SEQ ID NO: 59 Exemplified human ABACA3 (hABCA3) transgene SEQ ID NO: 60 codon‐optimised hABCA3 transgene SEQ ID NO: 61 Exemplified Human ABCA3 polypeptide SEQ ID NO: 62 Exemplary SFTPC transgene SEQ ID NO: 63 Exemplary SFTPC polypeptide SEQ ID NO: 64 Exemplary hGM‐CSF transgene SEQ ID NO: 65 Exemplary hGM‐CSF polypeptide SEQ ID NO: 66 Exemplary Human DCN (Decorin) transgene SEQ ID NO: 67 Exemplary Human Decorin polypeptide SEQ ID NO: 68 Exemplary Human TRIM72 transgene SEQ ID NO: 69 Exemplary Human TRIM72 polypeptide SEQ ID NO: 70 Exemplary AAT transgene (SERPINA1) SEQ ID NO: 71 Exemplary A1A1 polypeptide SEQ ID NO: 72 Exemplary FVIII transgene (N6) SEQ ID NO: 73 Exemplary FVIII transgene (V3) SEQ ID NO: 74 Exemplary FVIII polypeptide (N6) SEQ ID NO: 75 Exemplary FVIII polypeptide (V3)
SEQ ID NO: 76 corresponds to SLC34A2 enhancer + core SFTPB promoter of SEQ ID NO: 41 without restriction site SEQ ID NO: 77 corresponds to VEGFA enhancer + core SFTPB promoter of SEQ ID NO: 4241 without restriction site SEQ ID NO: 78 corresponds to CpG‐free CMV enhancer + core SFTPB promoter of SEQ ID NO: 43 without restriction site SEQ ID NO: 79 corresponds to ELF3 enhancer + core SFTPB promoter of SEQ ID NO: 44 without restriction site SEQ ID NO: 80 corresponds to SV40 enhancer + core SFTPB promoter of SEQ ID NO: 45 without restriction site SEQ ID NO: 81 corresponds to Alv‐01 (CMV forward enhancer + mSFPB promoter) of SEQ ID NO: 46 without restriction site SEQ ID NO: 82 corresponds to Alv‐02 (SLC34A2‐2 forward enhancer + mSFPB promoter) of SEQ ID NO: 47 without restriction site SEQ ID NO: 83 corresponds to Alv‐03 (VEGFA‐2 forward enhancer + mSFPB promoter) of SEQ ID NO: 48 without restriction site SEQ ID NOs: 84‐86 correspond to the exemplary SFTPC transgene of SEQ ID NO: 62 without either/both restriction site SEQ ID NO: 1 core SFTPB Promoter* (635bp) TATAGGGCTGTCTGGGAGCCACTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGGGTATTTTGTTTTCTGTTT TGTGTAAATGCTCTTCTGACTAATGCAAACCATGTGTCCATAGAACCAGAAGATTTTTCCAGGGGAAAAGGTA AGGAGGTGGTGAGAGTGTCCTGGGTCTGCCCTTCCAGGGCTTGCCCTGGGTTAAGAGCCAGGCAGGAAGCTC TCAAGAGCATTGCTCAAGAGTAGAGGGGGCCTGGGAGGCCCAGGGAGGGGATGGGAGGGGAACACCCAGG CTGCCCCCAACCAGATGCCCTCCACCCTCCTCAACCTCCCTCCCACGGCCTGGAGAGGTGGGACCAGGTATGG AGGCTTGAGAGCCCCTGGTTGGAGGAAGCCACAAGTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGG AGGCAGGAACAGGCCATCAGCCAGGACAGGTGGTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGG GATCAAGCACCTGGAGGGCTCTTCAGAGCAAAGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCA CGCCCCGCCCAGCTATAAGGGGCCATGCMCCAAGCAGGGTACCCAGGCTGCAGAGGGTGCC *wherein M is C or A, preferably M is C SEQ ID NO: 2 972bp SFTPB genomic promoter sequence – bases 4845‐5816 of NG_016967.1 (972bp) TTCTTTCTGCTGAACCATCGCAGCTATGCCCCAGCCCCTACCCTGGAGGGGTCCCCAGGGGCCATGGG CAGCACCTCCTGTATAGGGCTGTCTGGGAGCCACTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGG
GTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTAATGCAAACCATGTGTCCATAGAACCAGA AGATTTTTCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCTGGGTCTGCCCTTCCAGGGCTTGCC CTGGGTTAAGAGCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGTAGAGGGGGCCTGGGAGGCCC AGGGAGGGGATGGGAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCTCCACCCTCCTCAACCTCC CTCCCACGGCCTGGAGAGGTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTGGAGGAAGCCACAA GTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGCCAGGACAGGTG GTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCTTCAGAGCA AAGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGGCCAT GCACCAAGCAGGGTACCCAGGCTGCAGAGGTGCCATGGCTGAGTCACACCTGCTGCAGTGGCTGCTGC TGCTGCTGCCCACGCTCTGTGGCCCAGGCACTGGTGAGTCTCCCCCAGCCTCCCCTCTCCTAGGCAGC TCCACCACTCACTGAGCACTGCTTTGTGCTAGGCATTAACCCAAGTCTGTCCTCATTTTAAAGACAAG GCAGCTGGGGTTCAGAGAGGGTTCAGAGCTTATCCAAGGTCACACAGCTGGCGGGTCCAGGAGCAGGT GGAACCCAGAGCTGTCTGAC The bold and underlined 5’ and 3’ sequences were used in the art to design primers to this 972bp SFTPB promoter. Exon 1 of the SFTPB gene (as described by NG_016967.1) is dash‐underlined (corresponding to bases 5543‐5565 of NG_016967.1). The first A of this exon is the starting nucleotide of the mRNA generated by the SFTPB promoter. The bold and italicised ATG (corresponding to bases 5559‐5561 of NG_016967.1) encodes the first methionine of pre‐pro‐SFTPB. 3’ of the double‐underlined exon 1 is a partial portion of intron 1 (corresponding to bases 5626‐5816 of NG_016967.1). The first base (base 81 of SEQ ID NO: 2) and last base (base 710 of SEQ ID NO: 2) of the core SFTP promoter fragment of the invention are double‐underlined. The wavy‐underlined sequence is the predicted TATA box of the SFTPB promoter. SEQ ID NO: 3 5’ fragment portion of SFTPB exon 1 AGGCTGCAGAGG
SEQ ID NO: 4 full length SFTPB promoter sequence (1054bp) GGATCCTCCCTCCTCGGCCTCCCAAAGTGCCAGGATTACAGGAGTGAGCCACCACACCCAGCCCCATCTCTTTT CATCATGGTACTAATTCCTGCCCGTCCACCCACAAAAGCACTGTAGTCGTTCCCGAGTATAGAGGCCTGTGAG CCTCCACTAGGGAGAGGGCTCCTGCAGAGATCAGATAAATTGATCACAATGGCTGGGGTGGTGGCAATGTGC TAATGCTCTCTTTCTTCCACTCAAGATATCCTCTGTCTCCCTCAGCCTGTGAGCTTTTTCTCCAGTGTGCTCTGCC AGTGGGGGCCTTGCCTGAGAGCCCCTGCAGCTGCAGAGGACAGTTTCTTTCTGCTGAACCATCGCAGCTATGC CCCAGCCCCTACCCTGGAGGGGTCCCCAGGGGCCATGGGCAGCACCTCCTGTATAGGGCTGTCTGGGAGCCA CTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGGGTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTA ATGCAAACCATGTGTCCATAGAACCAGAAGATTTTTCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCT GGGTCTGCCCTTCCAGGGCTTGCCCTGGGTTAAGAGCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGT AGAGGGGGCCTGGGAGGCCCAGGGAGGGGATGGGAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCT CCACCCTCCTCAACCTCCCTCCCACGGCCTGGAGAGGTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTG GAGGAAGCCACAAGTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGC CAGGACAGGTGGTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCT TCAGAGCAAAGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGG CCATGCCCCAAGCAGGGTACCCAGGCTGCAGAGGGTGCC SEQ ID NO: 5 Exemplified hCEF promoter (562bp) GTTACATAACTTATGGTAAATGGCCTGCCTGGCTGACTGCCCAATGACCCCTGCCCAATGATGTCAATAATGAT GTATGTTCCCATGTAATGCCAATAGGGACTTTCCATTGATGTCAATGGGTGGAGTATTTATGGTAACTGCCCAC TTGGCAGTACATCAAGTGTATCATATGCCAAGTATGCCCCCTATTGATGTCAATGATGGTAAATGGCCTGCCTG GCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTATGTATTAGTCATTGCTATTAC CATGGGAATTCACTAGTGGAGAAGAGCATGCTTGAGGGCTGAGTGCCCCTCAGTGGGCAGAGAGCACATGG CCCACAGTCCCTGAGAAGTTGGGGGGAGGGGTGGGCAATTGAACTGGTGCCTAGAGAAGGTGGGGCTTGGG TAAACTGGGAAAGTGATGTGGTGTACTGGCTCCACCTTTTTCCCCAGGGTGGGGGAGAACCATATATAAGTGC AGTAGTCTCTGTGAACATTCAAGCTTCTGCCTTCTCCCTCCTGTGAGTTT SEQ ID NO: 6 Exemplified CMV promoter (855bp) CAATATTGGCCATTAGCCATATTATTCATTGGTTATATAGCATAAATCAATATTGGCTATTGGCCATTGCATACG TTGTATCTATATCATAATATGTACATTTATATTGGCTCATGTCCAATATGACCGCCATGTTGGCATTGATTATTG ACTAGTTATTAATAGTAATCAATTACGGGGTCATTAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACT TACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCC ATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAG TACATCAAGTGTATCATATGCCAAGTCCGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTAT
GCCCAGTACATGACCTTACGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGT GATGCGGTTTTGGCAGTACACCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCC ATTGACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAATAACCCCGCCCC GTTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCA GATCACTAGAAGCTTTATTGCGGTAGTTTATCACAGTTAAATTGCTAACGCAGTCAGTGCTTCTGACACAACAG TCTCGAACTTAAGCTGCAGAAGTTGGTCGTGAGGCACTGGGCAG SEQ ID NO: 7 Exemplary EF1aS promoter (236bp) CGTGAGGCTCCGGTGCCCGTCAGTGGGCAGAGCGCACATCGCCCACAGTCCCCGAGAAGTTGGGGGGAGGG GTCGGCAATTGAACCGGTGCCTAGAGAAGGTGGCGCGGGGTAAACTGGGAAAGTGATGTCGTGTACTGGCT CCGCCTTTTTCCCGAGGGTGGGGGAGAACCGTATATAAGTGCAGTAGTCGCCGTGAACGTTCTTTTTCGCAAC GGGTTTGCCGCCAGAACACAG SEQ ID NO: 8 Exemplary PGK promoter (529bp) GATCTTCGAATTCCCACGGGGTTGGGGTTGCGCCTTTTCCAAGGCAGCCCTGGGTTTGCGCAGGGACGCGGC TGCTCTGGGCGTGGTTCCGGGAAACGCAGCGGCGCCGACCCTGGGTCTCGCACATTCTTCACGTCCGTTCGCA GCGTCACCCGGATCTTCGCCGCTACCCTTGTGGGCCCCCCGGCGACGCTTCCTGCTCCGCCCCTAAGTCGGGA AGGTTCCTTGCGGTTCGCGGCGTGCCGGACGTGACAAACGGAAGCCGCACGTCTCACTAGTACCCTCGCAGA CGGACAGCGCCAGGGAGCAATGGCAGCGCGCCGACCGCGATGGGCTGTGGCCAATAGCGGCTGCTCAGCAG GGCGCGCCGAGAGCAGCGGCCGGGAAGGGGCGGTGCGGGAGGCGGGGTGTGGGGCGGTAGTGTGGGCCC TGTTCCTGCCCGCGCGGTGTTCCGCATTCTGCAAGCCTCCGGAGCGCACGTCGGCAGTCGGCTCCCTCGTTGA CCGAATCACCGACCTCTCTCCCCAGG SEQ ID NO: 9 Exemplary core SFTPC promoter (345bp) CAGGGCAGCAGGGGCAGGTGCCAGCAAGGAAGGCAGGCACGCCAGGAAGACACCCATGGTGAGAAGTGCA GATGGCCCGAGGGCAAGTTTGCTCAACTCACCCAGGTTTGCTCTTGCTGGGGCCAAGAGGACTCATGTGCCA GGGCCAAGGGCCCTTGGGGGCTCTCACAGGGGGCTTATCTGGGCTTCGGTTCTGGAGGGCCAGGAACAAAC AGGCTTCAAAGCCAAGGGCTTGGCTGGCACACAGGGGGCTTGGTCCTTCACCTCTGTCCCCTCTCCCTACGGA CACATATAAGACCCTGGTCACACCTGGGAGAGGAGGAGAGGAGAGCATAGCACCTGCAG SEQ ID NO: 10 full length SFTPC promoter sequence (3813bp) AGCTTGAAGACTGCTGCTCTCTACCACGTTAGCTCCCCTGTGGCGGAGGATGACTGTCCTAAGAGCTCATAGG CACGGGGCAAGGGGAGCCTGGCTGTCAGCTGCCTGGGCTTCTAGTCTTGTGCTTTTTGCACACCAGTCCAGGG AACGAAGACCACTGGCTTTAAGATGCTTCCCCAGCTGTCCCCAGACTCTGCCAGCAGGGGATTCTCTGGTCTG
AGCTTAAGTTGTGTTCTCCCAGCCAGGGATGCCCCTGCCCTTTGATGTCTCCTTGCTGCCACACATTTAGCCGC CCTCCCCATGCCAGCTTGGGGGGAGGGAAGCAGTGAGGGTAGGGAGGTGGCTGGGGCAGCTGGGCAACTG TCCCCACCCGTCCCTGGCACGGCTCTGCCCAGTACACAAAGAGCAAAGTGAATCTTGTCCCCACCCCTGCAGCT GAGGGGCTGGAGGAGGAAACGGGGAGGCCCACACAGAAGGGGTGGCCACCGTGGGGCTGTCCATCACTCA GGGCTCTCAGAGGGAGTCAACCCAGAAACAGACAAAGAGGGTGAGTCTGGGCTGTGTTCTTAGCTAGTGAG AGGTCCCCTAGAGGATGAAGTAGATGATGCTAATGAGGATGACTGGATGTCACACCCATGATGCTATTAGGT CCTCATAATAGCATAGTGAGGTGGACAGCTAGTACCTGACCCATCTCACAGATGAACATACTAATGCCTAACA AAGCAGAACAACTCACGCTGGGTCCCAGAGCTGGCCAGTGGAAGCACTGAGACCTCCACATACTGAAGGCAT GGACTATTGACCGCTGTTGGTATTGGTCTCATCATTGACTATCATTAAGTGTTGGCTGTTTGCATGCTTCCTGCC CAGTGGCAGGTTCAAAGAAGCCCGCAGGAAGCGTGCTCCTTTCTTTCCCAGGGCCCGCAATTGGGCTGGAAG ATAGAGCAACAAAAAGCGCCCATGTAACTCATGGGAACATTCATGTGTGCTGAATGGCAGGTGAAGGTGCCA CAGAGAGGCTGAGGATTTCAGAGGGCACCATGAACTGGAGTGAGGTCGCAGAGCAGGTGCCATTGGCTCTT GGCCTGTTTGGGTGGGTGGCATTCAGAGAGGTGGAAGGTCAGATGCACTGTTCACGCCTGTAATCTCAGCAC TTTGGAAGGCCAAGGTAGGAGGATCACACGAGGCCAGGAGATCAACGCTGCAGTGAGCTATGAAGCTGTGA TTGCACCACTGCACTGCAGCTTGGGTAACAGAGTGAGACCCTGTCTCTAAATAATTAAATAAATAAAATAAAA ATAAAACCGGAGAAGTGGAGAGGGATTGGAGGTGGGCTTTCACAGAGGGAGAAACGGCTTAAGTACAGGC CAAAAAGTGAGAAGGCTGCAGACAGGGCTGGTAGGGGGAGGGGGAAATTTGGCACACCCAGCTAAAGGTC CTTCTGTGTCTCCTTCTCCAAGGAACCCAAGACCTTCACTTGGTTGGTGTGAGCACTCCAGGAGGCAGGCACCC TCCCTCAGCCCTCAAGCAAGCAAAAATGGGTTTAAAAAAAGAAGGAGAAGAAGCAGCAGCAGCAGCCGCCA CAGAGCTTGTGACAGCTACAGCCTAAGGGCAACAGGCAGGGGAGACCAAGGACCAGAAAGAGCAGAGGCTT TTTCAAAGAAAGAGATCCCTCTCCCAGCACCCAGCGATGGCGGCAAACCCCACCCACAGTGCCTGCTAAGAAC AAGTCCCACGTGAGAACAACATGGCCCCCCGAGATGCCCACAGGGACCCCGAGATGCCTGCAGTGCTTGGCT CTCCCGCTGGCCAGCTGCCCACCTGGCTCAGGCCCAGTACTCGTGAGTCAGCCGATCAATCCCAATGTTGCCA GGATGATGGGGGCGGGAGTAAGGGCCCTGGGGGAGGGCAGGGGTGGGCACTGCAGGCGAGCTGTCTCCCA CATCTGGCACCTGCACACAGCTGAGGCCGAGCTGAGAGGATGCTTCTGTGGGCTCCCCCTCCTCCCGGCACCC CTCCCCTCCTTTCACTGTCCTAGGACACTCTCTGGCTGCTGGAGTCTTAGGCAAATATTTAAAGGGGCAGCAAG GGGGTGAGGAGGGTGGTGGGAGCAAACACTTCCCTCCCTTTCTTCCTCCCTGCGCTTCTCAGGGGCTCTCAGT TCAGATGCCATGCTGTTATGCAACCTTGGGGCTGAAGGCCCTCCAGATGGAGAGGGGGACAGGGGAACCTG CCAGCTCATGACCGAAGGGCAGGGCCCAGGTGGGAGGGGGCTGGGGCAGGGGACAGGAACTGGGGTGGC ATGTTAAAGGACAGGAGGCTGGTTGGGCATGGTGGCTCACACCTGTAATCCTAGCACTTTAGGAGGCCGAGA TCACTTGAGCCCAGGAGTTCAAGACCAGCCTGGGCAACATGGTGAAAACTCATCTCTATAAAACAAGCAAAAA TTAGCTGGGCACAATGGCATGCACTGGTAGTCCCAGCTACTTGGGATGCTGAGGTGTGAGGATCACCGGAGT CCAGGAGGTCAACGCTGCAGTGAGCAGTGATCTCGCTACTGCATGCCAGCCTGGGTGATAAAGTGAGACCCT GTCTCACAACAAAACAAAACAAAACAAAACAAAAGGATAGGAGGTTAAGGGAGCGAGCCCAGGCCTGGACT
CTGCCACAGTCACCTGAGTTTAGAGGTGAAGGGACTTTAGAGACCACCTGGCCCAAGGAGTGAAAGGGAACC TGAGAGAAAGGGTGAGCCAGCCCAAAGTCATTGGCAGATTTGACCTCGTAAATACATAGAGATGGCTTTGGG AAGGCACTAGGAAAGACAGAGAAAAGAGAAGGAGACAGTCCTCAAAGCTGATCGTATTTGGGTGAGTATTA TTCTCAGGGCAAATTTAGGATCAGAGGATGCAGAAAGGGGAGTCTAGAGGGGTAGAGTGTAGACCACAGGG TGAGTGAGCTGATTCGAGGATGGGGAGACTGGGAGCCCACCAGTGACCAGAGCCAGCCCTGTTCAGGGCTG TCCGGGCAGAAGAAAGCAGTGTCAGACCTGGAATCTGCCATCAGCACAGCCTGCAATTGACAGACAAGCCCA GAGCAAAGAAGGAAGCACTGCACATGAGTAAGAGCTTGCCACCAGTGGGGACAGAGTTTCCAGAATTAGGA AAATAATCACTGGGGGCAAGTTTGAGGTTGGTACCAGATATGTGGGAGGAGGCAAGGTAAGGGAAAGAGTA CTTGAAGTTGGAACTGGTCCTTGCAGGGAAATGCACATTTATGAAACCCCGAAAACTGATGTCAAAGCACCTC CTGCCTTGGGCAGAGTCCTCTCAGAGTCTACAGGTGCTGCCTCCAGAACCCTCTTCCTGGAGCGCATCCCTATG TATCTAGAAATTCTGCTGGGAAATATGATGGTCAGACCCTTGGCCACCTGAAAGGTTCAGGGTGGTAGAAGA AAAAGGAAAGCCACAGGGCAGCAGGGGCAGGTGCCAGCAAGGAAGGCAGGCACGCCAGGAAGACACCCAT GGTGAGAAGTGCAGATGGCCCGAGGGCAAGTTTGCTCAACTCACCCAGGTTTGCTCTTGCTGGGGCCAAGAG GACTCATGTGCCAGGGCCAAGGGCCCTTGGGGGCTCTCACAGGGGGCTTATCTGGGCTTCGGTTCTGGAGGG CCAGGAACAAACAGGCTTCAAAGCCAAGGGCTTGGCTGGCACACAGGGGGCTTGGTCCTTCACCTCTGTCCCC TCTCCCTACGGACACATATAAGACCCTGGTCACACCTGGGAGAGGAGGAGAGGAGAGCATAGCACCTGCAG SEQ ID NO: 11 CpG‐Free CMV Enhancer Forward (302bp) GTTACATAACTTATGGTAAATGGCCTGCCTGGCTGACTGCCCAATGACCCCTGCCCAATGATGTCAATAATGAT GTATGTTCCCATGTAATGCCAATAGGGACTTTCCATTGATGTCAATGGGTGGAGTATTTATGGTAACTGCCCAC TTGGCAGTACATCAAGTGTATCATATGCCAAGTATGCCCCCTATTGATGTCAATGATGGTAAATGGCCTGCCTG GCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTATGTATTAGTCATTGCTATTAC CATGG SEQ ID NO: 12 CpG‐Free CMV Enhancer Reverse (302bp) CCATGGTAATAGCAATGACTAATACATAGATGTACTGCCAAGTAGGAAAGTCCCATAAGGTCATGTACTGGGC ATAATGCCAGGCAGGCCATTTACCATCATTGACATCAATAGGGGGCATACTTGGCATATGATACACTTGATGT ACTGCCAAGTGGGCAGTTACCATAAATACTCCACCCATTGACATCAATGGAAAGTCCCTATTGGCATTACATG GGAACATACATCATTATTGACATCATTGGGCAGGGGTCATTGGGCAGTCAGCCAGGCAGGCCATTTACCATAA GTTATGTAAC SEQ ID NO: 13 ELF3‐1 enhancer forward (233bp) CAGGGGGCCTGGGGCTAGGGGACAGCGGGGCCTTTTTCTTACCTAAGAGGCTACAAAGGGGAGGGTGAGGT GCCTGAAAAGGTCGGCTCTGCTGCCACCGTGTGGCTAATACTAAGAACAGCAGCAAGCCTAGGCTGTTGGTCT
GGTGGTTGGGGAATTGGGAGAGCAGAAAGTAATATAGAGTCACCCAACTCACATTTTCTCAATCAGCTCCTTC TCATCCAATGCTCCCG SEQ ID NO: 14 ELF3‐1 enhancer reverse (233bp) CGGGAGCATTGGATGAGAAGGAGCTGATTGAGAAAATGTGAGTTGGGTGACTCTATATTACTTTCTGCTCTCC CAATTCCCCAACCACCAGACCAACAGCCTAGGCTTGCTGCTGTTCTTAGTATTAGCCACACGGTGGCAGCAGA GCCGACCTTTTCAGGCACCTCACCCTCCCCTTTGTAGCCTCTTAGGTAAGAAAAAGGCCCCGCTGTCCCCTAGC CCCAGGCCCCCTG SEQ ID NO: 15 ELF3‐2 enhancer forward (654bp) TAGGGATGGGCCGAGGCTGGCACTGATGCTAGACTTCCGTGCACAGGGCAAGTATGGACAAGCCCCAAGTG GCTTTGTGAGGCCCACACAGTGAAGCTTGGGAAATGGGAAGTGGGGCTGCGCCCAGATTCTGGTATCTATGA CAACTAAGGCCGCTGCACATCCTCATGGCTCTCCCAGAGACCTCAGGTGAGGCCCTTCTGTGTTCCTCAAGCAC CCATGCCACCTGCGGGGTGGGGCAGGACCCTCCTACCCAGCCCTGGGCCTCCTGGGGAACCATGGGTGCACA GGGGTAACCTGAGCCAGCCTCTCTGGGGCATGGGCGGGCGTGGTGGCGTGGGCCTGGCGCCAGAGAGTGG AGCAGAATGTCAGCTCTGTGAGCCGCACCGGGTGCCAGCACTCTGCAAACAGACTCTAGTCACCAGATAGACT GGAGTCACGAACCTAACAAAGCGCTCAGCTGGGCAACTGTAACTGCAGAGGGCGGGGCCGCACAGTGCTGC CTAGTGCCTCCTGCCTTGATCTGGTCAGGGTCATGGGAGGAGACAGGTTCCCCAGGAGGGCCCATTAACAAG ATTAATTGGGGTGGCCTGGGGGACTCTGGGATGCTCACTGGAAACATATCCTGGGGGAATGAGGTAGGTGG GGAGCCAG SEQ ID NO: 16 ELF3‐2 enhancer reverse (654bp) CTGGCTCCCCACCTACCTCATTCCCCCAGGATATGTTTCCAGTGAGCATCCCAGAGTCCCCCAGGCCACCCCAA TTAATCTTGTTAATGGGCCCTCCTGGGGAACCTGTCTCCTCCCATGACCCTGACCAGATCAAGGCAGGAGGCA CTAGGCAGCACTGTGCGGCCCCGCCCTCTGCAGTTACAGTTGCCCAGCTGAGCGCTTTGTTAGGTTCGTGACT CCAGTCTATCTGGTGACTAGAGTCTGTTTGCAGAGTGCTGGCACCCGGTGCGGCTCACAGAGCTGACATTCTG CTCCACTCTCTGGCGCCAGGCCCACGCCACCACGCCCGCCCATGCCCCAGAGAGGCTGGCTCAGGTTACCCCT GTGCACCCATGGTTCCCCAGGAGGCCCAGGGCTGGGTAGGAGGGTCCTGCCCCACCCCGCAGGTGGCATGG GTGCTTGAGGAACACAGAAGGGCCTCACCTGAGGTCTCTGGGAGAGCCATGAGGATGTGCAGCGGCCTTAGT TGTCATAGATACCAGAATCTGGGCGCAGCCCCACTTCCCATTTCCCAAGCTTCACTGTGTGGGCCTCACAAAGC CACTTGGGGCTTGTCCATACTTGCCCTGTGCACGGAAGTCTAGCATCAGTGCCAGCCTCGGCCCATCCCTA SEQ ID NO: 17 ELF3‐3 enhancer forward (681bp)
GGCAGGGCTTCTGGGGATTGGTGACACCCAGCTGACGTCAGGGAGGTGGAGAAGCCGCAGGTCCCTCTTATC CCCAGGGCAGTTCAGGGGCTTGACTGCATTTTAGCGATGATTGTAGTTACAGATTGCTCTCCGAAACACTGGG CAGGAAAAGCTTGCTGGTGTTCTGCTTTTGGCTCTGGGATTTGAACCCATGACTGGCTCAGGAGACTTGAGAT TACCAACGTGCTTTGTGTCCCCAACAAGTCACTTGTCCTCTCTGGGCCATCTCTAGAGCACAGCCTCCTAATTCA ATAAAATGAAGAGGCTGGAGGAGAGATTGCCAAGGACTCTTTCAAGAGGCCCAGATAGGAGAGCAGGAGGA TCGGGGGGTGGGGGGGTGTTCTGAGTTGGCCCTGTCATGGCCCTAATCTGTCCTTCCCCCATCCTTCACTCCCC CTCTTATCCCTGGTGGCCCCGGAGTGAGAACGTGACTCATCCAGCTCCAGGCACCCAGTTTCAGCCTTGCCCCA CCCCTGCCCCGGGCCTCTCATTTGCTGTTTACCCCCCAGGAGCAGGTGCACCTGGGTCACCTCACCGCAGGAA GGAGAGATATAAGGCTCTAGGGCACAGCCTGACTCCACACCCACAGATTGCCCAAGGCCAGCACAGGGGTTG GAGCTCTCTAGAAAGGTGAGGCAC SEQ ID NO: 18 ELF3‐3 enhancer reverse (681bp) GTGCCTCACCTTTCTAGAGAGCTCCAACCCCTGTGCTGGCCTTGGGCAATCTGTGGGTGTGGAGTCAGGCTGT GCCCTAGAGCCTTATATCTCTCCTTCCTGCGGTGAGGTGACCCAGGTGCACCTGCTCCTGGGGGGTAAACAGC AAATGAGAGGCCCGGGGCAGGGGTGGGGCAAGGCTGAAACTGGGTGCCTGGAGCTGGATGAGTCACGTTCT CACTCCGGGGCCACCAGGGATAAGAGGGGGAGTGAAGGATGGGGGAAGGACAGATTAGGGCCATGACAGG GCCAACTCAGAACACCCCCCCACCCCCCGATCCTCCTGCTCTCCTATCTGGGCCTCTTGAAAGAGTCCTTGGCA ATCTCTCCTCCAGCCTCTTCATTTTATTGAATTAGGAGGCTGTGCTCTAGAGATGGCCCAGAGAGGACAAGTG ACTTGTTGGGGACACAAAGCACGTTGGTAATCTCAAGTCTCCTGAGCCAGTCATGGGTTCAAATCCCAGAGCC AAAAGCAGAACACCAGCAAGCTTTTCCTGCCCAGTGTTTCGGAGAGCAATCTGTAACTACAATCATCGCTAAA ATGCAGTCAAGCCCCTGAACTGCCCTGGGGATAAGAGGGACCTGCGGCTTCTCCACCTCCCTGACGTCAGCTG GGTGTCACCAATCCCCAGAAGCCCTGCC SEQ ID NO: 19 ELF3‐4 enhancer forward (571bp) GTGGCTCAGGTGATAAGAGGAGACTCAGGACCAGTCCCTTAACAGGTTGGTGGAACTGTGAGTAGAAATTTT TTTTTGCTCCGCCCCTACCCAAAGTGGCTTCATTGAGATTCAGCCTATCTCCCACCCTTCCTGTGCTGTGATGAG GGACTCTGGGGACTGGGAGTGAATAGGCCCTGCTGAGTGCAGGGAAGACACACTCTGGCCTTCCTCACGCCA GGCTGGAGACCTGGCACTGGGCTCTGCTCCTAGCAGCCTGGAAAGCTGTTATTCTAGAACCTGATAACTGCTT GGGCTGGGCTGAAGATAAGAGGCAGCTGTTTGAAGTCCAGAGAGATAACAGCTTCAGAGAAATTTGATAAG AGGAGTGAGATGCACTGGGTTAAGCTCACCAGTCTCGCCTGTTTGCCTTTTGCCAGTTTATAATCATAACCAGC ACTGTACTCTCAAGTCAGACTTCCCAGCACCGTGAGGAACCTCAGTGTCACTGTCTCTTAATGGTAGTCTGTGC ATCTCTCCTGAGGTCCGCGCCTGGGCCTGCAGGCCTCTGTGCTTGTTCTATGAAATGACC SEQ ID NO: 20 ELF3‐4 enhancer reverse (571bp)
GGTCATTTCATAGAACAAGCACAGAGGCCTGCAGGCCCAGGCGCGGACCTCAGGAGAGATGCACAGACTACC ATTAAGAGACAGTGACACTGAGGTTCCTCACGGTGCTGGGAAGTCTGACTTGAGAGTACAGTGCTGGTTATG ATTATAAACTGGCAAAAGGCAAACAGGCGAGACTGGTGAGCTTAACCCAGTGCATCTCACTCCTCTTATCAAA TTTCTCTGAAGCTGTTATCTCTCTGGACTTCAAACAGCTGCCTCTTATCTTCAGCCCAGCCCAAGCAGTTATCAG GTTCTAGAATAACAGCTTTCCAGGCTGCTAGGAGCAGAGCCCAGTGCCAGGTCTCCAGCCTGGCGTGAGGAA GGCCAGAGTGTGTCTTCCCTGCACTCAGCAGGGCCTATTCACTCCCAGTCCCCAGAGTCCCTCATCACAGCACA GGAAGGGTGGGAGATAGGCTGAATCTCAATGAAGCCACTTTGGGTAGGGGCGGAGCAAAAAAAAATTTCTA CTCACAGTTCCACCAACCTGTTAAGGGACTGGTCCTGAGTCTCCTCTTATCACCTGAGCCAC SEQ ID NO: 21 actin enhancer forward (37bp) CGCCTCCGACCAGTGTTTGCCTTTTATGGTAATAACG SEQ ID NO: 22 actin enhancer reverse (37bp) CGTTATTACCATAAAAGGCAAACACTGGTCGGAGGCG SEQ ID NO: 23 LMO7‐1 enhancer forward (864bp) ATTTGTTATAAATTCACTAATATCTTTTAAAGTGGGAGGAATGAGAGAATGACTTTAACCCATCTCTCCAAGCC ACCCAGCCTGAGGCCTTTGGTTTGTTGAATAAACGAAAACGTGGCCACTTCAGAGAAGAGCAAGGCTTCCAG CACCTTCCCACAGCCTGAATTCCATTAGCATGAGACTTTGAAACCACATCTGTTTTTCCTTATGAAATTCAAACA ATTGTCTAACCTCTTGGCTCTTAAGTCAAATGTTCAAGGGCTCATATGACAACACTTTGTTGTATATAAAAATTT GTTAAATTTAAATTGCTTTTGCAAAAAAGAGGAAAAGGGGGAATAAATAAATAGCTTTAAGGCATCTGGTTAG GATCTGCACAAGGTTGCATTCTTTCCATGTCTCCAAAAGGTTTGTTCTATCTGTCACCCATCACTTCACTGGAGT TCCTGCAGGCAGCAAAATCATGCAGAAGTTCTTTGTATAGAAATACAGTGTCTGGAGCCTAGTCTGACTTCCT GTTTGGCTGGAGCTGAGCTAGTCCATGGATAGGCAGAAAGCAGCTGTAGGAATCTGTGTTTAGAGTAGGGCT GTTTAGTAGACACAAGGGCAACCCACAGGGACCTGTGACACATCTCTATTGAAATATGGCATGATGTGCTGAT TTGTTTGTCCAAGATTTAATTATAGATGTTTAGCAGGGTTGCATAAACAACAGAGGGATAGGAGAGATAAGG AGGGAAGATTCAGGGAAAAAAATCAGTGGCTTATATTTTTGATCAGATATATTATGTGCCAGAATGAATCTCT ATCATGCTTTGAGATTTGCATGTGTGTGTGTGTGTGTGTGTAGACACATCAAAGGTC SEQ ID NO: 24 LMO7‐1 enhancer reverse (864bp) GACCTTTGATGTGTCTACACACACACACACACACACATGCAAATCTCAAAGCATGATAGAGATTCATTCTGGCA CATAATATATCTGATCAAAAATATAAGCCACTGATTTTTTTCCCTGAATCTTCCCTCCTTATCTCTCCTATCCCTCT GTTGTTTATGCAACCCTGCTAAACATCTATAATTAAATCTTGGACAAACAAATCAGCACATCATGCCATATTTCA ATAGAGATGTGTCACAGGTCCCTGTGGGTTGCCCTTGTGTCTACTAAACAGCCCTACTCTAAACACAGATTCCT
ACAGCTGCTTTCTGCCTATCCATGGACTAGCTCAGCTCCAGCCAAACAGGAAGTCAGACTAGGCTCCAGACAC TGTATTTCTATACAAAGAACTTCTGCATGATTTTGCTGCCTGCAGGAACTCCAGTGAAGTGATGGGTGACAGA TAGAACAAACCTTTTGGAGACATGGAAAGAATGCAACCTTGTGCAGATCCTAACCAGATGCCTTAAAGCTATT TATTTATTCCCCCTTTTCCTCTTTTTTGCAAAAGCAATTTAAATTTAACAAATTTTTATATACAACAAAGTGTTGT CATATGAGCCCTTGAACATTTGACTTAAGAGCCAAGAGGTTAGACAATTGTTTGAATTTCATAAGGAAAAACA GATGTGGTTTCAAAGTCTCATGCTAATGGAATTCAGGCTGTGGGAAGGTGCTGGAAGCCTTGCTCTTCTCTGA AGTGGCCACGTTTTCGTTTATTCAACAAACCAAAGGCCTCAGGCTGGGTGGCTTGGAGAGATGGGTTAAAGTC ATTCTCTCATTCCTCCCACTTTAAAAGATATTAGTGAATTTATAACAAAT SEQ ID NO: 25 LMO7‐2 enhancer reverse (675bp) AAACACGTATTTGTTAAGGCCTTAGAGTTCAACACTCAAGATGGATTTTTGACCTCATAAATGTTAAAGTTATT GATGACAGCCACTTCAGTTATAAAATAACCTTTGGCCCTGGTGCCTGGCTTTTAAGGAGGCACGCTGACAAGT CAATTAAAATGAGTCACCTCCAACTTCCAAGATGACTGACTGAGGTGCTACAACTTCCACCAACGATTGTTTAA TTGGTATTTTATGTTAGCTTTTAAGCATTATCTGGTTAAATGGGAAATGCCTGATGAACACATTGGTTTTGATTT TAATAGAACTGATACAAAACATGTGATATAGTCACATACCTCTAACAGCTACCCCCAGTTTATTATAATGGAAT GGAACCACAACCTTAGTTTTATACACCATAAGGTCTTCAACTACCTCCTCTGGTTTTGAAACTGGTAACAGGAA ACAGCCTTTGATAAAGCATTCCTGGCATAGACACTGTACTAGGATTATTCAACCTGAGTCAGACTGTCATTAAA AGGACCAAAGGCTACAAAGACAGCCCTAGTACATAAGCCAAAGTCCCACATGGTTTTTTGTTTGTTTGTTTGTT TTTGAGACGGAGTTTTGCTCTTGTTGCCCAGGCTTGGCTCACCACAACCTCTGCCTCCCAGGTTCAAGCGACTC TCCTGCCTC SEQ ID NO: 26 SFTPB enhancer forward (697bp) GATGCTGATGTGACTGATTTGTAGGTGGCCAGAAGGAAAGGCGGGAAGAAAGATGTTGGTGGCAGCCAAGA CCGAGGAAGTCTGTGAGGCAGGAGGGGTGAGGGGTCCCTTGCCCTGTGACATAACCTTGTCGTGTTGGAGTT TCAGGCCCAGGTTACTTACTGAAGCATTGTCTTTTTATGTTGGACTTCTGTCTGGTCCCCCTGAGACGTTGTCTC TTCCGCAGGCCCTCGGGGGCTGGAGCACAGCTGTAGCCAACAGAACACAGGCTCTGTCTGCTGCCCTGAAGC ACAAAGGAAGTTGGTCTCTTGAGGTTTTTCTAGGAATGTTTTCCTGTGAGCAATCACAGGAGAAGCGGGAGTA AAACAAACAAAGGAGGGAAAAAGATGACACAAAAACATCATGGGAAGGCTGGCGGTTGGCAGAGGAGGCT CAGATTGAGGAAACAATCCCTCATGGAATGAATTCTTCTGTTAGTGGAGTGATGCATATTGACAGATCCAGGC ACATCTTCAGGCTGCCCTTCTGTCTGTCCCTCACCCATATATTCATTCTACAAATATTTGCTGAGAGCCTATTAT GTGTGAGGCATCATGCTAAGTGCATAACATAAATTACTTCATTTTATCTTCACAGTGTCCCTTTAAGGTAGGGT CTCTTATCCCCATATCTTACCCAAGGAACCCGAGGCTTAGAG SEQ ID NO: 27 SFTPC enhancer forward (590bp)
CAAGTTCACCTCTCAGCTTCATTTTTGCTCATCTCTAAATGAGGAGAATGTCGTTTTGTGTGAGGTTTAGAGAC GATGTATGAAAGTCCCAAGTATACAGGGGCTGCCACCAAGCAGGTAGTGGTTGGCTGGAAACAGGAGCACA GCTAGACCAGGGGCCCTCCACCCCGAAGTTGCCATTCTGCCTGCGGAGGCTTCATTTCTAAAAGTCCAGGGGA GCCACAGTCTGGATCAGTTTCTGCCTGGAAGAAGAGCGCGCCAGGATGAGTGGGCACAAACTCGGAGGGCC CAGTGGGCGGGTCTTATCACCCACATCCTGGGGAGCTGTGGAGAGGAGAGGGTGGTGGGTGAGGTGGGGC TGGGCTGGTGGCTCAGATAAGGCAGGGACACACAGCTGGGGGAGGGTGGGGCTGAGATGGAGGGCGAGC GGCTGGCTGGCGTCCGACCAGGGCCAGGGGCAGCTGAGCTGTCCCTTCCGTCACCTGGACCTGCCTTCCTCTG TGGTGCTCTCTCTTCCTCCCTCCCTCCCCTCTTCCTGCTGCAGCTCGCCCTTTCTATCTCTTTGTGCACGGAGGTC GCTGGGACCTTGG SEQ ID NO: 28 SFTPC enhancer reverse (590bp) CCAAGGTCCCAGCGACCTCCGTGCACAAAGAGATAGAAAGGGCGAGCTGCAGCAGGAAGAGGGGAGGGAG GGAGGAAGAGAGAGCACCACAGAGGAAGGCAGGTCCAGGTGACGGAAGGGACAGCTCAGCTGCCCCTGGC CCTGGTCGGACGCCAGCCAGCCGCTCGCCCTCCATCTCAGCCCCACCCTCCCCCAGCTGTGTGTCCCTGCCTTA TCTGAGCCACCAGCCCAGCCCCACCTCACCCACCACCCTCTCCTCTCCACAGCTCCCCAGGATGTGGGTGATAA GACCCGCCCACTGGGCCCTCCGAGTTTGTGCCCACTCATCCTGGCGCGCTCTTCTTCCAGGCAGAAACTGATCC AGACTGTGGCTCCCCTGGACTTTTAGAAATGAAGCCTCCGCAGGCAGAATGGCAACTTCGGGGTGGAGGGCC CCTGGTCTAGCTGTGCTCCTGTTTCCAGCCAACCACTACCTGCTTGGTGGCAGCCCCTGTATACTTGGGACTTT CATACATCGTCTCTAAACCTCACACAAAACGACATTCTCCTCATTTAGAGATGAGCAAAAATGAAGCTGAGAG GTGAACTTG SEQ ID NO: 29 SLC34A2‐1 enhancer forward (162bp) GACTTTCCCATCAGTCTGAACTCCTGGGTTTTCCCAGGCCTGGTGACTCAGAGGGTAAGGCACCGAGCCGAGG AAGAGAAAGGAGAACCTGGACTCCTGGGTGCCTTTGTCTCACCACACAATCACCTCCCAGGAAGGCTGTCATC CAATCACTCCTGTGGA SEQ ID NO: 30 SLC34A2‐1 enhancer reverse (162bp) TCCACAGGAGTGATTGGATGACAGCCTTCCTGGGAGGTGATTGTGTGGTGAGACAAAGGCACCCAGGAGTCC AGGTTCTCCTTTCTCTTCCTCGGCTCGGTGCCTTACCCTCTGAGTCACCAGGCCTGGGAAAACCCAGGAGTTCA GACTGATGGGAAAGTC SEQ ID NO: 31 SLC34A2‐2 Enhancer Forward (403bp) AGCTAACCATTCAGACTGTCAGTTTAGATTATTGAAAGGAACAGAAGAGAAATCTTGCCTATAAAACACGGAA CTGAGAAAAAAGTGTATACCTCACTTAGAGCGTCATTAAGTTAGAGAACAACTCCCACACTCTGCTCTAAGAT
GAGGCCAACATCCAGATTCTCACTGCAAACTTGGCTCCCCTAGAATTCTGTTCTTCTTTCCCCTGTCCCCGCTGC TTTCTAAACGTGAAATATCCACAGCTGCACCGTTTCTTTACTTTTTTTTTTTTTTCTGGAGAGGTTAACCTTCCTT CTTCAGTTGTATGTTGTGCAATACCCAAGAGGCCAACACTACCATTTGAGATATTTAAATATGACTTTGAGAAA TGAGTCTGTCTTAGATATAAAAGCCCACAGTC SEQ ID NO: 32 SLC34A2‐2 enhancer reverse (403bp) GACTGTGGGCTTTTATATCTAAGACAGACTCATTTCTCAAAGTCATATTTAAATATCTCAAATGGTAGTGTTGG CCTCTTGGGTATTGCACAACATACAACTGAAGAAGGAAGGTTAACCTCTCCAGAAAAAAAAAAAAAAGTAAA GAAACGGTGCAGCTGTGGATATTTCACGTTTAGAAAGCAGCGGGGACAGGGGAAAGAAGAACAGAATTCTA GGGGAGCCAAGTTTGCAGTGAGAATCTGGATGTTGGCCTCATCTTAGAGCAGAGTGTGGGAGTTGTTCTCTA ACTTAATGACGCTCTAAGTGAGGTATACACTTTTTTCTCAGTTCCGTGTTTTATAGGCAAGATTTCTCTTCTGTT CCTTTCAATAATCTAAACTGACAGTCTGAATGGTTAGCT SEQ ID NO: 33 SLC34A2‐3 enhancer reverse (705bp) TCAGGAGGCTGAGGTGGGAGGATCACTTGAGCCCAGGAGTTCGAGACTGCAGTGAGCTGTAATCACACTACT GCATTCCAGCTAGGGTGACAGTCTCATAAGGGAAAAAAAAAAAAGATTTAAAAAACCTCTCTCCAGCATTGAA ACTTCCTGTAGGCTAGTTGACCAAATCTCAACATTAGCTGCCAGGTTATGTGTGTCCAAGCAATGCAGGTTGTG GTTTTAATTAGCATGATGTCTCCAAGGATTGTGTCTGTCCTCCTGTGTATGGCGATTCCAATTTACTGGTTTTAA GATGTGGTGAACAAAAGTCTGTTCTTGTAATTGACCAAGGGGAAGTAGGCCACGAAGGTGTGAATCTTAAGT ACGTGACCCCCTAACTCACGCTGAAGCCATGCTCACTGCATCTGAGCAGGTGCCGCTTTCGCTTCTTGTTTTTTA TTTTTAAATTTCCTAAGTTTTGTTTGGTTTTTACCTTCCCCTGATTTGGCAGATGCGTTTGTACTTTTTGGAGAGA TATTTATACCATTCGTGTTAAAGGTATTTTTTTTTTAAATGGTGTCCTGTCAACAGTACATTGAGGATACAATTT AGGAACTCAAAATGTGTTTAAACACAACTTAGCAAATGTTATGTCTCATCATTTGGCATTTGGCCTCTGCAGAT TCATTTCATTTAATGCCATAACCAATATTCCTGATTTGA SEQ ID NO: 34 SV40 forward enhancer (237bp) CGATGGAGCGGAGAATGGGCGGAACTGGGCGGAGTTAGGGGCGGGATGGGCGGAGTTAGGGGCGGGACT ATGGTTGCTGACTAATTGAGATGCATGCTTTGCATACTTCTGCCTGCTGGGGAGCCTGGGGACTTTCCACACCT GGTTGCTGACTAATTGAGATGCATGCTTTGCATACTTCTGCCTGCTGGGGAGCCTGGGGACTTTCCACACCCTA ACTGACACACATTCCACAGC Also referred to interchangeably as SV40 Forward in the Examples/Figures herein. SEQ ID NO: 35 SV40 reverse enhancer (237bp)
GCTGTGGAATGTGTGTCAGTTAGGGTGTGGAAAGTCCCCAGGCTCCCCAGCAGGCAGAAGTATGCAAAGCAT GCATCTCAATTAGTCAGCAACCAGGTGTGGAAAGTCCCCAGGCTCCCCAGCAGGCAGAAGTATGCAAAGCAT GCATCTCAATTAGTCAGCAACCATAGTCCCGCCCCTAACTCCGCCCATCCCGCCCCTAACTCCGCCCAGTTCCGC CCATTCTCCGCTCCATCG SEQ ID NO: 36 VEGFA‐1 enhancer forward (650bp) AATTGGGTATGTAGATCATCAGGGTGGAAGGTGAGCTTCCTTTCTAGGTGGGGAAACAGGCTCAGAGAGGG GCCATGACTTACTGGTGTTTCACAATGAGCCAGTGGTGGAGCTCACAGATCCTGGCCCCAGACCAGGCACCCT CCTCCCCCTAACCCCCATCTCCAGCTCTCAATGCCAAGAGGCTGGAGCGGGCCCTGTGCACTCCAGGCAAGGC TGCGGTTTCTTCTGTGGGACTGTGGAGCCACCCAGGGTGTGTGGGTGCTGCGCACGCGCACCACAGTTGTGT AACAGGCTCTGGGGCGTGTGCACGCGCTCAGTAACCTGATGCAACAGGGAAGCTGTGGTCCCGTGAGCTGG GGTGTGGGCGCTTGGCCGGTGCCCAAACCACAAAGAGTATTTGGCTTGTGGCTAAGGAGGCCAGGCATGGG GGTCGCAGAGTTGTCTACGTCAGGGCTTGTCTATCCCTAGTTGCCCACATCCTGGTCCCTGTGGGCGCAGGGC TGGGAATGGGGCTGCCTGGCCAGGTGATACTGTGGAGGAGGGGTTGGGGCCTTGCCCCTTCCTCCTTCTGTTT GCCACGATGCCTATTGTGATGAGGATGGGGGGCAGGCGGGGCTTCCTCCCTGTGGTTCCATCTCTTCCTGGCT SEQ ID NO: 37 VEGFA‐1 enhancer reverse (650bp) AGCCAGGAAGAGATGGAACCACAGGGAGGAAGCCCCGCCTGCCCCCCATCCTCATCACAATAGGCATCGTGG CAAACAGAAGGAGGAAGGGGCAAGGCCCCAACCCCTCCTCCACAGTATCACCTGGCCAGGCAGCCCCATTCC CAGCCCTGCGCCCACAGGGACCAGGATGTGGGCAACTAGGGATAGACAAGCCCTGACGTAGACAACTCTGCG ACCCCCATGCCTGGCCTCCTTAGCCACAAGCCAAATACTCTTTGTGGTTTGGGCACCGGCCAAGCGCCCACACC CCAGCTCACGGGACCACAGCTTCCCTGTTGCATCAGGTTACTGAGCGCGTGCACACGCCCCAGAGCCTGTTAC ACAACTGTGGTGCGCGTGCGCAGCACCCACACACCCTGGGTGGCTCCACAGTCCCACAGAAGAAACCGCAGC CTTGCCTGGAGTGCACAGGGCCCGCTCCAGCCTCTTGGCATTGAGAGCTGGAGATGGGGGTTAGGGGGAGG AGGGTGCCTGGTCTGGGGCCAGGATCTGTGAGCTCCACCACTGGCTCATTGTGAAACACCAGTAAGTCATGG CCCCTCTCTGAGCCTGTTTCCCCACCTAGAAAGGAAGCTCACCTTCCACCCTGATGATCTACATACCCAATT SEQ ID NO: 38 VEGFA‐2 forward enhancer (680bp) GAGGCACAAAGCGATCCCCATCACTGCTCCACAATCATTCATTAGCTAACAAGACAGAGCAGCTCATAAAAAA AAAGCCGTTAAAAAAATTCCGGGGAAATGGAAAGCAGGAGGTGATGCAAGCCCTGGTTAACAAAGGCTGAG GGTTGGGGGGAGGCATGAGAGGGTGTGAGTGGAATAACCCAAGCCTGATAAGCCACAAAGCAGCGCCTGCT GCTCCCTCCTCCCTCTGCCGCCTGAGTCAGAGAAGCCGGGATGTGTTCAAAATCAAGCAATGTAATTAGCAGG TTGCATCATGCCTCTCGATTATAAATTACAATGGCCACAAAGAGGGTGGATGGAGGAGCAGGGATTAGATGG GGTAACCCAATCCCAGTGCCGGGAAGGGGGACAGAGGGCCCGGGAGCTGCCCCCCGCCACTAGGAGCCACT
CGGAGTGGCCTCTGTCTACTGTGTCTAGGGACAGCAACAAGGCCTACTCAATACCTGCTTTGTGTAGCCTGGG ATAATGATTTACAAACTCAAGGAAATCAGGAGGTGTGAAAAGAAGATAGCCCTCTTAACCCAGCCATTGTCTA CCCTCTGCTGGAATGTCCCTCAGCTCCACCCACCTGGTGCACCTGTGCACACGGATGCTGTTTGAGGATGGTCC ACGTGAGGAGCCCTCATAGGGGCTGGACT SEQ ID NO: 39 VEGFA‐3 enhancer forward (801bp) GGAGGCTTGGCCTGCAGTTTGAGGATCAAACTGTCTCCTGGAAACTACCCCCATGCCCTGGACCCAGCCTCTG TGGCCCAGGACAGTAGTGGGGACAGGTTAAGGCTAGGGAGTGCCTTTGCCAGGCACTGGGAAAAGACACTG CAGGCCCCAAATTGTGGTTAATCTCCAAGGACACGCAAGAATGCAGCAGGAGAGACTAGGATTAGACAACAG GGAGAACTCGATGACAGCTAGAGCTCCGGGATTCATTCCCTGCAAGGGCAGGGAGGTTGAGGGGGCAGAGG AGGAAGGAATCTGCAAGGAGGTTGCTGACGTCCTCCTGGAAGGCTAGAAATGAAGCACAGGCAGCCGAGTC TTCTCCCACCTCCCTCCACCCCCAGCGATGAGTCAGGGATTCCGCAGAATGCTGAATTCCTGTCTTCCTTTGGG CAATGATACTTGACACTCTCCTAGGTTTTCTCTCCTGCCACAGAAGTAAAATCTTCAACCATTCTCAAGGACTTC TTTCTCCTCAGCTCCCCAGACAGGCCCCTGCTTTCCTGTGAGTCACTGTATCTTCAAAGCCCCTTCCTGCCTCAG CCTACCCCAATAGCTCTCAGGCTCCTGGCTTTGCACGTGTTGTTCTGTCTGCCTGGAACACTCTTCCCAGAACTA GCTTCTTCTCATCCACCAGAAACCAGTTTAGATGGCACCTGCTGAAACAGATCAGCTCACAAAGGCCCGGCAC AACCTGGCCCACAGAAGCATTCCTGAGTGCTGTAAATGAATGAACTGTGAATCAATGAATGCAGAAGATTC SEQ ID NO: 40 VEGFA‐3 enhancer reverse (801bp) GAATCTTCTGCATTCATTGATTCACAGTTCATTCATTTACAGCACTCAGGAATGCTTCTGTGGGCCAGGTTGTG CCGGGCCTTTGTGAGCTGATCTGTTTCAGCAGGTGCCATCTAAACTGGTTTCTGGTGGATGAGAAGAAGCTAG TTCTGGGAAGAGTGTTCCAGGCAGACAGAACAACACGTGCAAAGCCAGGAGCCTGAGAGCTATTGGGGTAG GCTGAGGCAGGAAGGGGCTTTGAAGATACAGTGACTCACAGGAAAGCAGGGGCCTGTCTGGGGAGCTGAG GAGAAAGAAGTCCTTGAGAATGGTTGAAGATTTTACTTCTGTGGCAGGAGAGAAAACCTAGGAGAGTGTCAA GTATCATTGCCCAAAGGAAGACAGGAATTCAGCATTCTGCGGAATCCCTGACTCATCGCTGGGGGTGGAGGG AGGTGGGAGAAGACTCGGCTGCCTGTGCTTCATTTCTAGCCTTCCAGGAGGACGTCAGCAACCTCCTTGCAGA TTCCTTCCTCCTCTGCCCCCTCAACCTCCCTGCCCTTGCAGGGAATGAATCCCGGAGCTCTAGCTGTCATCGAGT TCTCCCTGTTGTCTAATCCTAGTCTCTCCTGCTGCATTCTTGCGTGTCCTTGGAGATTAACCACAATTTGGGGCC TGCAGTGTCTTTTCCCAGTGCCTGGCAAAGGCACTCCCTAGCCTTAACCTGTCCCCACTACTGTCCTGGGCCAC AGAGGCTGGGTCCAGGGCATGGGGGTAGTTTCCAGGAGACAGTTTGATCCTCAAACTGCAGGCCAAGCCTCC SEQ ID NO: 41 SLC34A2‐2 Enhancer & core SFTPB Promoter (1039bp)* agctaaccat tcagactgtc agtttagatt attgaaagga acagaagaga aatcttgcct 60 ataaaacacg gaactgagaa aaaagtgtat acctcactta gagcgtcatt aagttagaga 120 acaactccca cactctgctc taagatgagg ccaacatcca gattctcact gcaaacttgg 180
ctcccctaga attctgttct tctttcccct gtccccgctg ctttctaaac gtgaaatatc 240 cacagctgca ccgtttcttt actttttttt ttttttctgg agaggttaac cttccttctt 300 cagttgtatg ttgtgcaata cccaagaggc caacactacc atttgagata tttaaatatg 360 actttgagaa atgagtctgt cttagatata aaagcccaca gtcagatctt atagggctgt 420 ctgggagcca ctccagggcc acagaaatct tgtctctgac tcagggtatt ttgttttctg 480 ttttgtgtaa atgctcttct gactaatgca aaccatgtgt ccatagaacc agaagatttt 540 tccaggggaa aaggtaagga ggtggtgaga gtgtcctggg tctgcccttc cagggcttgc 600 cctgggttaa gagccaggca ggaagctctc aagagcattg ctcaagagta gagggggcct 660 gggaggccca gggaggggat gggaggggaa cacccaggct gcccccaacc agatgccctc 720 caccctcctc aacctccctc ccacggcctg gagaggtggg accaggtatg gaggcttgag 780 agcccctggt tggaggaagc cacaagtcca ggaacatggg agtctgggca gggggcaaag 840 gaggcaggaa caggccatca gccaggacag gtggtaaggc aggcaggagt gttcctgctg 900 ggaaaaggtg ggatcaagca cctggagggc tcttcagagc aaagacaaac actgaggtcg 960 ctgccactcc tacagagccc ccacgccccg cccagctata aggggccatg cmccaagcag 1020 ggtacccagg ctgcagagg 1039 *wherein M is C or A, preferably M is C Bases in bold and underlined represent a BglII restriction enzyme site created to operably link the enhancer and promoter sequences. For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 76). SEQ ID NO: 42 VEGFA‐2 Enhancer & core SFTPB Promoter (1316bp)* gaggcacaaa gcgatcccca tcactgctcc acaatcattc attagctaac aagacagagc 60 agctcataaa aaaaaagccg ttaaaaaaat tccggggaaa tggaaagcag gaggtgatgc 120 aagccctggt taacaaaggc tgagggttgg ggggaggcat gagagggtgt gagtggaata 180 acccaagcct gataagccac aaagcagcgc ctgctgctcc ctcctccctc tgccgcctga 240 gtcagagaag ccgggatgtg ttcaaaatca agcaatgtaa ttagcaggtt gcatcatgcc 300 tctcgattat aaattacaat ggccacaaag agggtggatg gaggagcagg gattagatgg 360 ggtaacccaa tcccagtgcc gggaaggggg acagagggcc cgggagctgc cccccgccac 420 taggagccac tcggagtggc ctctgtctac tgtgtctagg gacagcaaca aggcctactc 480 aatacctgct ttgtgtagcc tgggataatg atttacaaac tcaaggaaat caggaggtgt 540 gaaaagaaga tagccctctt aacccagcca ttgtctaccc tctgctggaa tgtccctcag 600 ctccacccac ctggtgcacc tgtgcacacg gatgctgttt gaggatggtc cacgtgagga 660 gccctcatag gggctggact agatcttata gggctgtctg ggagccactc cagggccaca 720 gaaatcttgt ctctgactca gggtattttg ttttctgttt tgtgtaaatg ctcttctgac 780 taatgcaaac catgtgtcca tagaaccaga agatttttcc aggggaaaag gtaaggaggt 840 ggtgagagtg tcctgggtct gcccttccag ggcttgccct gggttaagag ccaggcagga 900 agctctcaag agcattgctc aagagtagag ggggcctggg aggcccaggg aggggatggg 960
aggggaacac ccaggctgcc cccaaccaga tgccctccac cctcctcaac ctccctccca 1020 cggcctggag aggtgggacc aggtatggag gcttgagagc ccctggttgg aggaagccac 1080 aagtccagga acatgggagt ctgggcaggg ggcaaaggag gcaggaacag gccatcagcc 1140 aggacaggtg gtaaggcagg caggagtgtt cctgctggga aaaggtggga tcaagcacct 1200 ggagggctct tcagagcaaa gacaaacact gaggtcgctg ccactcctac agagccccca 1260 cgccccgccc agctataagg ggccatgcmc caagcagggt acccaggctg cagagg 1316 *wherein M is C or A, preferably M is C Bases in bold and underlined represent a BglII restriction enzyme site created to operably link the enhancer and promoter sequences. For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 77). SEQ ID NO: 43 CpG‐Free CMV Enhancer & core SFTPB Promoter (938bp)* gttacataac ttatggtaaa tggcctgcct ggctgactgc ccaatgaccc ctgcccaatg 60 atgtcaataa tgatgtatgt tcccatgtaa tgccaatagg gactttccat tgatgtcaat 120 gggtggagta tttatggtaa ctgcccactt ggcagtacat caagtgtatc atatgccaag 180 tatgccccct attgatgtca atgatggtaa atggcctgcc tggcattatg cccagtacat 240 gaccttatgg gactttccta cttggcagta catctatgta ttagtcattg ctattaccat 300 ggagatctta tagggctgtc tgggagccac tccagggcca cagaaatctt gtctctgact 360 cagggtattt tgttttctgt tttgtgtaaa tgctcttctg actaatgcaa accatgtgtc 420 catagaacca gaagattttt ccaggggaaa aggtaaggag gtggtgagag tgtcctgggt 480 ctgcccttcc agggcttgcc ctgggttaag agccaggcag gaagctctca agagcattgc 540 tcaagagtag agggggcctg ggaggcccag ggaggggatg ggaggggaac acccaggctg 600 cccccaacca gatgccctcc accctcctca acctccctcc cacggcctgg agaggtggga 660 ccaggtatgg aggcttgaga gcccctggtt ggaggaagcc acaagtccag gaacatggga 720 gtctgggcag ggggcaaagg aggcaggaac aggccatcag ccaggacagg tggtaaggca 780 ggcaggagtg ttcctgctgg gaaaaggtgg gatcaagcac ctggagggct cttcagagca 840 aagacaaaca ctgaggtcgc tgccactcct acagagcccc cacgccccgc ccagctataa 900 ggggccatgc mccaagcagg gtacccaggc tgcagagg 938 *wherein M is C or A, preferably M is C Bases in bold and underlined represent a BglII restriction enzyme site created to operably link the enhancer and promoter sequences. For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 78). SEQ ID NO: 44 ELF3 Enhancer & core SFTPB Promoter (1290bp)*
tagggatggg ccgaggctgg cactgatgct agacttccgt gcacagggca agtatggaca 60 agccccaagt ggctttgtga ggcccacaca gtgaagcttg ggaaatggga agtggggctg 120 cgcccagatt ctggtatcta tgacaactaa ggccgctgca catcctcatg gctctcccag 180 agacctcagg tgaggccctt ctgtgttcct caagcaccca tgccacctgc ggggtggggc 240 aggaccctcc tacccagccc tgggcctcct ggggaaccat gggtgcacag gggtaacctg 300 agccagcctc tctggggcat gggcgggcgt ggtggcgtgg gcctggcgcc agagagtgga 360 gcagaatgtc agctctgtga gccgcaccgg gtgccagcac tctgcaaaca gactctagtc 420 accagataga ctggagtcac gaacctaaca aagcgctcag ctgggcaact gtaactgcag 480 agggcggggc cgcacagtgc tgcctagtgc ctcctgcctt gatctggtca gggtcatggg 540 aggagacagg ttccccagga gggcccatta acaagattaa ttggggtggc ctgggggact 600 ctgggatgct cactggaaac atatcctggg ggaatgaggt aggtggggag ccagagatct 660 tatagggctg tctgggagcc actccagggc cacagaaatc ttgtctctga ctcagggtat 720 tttgttttct gttttgtgta aatgctcttc tgactaatgc aaaccatgtg tccatagaac 780 cagaagattt ttccagggga aaaggtaagg aggtggtgag agtgtcctgg gtctgccctt 840 ccagggcttg ccctgggtta agagccaggc aggaagctct caagagcatt gctcaagagt 900 agagggggcc tgggaggccc agggagggga tgggagggga acacccaggc tgcccccaac 960 cagatgccct ccaccctcct caacctccct cccacggcct ggagaggtgg gaccaggtat 1020 ggaggcttga gagcccctgg ttggaggaag ccacaagtcc aggaacatgg gagtctgggc 1080 agggggcaaa ggaggcagga acaggccatc agccaggaca ggtggtaagg caggcaggag 1140 tgttcctgct gggaaaaggt gggatcaagc acctggaggg ctcttcagag caaagacaaa 1200 cactgaggtc gctgccactc ctacagagcc cccacgcccc gcccagctat aaggggccat 1260 gcmccaagca gggtacccag gctgcagagg 1290 *wherein M is C or A, preferably M is C Bases in bold and underlined represent a BglII restriction enzyme site created to operably link the enhancer and promoter sequences. For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 79). SEQ ID NO: 45 SV40 Enhancer & mSPB Promoter (873bp)* cgatggagcg gagaatgggc ggaactgggc ggagttaggg gcgggatggg cggagttagg 60 ggcgggacta tggttgctga ctaattgaga tgcatgcttt gcatacttct gcctgctggg 120 gagcctgggg actttccaca cctggttgct gactaattga gatgcatgct ttgcatactt 180 ctgcctgctg gggagcctgg ggactttcca caccctaact gacacacatt ccacagcaga 240 tcttataggg ctgtctggga gccactccag ggccacagaa atcttgtctc tgactcaggg 300 tattttgttt tctgttttgt gtaaatgctc ttctgactaa tgcaaaccat gtgtccatag 360 aaccagaaga tttttccagg ggaaaaggta aggaggtggt gagagtgtcc tgggtctgcc 420 cttccagggc ttgccctggg ttaagagcca ggcaggaagc tctcaagagc attgctcaag 480 agtagagggg gcctgggagg cccagggagg ggatgggagg ggaacaccca ggctgccccc 540
aaccagatgc cctccaccct cctcaacctc cctcccacgg cctggagagg tgggaccagg 600 tatggaggct tgagagcccc tggttggagg aagccacaag tccaggaaca tgggagtctg 660 ggcagggggc aaaggaggca ggaacaggcc atcagccagg acaggtggta aggcaggcag 720 gagtgttcct gctgggaaaa ggtgggatca agcacctgga gggctcttca gagcaaagac 780 aaacactgag gtcgctgcca ctcctacaga gcccccacgc cccgcccagc tataaggggc 840 catgcmccaa gcagggtacc caggctgcag agg 873 *wherein M is C or A, preferably M is C Bases in bold and underlined represent a BglII restriction enzyme created to operably link the enhancer and promoter sequences. For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 80). SEQ ID NO: 46 Alv‐01 (CMV forward enhancer + mSFPB promoter) (955bp) AGATCTGTTACATAACTTATGGTAAATGGCCTGCCTGGCTGACTGCCCAATGACCCCTGCCCAATGATGTCAAT AATGATGTATGTTCCCATGTAATGCCAATAGGGACTTTCCATTGATGTCAATGGGTGGAGTATTTATGGTAACT GCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTATGCCCCCTATTGATGTCAATGATGGTAAATGGCC TGCCTGGCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTATGTATTAGTCATTGC TATTACCATGGAGATCTTATAGGGCTGTCTGGGAGCCACTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGG GTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTAATGCAAACCATGTGTCCATAGAACCAGAAGATTTT TCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCTGGGTCTGCCCTTCCAGGGCTTGCCCTGGGTTAAGA GCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGTAGAGGGGGCCTGGGAGGCCCAGGGAGGGGATGG GAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCTCCACCCTCCTCAACCTCCCTCCCACGGCCTGGAGAG GTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTGGAGGAAGCCACAAGTCCAGGAACATGGGAGTCTG GGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGCCAGGACAGGTGGTAAGGCAGGCAGGAGTGTTCCT GCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCTTCAGAGCAAAGACAAACACTGAGGTCGCTGCCAC TCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGGCCATGCMCCAAGCAGGGTACCCAGGCTGCAGAG GGTGCCGCTAGC *wherein M is C or A, preferably M is C (C is present in the exemplified Alv‐01) The bold and underlined 3’ sequence is an NheI restriction site. For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 81).
SEQ ID NO: 47 Alv‐02 (SLC34A2‐2 forward enhancer + mSFPB promoter) (1056bp) AGATCTAGCTAACCATTCAGACTGTCAGTTTAGATTATTGAAAGGAACAGAAGAGAAATCTTGCCTATAAAAC ACGGAACTGAGAAAAAAGTGTATACCTCACTTAGAGCGTCATTAAGTTAGAGAACAACTCCCACACTCTGCTC TAAGATGAGGCCAACATCCAGATTCTCACTGCAAACTTGGCTCCCCTAGAATTCTGTTCTTCTTTCCCCTGTCCC CGCTGCTTTCTAAACGTGAAATATCCACAGCTGCACCGTTTCTTTACTTTTTTTTTTTTTTCTGGAGAGGTTAACC TTCCTTCTTCAGTTGTATGTTGTGCAATACCCAAGAGGCCAACACTACCATTTGAGATATTTAAATATGACTTTG AGAAATGAGTCTGTCTTAGATATAAAAGCCCACAGTCAGATCTTATAGGGCTGTCTGGGAGCCACTCCAGGGC CACAGAAATCTTGTCTCTGACTCAGGGTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTAATGCAAACC ATGTGTCCATAGAACCAGAAGATTTTTCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCTGGGTCTGCC CTTCCAGGGCTTGCCCTGGGTTAAGAGCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGTAGAGGGGGC CTGGGAGGCCCAGGGAGGGGATGGGAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCTCCACCCTCCT CAACCTCCCTCCCACGGCCTGGAGAGGTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTGGAGGAAGC CACAAGTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGCCAGGACAG GTGGTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCTTCAGAGCA AAGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGGCCATGCMC CAAGCAGGGTACCCAGGCTGCAGAGGGTGCCGCTAGC *wherein M is C or A, preferably M is C (C is present in the exemplified Alv‐02) The bold and underlined 3’ sequence is an NheI restriction site. For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 82). SEQ ID NO: 48 Alv‐03 (VEGFA‐2 forward enhancer + mSFPB promoter) (1333bp) AGATCTGAGGCACAAAGCGATCCCCATCACTGCTCCACAATCATTCATTAGCTAACAAGACAGAGCAGCTCAT AAAAAAAAAGCCGTTAAAAAAATTCCGGGGAAATGGAAAGCAGGAGGTGATGCAAGCCCTGGTTAACAAAG GCTGAGGGTTGGGGGGAGGCATGAGAGGGTGTGAGTGGAATAACCCAAGCCTGATAAGCCACAAAGCAGC GCCTGCTGCTCCCTCCTCCCTCTGCCGCCTGAGTCAGAGAAGCCGGGATGTGTTCAAAATCAAGCAATGTAATT AGCAGGTTGCATCATGCCTCTCGATTATAAATTACAATGGCCACAAAGAGGGTGGATGGAGGAGCAGGGATT AGATGGGGTAACCCAATCCCAGTGCCGGGAAGGGGGACAGAGGGCCCGGGAGCTGCCCCCCGCCACTAGGA GCCACTCGGAGTGGCCTCTGTCTACTGTGTCTAGGGACAGCAACAAGGCCTACTCAATACCTGCTTTGTGTAG CCTGGGATAATGATTTACAAACTCAAGGAAATCAGGAGGTGTGAAAAGAAGATAGCCCTCTTAACCCAGCCAT TGTCTACCCTCTGCTGGAATGTCCCTCAGCTCCACCCACCTGGTGCACCTGTGCACACGGATGCTGTTTGAGGA TGGTCCACGTGAGGAGCCCTCATAGGGGCTGGACTAGATCTTATAGGGCTGTCTGGGAGCCACTCCAGGGCC
ACAGAAATCTTGTCTCTGACTCAGGGTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTAATGCAAACCA TGTGTCCATAGAACCAGAAGATTTTTCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCTGGGTCTGCCC TTCCAGGGCTTGCCCTGGGTTAAGAGCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGTAGAGGGGGCC TGGGAGGCCCAGGGAGGGGATGGGAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCTCCACCCTCCTC AACCTCCCTCCCACGGCCTGGAGAGGTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTGGAGGAAGCC ACAAGTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGCCAGGACAGG TGGTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCTTCAGAGCAA AGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGGCCATGCMCC AAGCAGGGTACCCAGGCTGCAGAGGGTGCCGCTAGC *wherein M is C or A, preferably M is C (C is present in the exemplified Alv‐03) The bold and underlined 3’ sequence is an NheI restriction site. For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 83). SEQ ID NO: 49 ATTTGAGCTCTTCTTTCTGCTGAACCATCG Underlined sequence is an SacI restriction enzyme site SEQ ID NO: 50 SFTPB promoter reverse primer TCTTAGATCTGTCAGACAGCTCTGGGTTCC Underlined sequence is a BgIII restriction enzyme site SEQ ID NO: 51 Exemplary linker between SFTPB promoter fragment and enhancer agatct SEQ ID NO: 52 Exemplified WPRE component (mWPRE) 1 GGGCCCAATC AACCTCTGGA TTACAAAATT TGTGAAAGAT TGACTGGTAT TCTTAACTAT 61 GTTGCTCCTT TTACGCTATG TGGATACGCT GCTTTAATGC CTTTGTATCA TGCTATTGCT 121 TCCCGTATGG CTTTCATTTT CTCCTCCTTG TATAAATCCT GGTTGCTGTC TCTTTATGAG 181 GAGTTGTGGC CCGTTGTCAG GCAACGTGGC GTGGTGTGCA CTGTGTTTGC TGACGCAACC 241 CCCACTGGTT GGGGCATTGC CACCACCTGT CAGCTCCTTT CCGGGACTTT CGCTTTCCCC 301 CTCCCTATTG CCACGGCGGA ACTCATCGCC GCCTGCCTTG CCCGCTGCTG GACAGGGGCT 361 CGGCTGTTGG GCACTGACAA TTCCGTGGTG TTGTCGGGGA AATCATCGTC CTTTCCTTGG
421 CTGCTCGCCT GTGTTGCCAC CTGGATTCTG CGCGGGACGT CCTTCTGCTA CGTCCCTTCG 481 GCCCTCAATC CAGCGGACCT TCCTTCCCGC GGCCTGCTGC CGGCTCTGCG GCCTCTTCCG 541 CGTCTTCGCC TTCGCCCTCA GACGAGTCGG ATCTCCCTTT GGGCCGCCTC CCCGCAAGCT SEQ ID NO: 53 Exemplified SFTPB transgene (NM_000542.5) (1146bp) ATGGCTGAGTCACACCTGCTGCAGTGGCTGCTGCTGCTGCTGCCCACGCTCTGTGGCCCAGGCACTGCTGCCT GGACCACCTCATCCTTGGCCTGTGCCCAGGGCCCTGAGTTCTGGTGCCAAAGCCTGGAGCAAGCATTGCAGTG CAGAGCCCTAGGGCATTGCCTACAGGAAGTCTGGGGACATGTGGGAGCCGATGACCTATGCCAAGAGTGTG AGGACATCGTCCACATCCTTAACAAGATGGCCAAGGAGGCCATTTTCCAGGACACGATGAGGAAGTTCCTGG AGCAGGAGTGCAACGTCCTCCCCTTGAAGCTGCTCATGCCCCAGTGCAACCAAGTGCTTGACGACTACTTCCC CCTGGTCATCGACTACTTCCAGAACCAGACTGACTCAAACGGCATCTGTATGCACCTGGGCCTGTGCAAATCCC GGCAGCCAGAGCCAGAGCAGGAGCCAGGGATGTCAGACCCCCTGCCCAAACCTCTGCGGGACCCTCTGCCAG ACCCTCTGCTGGACAAGCTCGTCCTCCCTGTGCTGCCCGGGGCCCTCCAGGCGAGGCCTGGGCCTCACACACA GGATCTCTCCGAGCAGCAATTCCCCATTCCTCTCCCCTATTGCTGGCTCTGCAGGGCTCTGATCAAGCGGATCC AAGCCATGATTCCCAAGGGTGCGCTAGCTGTGGCAGTGGCCCAGGTGTGCCGCGTGGTACCTCTGGTGGCGG GCGGCATCTGCCAGTGCCTGGCTGAGCGCTACTCCGTCATCCTGCTCGACACGCTGCTGGGCCGCATGCTGCC CCAGCTGGTCTGCCGCCTCGTCCTCCGGTGCTCCATGGATGACAGCGCTGGCCCAAGGTCGCCGACAGGAGA ATGGCTGCCGCGAGACTCTGAGTGCCACCTCTGCATGTCCGTGACCACCCAGGCCGGGAACAGCAGCGAGCA GGCCATACCACAGGCAATGCTCCAGGCCTGTGTTGGCTCCTGGCTGGACAGGGAAAAGTGCAAGCAATTTGT GGAGCAGCACACGCCCCAGCTGCTGACCCTGGTGCCCAGGGGCTGGGATGCCCACACCACCTGCCAGGCCCT CGGGGTGTGTGGGACCATGTCCAGCCCTCTCCAGTGTATCCACAGCCCCGACCTTTGA SEQ ID NO: 54 Exemplified human SFTPB (hSP‐B) transgene (1162bp) GCTAGCCACCATGGCTGAGTCACACCTGCTGCAGTGGCTGCTGCTGCTGCTGCCCACGCTCTGTGGCCCAGGC ACTGCTGCCTGGACCACCTCATCCTTGGCCTGTGCCCAGGGACCTGAGTTCTGGTGCCAAAGCCTGGAGCAAG CATTGCAGTGCAGAGCCCTAGGGCATTGCCTACAGGAAGTCTGGGGACATGTGGGAGCCGATGACCTATGCC AAGAGTGTGAGGACATCGTCCACATCCTTAACAAGATGGCCAAGGAGGCCATTTTCCAGGACACGATGAGGA AGTTCCTGGAGCAGGAGTGCAACGTCCTCCCCTTGAAGCTGCTCATGCCCCAGTGCAACCAAGTGCTTGACGA CTACTTCCCCCTGGTCATCGACTACTTCCAGAACCAGACTGACTCAAACGGCATCTGTATGCACCTGGGCCTGT GCAAATCCCGGCAGCCAGAGCCAGAGCAGGAGCCAGGGATGTCAGACCCCCTGCCCAAACCTCTGCGGGACC CTCTGCCAGACCCTCTGCTGGACAAGCTCGTCCTCCCTGTGCTGCCCGGGGCTCTCCAGGCGAGGCCTGGGCC TCACACACAGGATCTCTCCGAGCAGCAATTCCCCATTCCTCTCCCCTATTGCTGGCTCTGCAGGGCTCTGATCA AGCGGATCCAAGCCATGATTCCCAAGGGTGCCCTAGCTGTGGCAGTGGCCCAGGTGTGCCGCGTGGTGCCTC TGGTGGCGGGCGGCATCTGCCAGTGCCTGGCTGAGCGCTACTCCGTCATCCTGCTCGACACGCTGCTGGGCC GCATGCTGCCCCAGCTGGTCTGCCGCCTCGTCCTCCGGTGCTCCATGGATGACAGCGCTGGCCCAAGGTCGCC
GACAGGAGAATGGCTGCCGCGAGACTCTGAGTGCCACCTCTGCATGTCCGTGACCACCCAGGCCGGGAACAG CAGCGAGCAGGCCATACCACAGGCAATGCTCCAGGCCTGTGTTGGCTCCTGGCTGGACAGGGAAAAGTGCAA GCAATTTGTGGAGCAGCACACGCCCCAGCTGCTGACCCTGGTGCCCAGGGGCTGGGATGCCCACACCACCTG CCAGGCCCTCGGGGTGTGTGGGACCATGTCCAGCCCTCTCCAGTGTATCCACAGCCCCGACCTTTGAGGGCCC SEQ ID NO: 55 Exemplified codon‐optimised human SFTPB (cohSP‐B) transgene (1162bp) GCTAGCCACCATGGCCGAGAGTCACCTGCTTCAGTGGCTCTTGTTGCTCTTGCCAACGCTCTGTGGTCCAGGA ACAGCGGCCTGGACAACTTCATCACTGGCTTGTGCTCAGGGACCCGAGTTTTGGTGCCAATCCCTGGAGCAGG CTCTTCAGTGTCGGGCTCTGGGACACTGCTTGCAGGAGGTATGGGGACACGTCGGCGCCGATGACCTTTGCC AGGAATGTGAAGATATAGTCCACATTCTGAATAAGATGGCCAAGGAAGCTATTTTTCAGGATACGATGCGCA AGTTTTTGGAACAAGAGTGCAACGTGCTCCCTCTTAAACTCCTTATGCCGCAGTGTAATCAGGTCCTTGATGAC TACTTTCCCCTGGTTATCGACTATTTCCAGAATCAGACAGACTCCAATGGAATCTGCATGCACCTGGGTTTGTG CAAGTCTAGACAGCCAGAACCGGAACAAGAGCCTGGCATGTCCGACCCGCTGCCTAAGCCTCTCCGGGACCC TCTCCCGGACCCGCTGTTGGACAAATTGGTATTGCCAGTGTTGCCAGGAGCACTCCAAGCGCGACCAGGACCG CACACGCAGGATCTTTCCGAACAACAGTTTCCGATACCACTTCCATACTGCTGGTTGTGCCGAGCTTTGATAAA ACGGATTCAGGCGATGATTCCAAAGGGAGCGCTGGCTGTCGCGGTGGCGCAAGTCTGCCGCGTTGTACCATT GGTCGCTGGAGGTATCTGCCAATGCTTGGCGGAGCGCTATTCAGTGATTCTGCTCGACACCTTGCTTGGTCGC ATGCTCCCGCAACTCGTGTGCCGCCTTGTTCTCCGATGCAGTATGGACGACAGTGCGGGGCCACGGAGCCCCA CCGGTGAATGGTTGCCACGAGATTCAGAATGCCATCTGTGTATGTCTGTTACTACGCAGGCCGGAAACAGCAG CGAACAAGCGATCCCCCAAGCCATGTTGCAAGCATGCGTAGGCTCCTGGCTTGACAGGGAAAAATGCAAGCA GTTCGTGGAGCAACATACGCCGCAGTTGCTGACTCTTGTGCCACGGGGTTGGGACGCTCATACAACCTGTCAG GCATTGGGAGTGTGCGGAACAATGTCTTCACCGCTGCAATGCATTCATAGTCCCGATTTGTGAGGGCCC SEQ ID NO: 56 Homo sapiens surfactant protein B (SFTPB), RefSeqGene on chromosome 2. (NCBI Reference Sequence: NG_016967.1) (18425bp) ttggctgtca ctctctacca agctccccaa cactctgaac ctctatttgt ggtgcagagt 60 ggtgaactct ccctgtacca gccccatgct ctctcttttc ctaggccttt gcatgaacta 120 tttatctgcc tggaaatcct cctcaatcaa tcccgcccca ccatttgacc aggccaattt 180 ctatgcctcc ccaagtctcc ccgtagctgc cccatctccc ccagcgtctc ctcttcctct 240 ccaggcctgg gtagtattcc actctgatta ctggtctgtg ccctcatcca tagatattaa 300 gtccaccagg gtggtgacca ttgtcatccc aactcctagc tcattaacca ttttttggtc 360 aaacaaatga attgtgaaag gtctttataa actgtaagtt tcattgccaa gggccctacc 420 tccctacccc ttcccacgga gaaattcctt cttttgaaag aaagacataa gatagagcta 480 tgctaggaaa ggctgctacc atgaggcctc ctgaggtgca ctggccttgt atttctcatg 540 ggttggttct accccagaaa cagaaactac actcaaataa attcaactac gaagagtggg 600 gataggagag gggaggaaag tcatgggtcc ttaggctcag gccccaagca gagtcagaag 660
aatgcagccc aaaatacaac aaaaggaact aaaatgaatt ctgctcttcc gtgcatctct 720 gcaaggacat aaaagagatt ttaatatcaa ctgtcccatt tggctgatct cttttttccc 780 agccccaggt gtctcaagtc ccacacttgg gagtctcaca ctttccacag tcaagaagaa 840 aaacaagact cttctccacc tgcagcctgg tttgtttctc tctcttcccc tttctttttc 900 aattgtgatg aaatatacaa atatagataa cataatactg ccttaaccat ttttaagtgt 960 acagctcggg gacgtcaagc ccacccacat tatgcaacca aaaccgtcgt tcatctccag 1020 aacttttcat cacaggatga aactttgttt ccattaacca gtacctccgc gttccctctt 1080 ccccagcccc tggtaacctc tgacctactt tctgtcccta cgagtttgac tactctgcat 1140 acttcctata aatggaatca cacaatatct gtccttttgt gactggatta tttcactgag 1200 cgtcatgttt tcaagattca tccatgttgt agcatgtgtc caaatttctt tccttttttt 1260 tcctaccccc acccacccca gacggagtat tgctctgttg cccaggctag agtgcagtgg 1320 cgcaatctca gttctctgta acctccgcct cccgggttca agcgattctc cagcctcagc 1380 ctcccgagta gctgggatta caggcatgca ccaccaggcc cggctaattt ttgtattttt 1440 agtagagacg gggtttcacc gtgttggtaa ggctggtctc gaactcctga cctcgtgatc 1500 tgcccacctt ggcctcccaa agtcctggga ttacaggcat gagccactgc gcctggctga 1560 atttccttcc taaaaaggct caataatatt cctttgagtg tgtacaacac attttgttta 1620 cccattcatc cgtggatgaa cacttgtttc taccttttgg ctattatgaa taatgctgca 1680 atgaaaactg acatacaaat atctgttccg gtacctgttt tccattctct tgaatacgta 1740 cgtaggagtg gaattgctgg gtctgaaagt aattcaatat tcaaactttt cagaaactgt 1800 tttgtaaaaa ctattttcca cacaggtttc cacagtgact gcaccatttc acattctcgg 1860 cagcaaagca cctagggttc caatttctcc acatcctctc aacatgtatt attttctggt 1920 ttttcataat agccgtccta gagggtgtga ggtatggttt tgatttgctc cattttgcat 1980 taagtggcca aggaccccct tggcaatgca ccattaactt gactttgatc cttcacccca 2040 caatcagctg acttggctct catctcccaa tttaaaaatc caattggttc agccaagcat 2100 gcagatggac tcttttgggg tctggtgtgt ctacctctgg tccagtcggc tctggccagg 2160 gggatgaagg gaggtggtcc atgaggcttc cctttccagg gatgtggtgc gtggctttct 2220 agacagtatt actggggcag gcaggggcag acctgtgtgg tgtagaccca tggacacaat 2280 gctaggcggg acatgtgtct tctatatttt gatgacaagg aactggctct cagaaacatg 2340 agtttgtcca cacaccaaga gaagaggcat ctcctgtgaa taaacttgag ttaatcaagt 2400 tctccatccc cagagtagct atcagaatgt tctttttttt tttttttttt tcccagaatg 2460 tttctttatt catttccagt gagaagcagg taggccagga gaagaaccgt cagtagaggc 2520 agatccatgc aggttccccc tgcagccacg cttcacagga agcccgattt tcatcagaga 2580 ctgtggtgag gatggcatca aagcaagggg tgccaggaga ctgcccatcc cagggaccag 2640 gagccaggac accctcagct cagtcaacag ctggcttttt gtgtctcctt tattggtgct 2700 cttggatggc agctctcgac agtcagatac caatccctct ttgaggttct gataaaagcc 2760 acggacccac tacctagaaa accccacata agcacctggt tttgcattca gcattatggc 2820 ttttagacaa tgctatgtgt ccaccaactc catctccctt ccctcctggg catgcaggcg 2880 cattgcactt cccagctttc cctggggtta gtgagggtgt gagactgaac tgagagggaa 2940 gaaatgacgg acaccactcc caggcctcgt ccctaaaacc tcctccacca tcttccactc 3000 tttctctctt ctccctgttg gccagatgca ggggacccag gggaggactc tgagccctcc 3060 aagcaagcag aagggggact ccgagcacct gaatgacggt acggacctga gtctcaccac 3120
ccggagccaa cctgtgctgc actgtaaccg agtaagaagg aaacatgtgg gccaggcgcg 3180 gtggctcaca cctataatcc cagcactttg ggaggccaag gcaggcggat cacctgagat 3240 caggagttcg aaaccaacct ggtcaacatg gtgaaacccc atctctacta aaaaaataca 3300 aaaattagct gggcgtgatc acgggtgcct gtactcccca gctattcagg aggctgaggc 3360 aggagaatcg cttgaacccg ggaggcagag gttgcagtga gccaagatcg taccactgca 3420 ctccagcctg ggcaacagag tgagactcca tctcaaaaaa aaaaagaagg aaacatgtgg 3480 ctaggcacag tggcttatgc ctgtaatccc agcactttgg gaggctgagg tgggtggatc 3540 acctgagctc aggagttcga aacgagcctg ggcaatatgg caaaaccctg tctctaccaa 3600 aaatatcata gaaaaaaatt agccagacgt ggtggcgtat gcctgtggtc cgagctactc 3660 tggagactta ggtgggagga tcacttgagc ctgggagaca gaggttgcag tgagccaaga 3720 tcacgccact gcactccagc ctgagcgcca gagtgagacc ccatgtcaaa aaacaaaaaa 3780 aaaaagaaaa agaaacttgc atgaagccca tagactggag ggctctctgt tatggtagtg 3840 ggaccaccct gaccaatacg cccctggagg cagccagcca caaatcccca ctctcttgcg 3900 cagcctgtcc tgattgtatt caaacgagat accactcgct gatttgttta atacatactg 3960 aataacaagc acgtttcagt tccaccttcc tgaggttcta ccaaggtcac tcagttccca 4020 gtctttggtt actcattttt acacagtctt ggtggactcc catgatggca gctgtgtttg 4080 ttgtactgac aaggacttga ggccaacaga aactacataa atccttgggt aaccttgagt 4140 gtggcattag ccaccttgta tctaattctt tttttttttt tttttttgag acagagtgtt 4200 acctggccaa catggtgtca ggagactgcc catcccaggg accaggagcc aggacaccct 4260 cagctcagtc acatgtttaa tggccacatg tttaatggcc tcacactccc attccagagt 4320 gcagcggtgc aatcatggct cactgcagcc gcgacctggt gggctcaagc gatcctcctg 4380 ccccagcttt ctgagtagct gggaccagag acatgcacca ccacacaagg ttaatttttt 4440 aatttttttg tagttatgga gtcttactat attggccagg ctggtctcaa actccagggc 4500 tcaaaggatc ctccctcctc ggcctcccaa agtgccagga ttacaggagt gagccaccac 4560 acccagcccc atctcttttc atcatggtac taattcctgc ccgtccaccc acaaaagcac 4620 tgtagtgctt cccgagtata gaggcctgtg agcctccact agggagaggg ctcctgcaga 4680 gatcagataa attgatcaca atggctgggg tggtggcaat gtgctaatgc tctctttctt 4740 ccactcaaga tatcctctgt ctccctcagc ctgtgagctt tttctccagt gtgctctgcc 4800 agtgggggcc ctgcctgaga gcccctgcag ctgcagagga cagtttcttt ctgctgaacc 4860 atcgcagcta tgccccagcc cctaccctgg aggggtcccc aggggccatg ggcagcacct 4920 cctgtatagg gctgtctggg agccactcca gggccacaga aatcttgtct ctgactcagg 4980 gtattttgtt ttctgttttg tgtaaatgct cttctgacta atgcaaacca tgtgtccata 5040 gaaccagaag atttttccag gggaaaaggt aaggaggtgg tgagagtgtc ctgggtctgc 5100 ccttccaggg cttgccctgg gttaagagcc aggcaggaag ctctcaagag cattgctcaa 5160 gagtagaggg ggcctgggag gcccagggag gggatgggag gggaacaccc aggctgcccc 5220 caaccagatg ccctccaccc tcctcaacct ccctcccacg gcctggagag gtgggaccag 5280 gtatggaggc ttgagagccc ctggttggag gaagccacaa gtccaggaac atgggagtct 5340 gggcaggggg caaaggaggc aggaacaggc catcagccag gacaggtggt aaggcaggca 5400 ggagtgttcc tgctgggaaa aggtgggatc aagcacctgg agggctcttc agagcaaaga 5460 caaacactga ggtcgctgcc actcctacag agcccccacg ccccgcccag ctataagggg 5520 ccatgcacca agcagggtac ccaggctgca gaggtgccat ggctgagtca cacctgctgc 5580
agtggctgct gctgctgctg cccacgctct gtggcccagg cactggtgag tctcccccag 5640 cctcccctct cctaggcagc tccaccactc actgagcact gctttgtgct aggcattaac 5700 ccaagtctgt cctcatttta aagacaaggc agctggggtt cagagagggt tcagagctta 5760 tccaaggtca cacagctggc gggtccagga gcaggtggaa cccagagctg tctgacgtcc 5820 acatgtttaa tggcctcaca ctcccagcaa aactgggtct agagggtggg tgaaatcatg 5880 atgccaggtg tgtagcctgg atcctgatta aggttgctct ggccccaaac cacagctgcc 5940 tggaccacct catccttggc ctgtgcccag ggccctgagt tctggtgcca aagcctggag 6000 caagcattgc agtgcagagc cctagggcat tgcctacagg aagtctgggg acatgtggga 6060 gccgtgagta ccaccaagga tgcatggcaa ctgggggtct gaaatgaagg gtgctgggtg 6120 ggctctggat gggcaggagg agagtggagc ccccataggg gatggatgag atgaaatggg 6180 atgagatgaa atgagatagg ataaaatgga atgggatgga tgcgatggga tacgatgaca 6240 tagaatagat ggagtcggat gaatgggatg ggatgggatg gatgggaggg gaagggatag 6300 gataggatga catagaataa agatggatgg gatgggatgg gatgggatgg gatgacacag 6360 aataaagatg gatggattgg gatggatgaa tagaagagat ggatgggata aattgatatg 6420 gatgagatgg gacaagttgg gctggtgggc agctgcatgt gccttggagt gctctgttgg 6480 cctcttccta agagaacctc cccattggag ctgggagcct cccccactca tgtgtcctcc 6540 accttggggc ccctccctcc ccaggatgac ctatgccaag agtgtgagga catcgtccac 6600 atccttaaca agatggccaa ggaggccatt ttccaggtaa tgatgcccag atcctggatg 6660 aaggttgggg cccaagagat gagggacaga gcagggaaga gctgagcccc ctaaaggggc 6720 catttccagg ctgaggagga ggcctgggtg cctgggaagt cccagctcct cctggctggg 6780 agcaggtcat ggccctgagc tcaatagcac agccagagat ggtcttccct gaggggaagg 6840 gcccctacat gtgcccaact acttaactcc ttggcactcg tgaactccag caccctgggg 6900 gattaggggt cagtctgccc tggtggggcc ttgtgtccag ggacttgggc ggggtagacc 6960 tcagagaggc ccagctgacg gccccctctg gcctcccagg acacgatgag gaagttcctg 7020 gagcaggagt gcaacgtcct ccccttgaag ctgctcatgc cccagtgcaa ccaagtgctt 7080 gacgactact tccccctggt catcgactac ttccagaacc agactgtgag ggctgcaagc 7140 tcacctcctg cctgcctccc cacgcaggcc cctgtgccca cccatgggga gccacacaca 7200 cagcacccca gccagccaga cacacacaca cacacacaca cacacacagc acccaagccg 7260 gccagacaca aacacacagc accccagcca gccggacaca cacacacaca cacacacaac 7320 accccagctg gccggacaca cacacacaca gtaccccagc tggccggaca cacacacaca 7380 cagcacccta tccagacaca tacacacaca cagtacccca gccagctgga aacacacaca 7440 cacacagcac tccatccaga cacataccca cacagtaccc cagccagcca gacacacaca 7500 cacacacaca cacacacaca cacacagcac acacacagca ccccagctgg ccacacacac 7560 acacacacac accctgtcca caaagggcct aggaaactac gtgcccttca gccatgcacc 7620 cgaccatggg cccccaggtt caggtgcaca cggtgggcct gtacgctcac acacccttac 7680 accctcactc tcacacacat gcttacacac ttattcattc tcacatatat gctcatgctc 7740 attcacacac aatcccggcc acctgcccta aagtccccac acagccctat ctttgccttt 7800 tgtcccccca catagagttc taaaccacag cacccccact aggcctgctt cctcccattc 7860 cagtggtccc tgagcccttg ggccggcctg aataggggtg ggcttccctc ccagacccta 7920 acactcccac cctgtgctgt gccccaggac tcaaacggca tctgtatgca cctgggcctg 7980 tgcaaatccc ggcagccaga gccagagcag gagccaggga tgtcagaccc cctgcccaaa 8040
cctctgcggg accctctgcc agaccctctg ctggacaagc tcgtcctccc tgtgctgccc 8100 ggggccctcc aggcgaggcc tgggcctcac acacaggtga gggaggcccc cacagccagt 8160 aaagtggaga tccagagggc tagagccacc tccgaagccc atgggcactg ggccctggga 8220 gaggcagagc cgggaaggtg ataggaagct ccaggcaggg cctaagggag gagggagaga 8280 aagggaggaa gagagagggg aggagagcct ggaggactct tctcccagca cccagcctgg 8340 cctccacctg attctttccc caggatctct ccgagcagca attccccatt cctctcccct 8400 attgctggct ctgcagggct ctgatcaagc ggatccaagc catgattccc aaggtgaggc 8460 atccagggcc tcacagagcc caggagcaca cgcatacctg tagctccctg cagctcccac 8520 ctctctccca actcacaccc ccgtcaggac ccagctggct gccagaagtt aggaggggag 8580 agagccgctt gtgcattgcc cccacccagg ggaccctggg gctcaggctc aggcctggta 8640 ggtgccaggc ctacagttca tgcaacaaac attaagcccc cactgtatgg aggtgccacg 8700 ccaggagcca aagtacaaaa acggacaaga cgcagctttg tcctccagca gctcaccatc 8760 tgatggagaa agatccccag aggtctctgt agaaaggttg ctttgatctt tcaagagggg 8820 aatttccaca gatagattcc ccatccttgc ctgagtccaa cttggagtct tccagacctg 8880 cagtggctat tgtccaatgg ccccgccagc ccagggctac cttgcccaaa ttggggccca 8940 aatgaggaaa ggccctgccc cctcagcctt tcccagatag ggttgcgtgg gccaccaggg 9000 gcacaaggca gcaggtgagg ttcctgctga ggcaggtggt tcacttgagc ccaggagttc 9060 aagaccagct tgggcaacat ggcgaaaccc cgtctctact aagaatacaa aaattagcca 9120 gatgtgacag gtgcctgtag tcccagctac tcgggaggct gaggcaggag aatcacttga 9180 acccaggagg cggaggttgc agtgagccga catcacgcca ctgtactcta gcctgggtga 9240 cagagcaaga ctctgtctca aaaaaaaaga aagaaggaaa gatcactgca gagattgcag 9300 tgagaggtga tgggacaggg acggagctga gggctggcct ggggatgcat ttgggaggtg 9360 ggcccactgc tattgggcat ggatgggcct ggagcgtgag gaccagggag gactccaaag 9420 tgacttttac acactggcca gagcaaccag ccctctgtaa tgccagcagc tgagatgggg 9480 agactaaaga agaaaacagg tttgagcaaa aaaacagaga gctccctcct ggccatgttg 9540 agttcaagat gcctgtgtga agtgcaggag aggagagtca ggcaagcagc tgaatcccaa 9600 gcattggggg aaggtcaggt ccaccatgtc agtctgagag tcactagctg tgggccagag 9660 cctttggggc cagacgtagg tctgaagctg gctcctacac tcagtgaccc tgtgtgagtc 9720 ccctgcatcc cctggactct ctgatcccca gtgtccttat ttgtgaatag ccttgccctc 9780 ccttctagaa gagaatgagg gaatgcgtag gaagtgccca gctgggtgct gggcagagag 9840 tggaggcttg ccaagtgaag gtcccatgct ggcctctctc cgcccccgcc ccagggtgcg 9900 ctagctgtgg cagtggccca ggtgtgccgc gtggtacctc tggtggcggg cggcatctgc 9960 cagtgcctgg ctgagcgcta ctccgtcatc ctgctcgaca cgctgctggg ccgcatgctg 10020 ccccagctgg tctgccgcct cgtcctccgg tgctccatgg atgacagcgc tggcccaagt 10080 gagcccactg cccactcctt agcccaatgc ctgctctcct cctcccccta ccctgccact 10140 gcatgaccct ctccctctgt ggtcccactg caatgcacca aggaggacag aaaccaaaca 10200 cctctgtagg gtggccttgc ctgctttccc cctaatgctc acatctccag ggtcgccgac 10260 aggagaatgg ctgccgcgag actctgagtg ccacctctgc atgtccgtga ccacccaggc 10320 cgggaacagc agcgagcagg ccataccaca ggcaatgctc caggcctgtg ttggctcctg 10380 gctggacagg gaaaaggtat gggctgggca catggggact catggtcagg gcccgttcaa 10440 ggcagaaggc tgagcccagg aaaggctttg cagccagaga cacctaggat gggccagaat 10500
ggagcacaga caggcagaca ggatgtgggg cagacaatgg tgggactgta agttagggca 10560 gagcctgcta aaggttagga gtcgcctctg gacaaagggc tgtgggctcc agaggaccag 10620 caggccctct tcacgggctg agtgagcacc aggcaagcct tcagaggcct ggttatctac 10680 caggagatga gtaatgctag ggccagttca agccaggaaa gggactagcc ttctctccag 10740 ggtcctgatc cctttactgc ccccacactc ctcaaggtgt gactcactca ggacaaaccc 10800 attggcaaaa ggagagggct ggacttgaag gtcctagggc ccttgccaat actcagtcaa 10860 tgacaggaaa ttcccttttt tttttttttt tttttttttt tgagatggag ttttgctctt 10920 gttgcccagg ctggagtgca atggcacaat cttggctcac tgcaacctct gcctccgggt 10980 tcaggcgatt ctcctgcctc agcctcttga gtagctggga ttacaggcat gtgctaccag 11040 gcccggctaa tttttgtatt tttagtagag acaaggtttc accatattgg tcaggctggt 11100 ctcgaacccc tgacctgaag tgatctgccc gccttggcct cccaaagtgc tgggattaca 11160 ggcataagcc actgcacccg gacaggaaat tcccttctta aagcgagatc ctgtcctgag 11220 gaaagccagc tgatgctctt cccaggaggc agctgtccac actgtgctcc ctgctcagca 11280 actcccaagc ctcccgactg cccatcacat ctggtctcaa ggaccagatg aacgttaagt 11340 ttccttctag aactgaaatg gaggtggagg gaggggaggg tggtggctga gattccaccc 11400 ctctgcctga gtcctccgtc tccagtgtcg cctgcttttc tgatggaagt cctccatttc 11460 agctggctcc agtttgttaa gggtttcaac tgcagccaga ggtgttccgt gagggctgat 11520 ggaggagtcg ggagggagcc ctagagtgat ccagagatgt ggagaggcca ggaccacacg 11580 acaggagagt cctgcaaagg gaccctccac agctgtgtgt ctccttcctc agtgcaagca 11640 atttgtggag cagcacacgc cccagctgct gaccctggtg cccaggggct gggatgccca 11700 caccacctgc caggtacacc cacccctccc agttggtcct aggacttccc ttggctccca 11760 gagcccccac cctttgggcc cgtgatcctc agaggcctca ctcccctggg tccaaggtgg 11820 tcccaggtgc acgggccagg gactgggagg cacccctctc tgtttcagtg taaaaaatca 11880 tgagagcatg gaaaaggggg atgggaaggg agggatggcc tgaggagtgc ggctggatgt 11940 ccattatagg atggggctgt gttccctggc cagtgtgtgc tggtggggtg ggggtacaaa 12000 gtgggtgttc tggagtgaac atctcacctc ctcaggctct aaaccctaag gcctgtggct 12060 cagggagtgg cccgaggggt ctacagagtc acactggtag cacccactag gcgggaggtg 12120 gagtgagtgc tgttctttcc cggaagagct gggtgtgggg agctgagggg gcccaggcct 12180 cagccctggt gctgtccctg tgacaggccc tcggggtgtg tgggaccatg tccagccctc 12240 tccagtgtat ccacagcccc gacctttgat gagaactcag ctgtccaggt gagtccaggc 12300 ccccagttgc ggggaggtaa gggggcaggt cctgaccatc agggcatggg aggcccttct 12360 gctccccaag caggaagagg cggccactcc tgccggctgc tccatcctcc ctctcaccgc 12420 acagctggag gctcctgagg gcttctggct ggccatcagg aaaacaccct ttccggaccc 12480 cgagcactgc cccgcccaga accccagtca ctgagtgccc aacccccagc ttccccccca 12540 accccccgcc ctgccctgtc ccaggcctcc ctctcagagc ttgccccagg gactctctgg 12600 ccctcagggt tcaatgtatt ctgaccaagg ccaagctttc ctggggctca gggaaaatca 12660 cactttgcta cccgaagctg tatcccctca gatgccagga aggccgtgat catctgactc 12720 caccctcctg agacacattc tctccctgac tgtcctgttc taagtcagcg gagcacctta 12780 ggatggaggg gtggaggcga ggccagatgc agcctctgtg aacaggtgcc tggaggctgg 12840 gaaatgaccc tgagagggca ggacacagca accgtgggct taaggtgacc ttgagagcaa 12900 gcttggccca ctttacaatt ctgttcagag ccagccccta acatggtggt catttattca 12960
tttgttccct cattttaaaa aatgtaaggc caggcatggt ggctcacgcc ggtaatccca 13020 gcactttggg aggccgaggc aggcagatca cctgaggtca ggagttcgag actagcctgg 13080 ccaacatggc gaaaccctgt ctctactaaa aatatttttt aaaaattagc tgagcatggt 13140 ggcaggtgcc tgtaatccca gctactcagg acgcttaggc aggagaatca cttgaacctg 13200 ggaggcgaag gttgcggtgt gccgagatcg tgccactgca ctctagccta ggcaacagag 13260 cacaactctg tctcaggaaa aaaaaaaaaa aaaaaaaggt atttctttgc tgggcgcagt 13320 ggctcacacc tgtaatccca gcactttggg agaccgaggc gagtggatca cttgaggtca 13380 ggagttcaag accagcctta ccaacatgat gaaaccccgt atctactaaa aaaaaaaaaa 13440 aaaaaaaaaa aaattagcca gatgtggtgg cacacacctg taatcccagc tacttgggag 13500 gctgaggagg agaattgctt gaacctggga ggcggagatt gcagcgagcc aagattgcgc 13560 ctctgcactc cagcctgggt gacagagtga gactccgtct caaaaaaaaa aaaaaaaaag 13620 tagtgggtgc ctgtggccag gccacatcct agggtagggg ctatggctga gccctgccct 13680 cctggagctc acagccaagt ccacttcttc catctgaggc ggggaagcca gccctgttcc 13740 tgaaaccctg catcacaagc ccctgtggga ggcagtgggg aggggaggtc ctcccccact 13800 cagacctgac ccacagggac cagtttaatg tgtccttgcc ccagtgatga cagctgggga 13860 tctgggggtg gggagtcacc caggacccgg gcagtcgcct ttccccagct cctagggctc 13920 ccggccttcc ctgctgaaac agcaagacca gtgggttggc gtgggaggcc tgggcttcaa 13980 accacctctg ctatcacctg gctgtgggtc cccaggcagg acatacacac agtccctctc 14040 tggccctcat cctcctcagc tgcaaaggaa aagccaagtg agacgggctc tgggaccatg 14100 gtgaccaggc tcttcccctg ctccctggcc ctcgccagct gccaggctga aaagaagcct 14160 cagctcccac accgccctcc tcaccgccct tcctcggcag tcacttccac tggtggacca 14220 cgggccccca gccctgtgtc ggccttgtct gtctcagctc aaccacagtc tgacaccaga 14280 gcccacttcc atcctctctg gtgtgaggca cagcgagggc agcatctgga ggagctctgc 14340 agcctccaca cctaccacga cctcccaggg ctgggctcag gaaaaaccag ccactgcttt 14400 acaggacagg gggttgaagc tgagccccgc ctcacaccca cccccatgca ctcaaagatt 14460 ggattttaca gctacttgca attcaaaatt cagaagaata aaaaatggga acatacagaa 14520 ctctaaaaga tagacatcag aaattgttaa gttaagcttt ttcaaaaaat cagcaattcc 14580 ccagcgtagt caagggtgga cactgcacgc tctggcatga tgggatggcg accgggcaag 14640 ctttcttcct cgagatgctc tgctgcttga gagctattgc tttgttaaga tataaaaagg 14700 ggtttctttt tgtctttctg taaggtggac ttccagcttt tgattgaaag tcctagggtg 14760 attctatttc tgctgtgatt tatctgctga aagctcagct ggggttgtgc aagctaggga 14820 cccattcctg tgtaatacaa tgtctgcacc aatgctaata aagtcctatt ctcttttatg 14880 agaaagaaaa agacaccgtc ctttaaagtg ctgcagtatg gccagacgtg gtggctcaca 14940 cctgcaatcc cagcacctta ggaggccgag gcaggaggat ccttgaggtc aggagttcga 15000 gaccagcctc gccaacatgg tgaaacccca tttctactaa aaatacaaaa aattagccaa 15060 gtgtggtggc atatgcctgt aatcccaact actcagaagg ccgaggcagg agaattactt 15120 gaacgcagga gaatcactgc agcccaggag gcagaggttg cagtgagccg agattgcacc 15180 actgcactcc agcctgggtg acagagcaag actccatctc agtaaataaa taaataaata 15240 aaaagcgctg cagtagctgt ggcctcaccc tgaagtcagc gggcccaggc ctacctcact 15300 ctctcccttg gcagagaagc agacgtccat agctcctctc cctcacaagc gctcccagcc 15360 tgccctccag ctgctgctct cccctcccag tctctactca ctgggatgag gttaggtcat 15420
gaggacacca aaaacctaaa aataaacaaa aagccaaaca agccttagct tttcttaaag 15480 actgaaatgc ctggaagtgt ccctttattt ataaaataac ttttgtcata tttcttatac 15540 atgtttcttg taagaaattc agaaactaca gacaaagaga gtggaaatta cccactgtca 15600 ggcctctgag cccaagctaa gccatcatat cccctgtgcc ctgcacgtat acacccagat 15660 ggcctgaagc aactgaagat ccacaaaaga agtgaaaata gccagttcct gccttaactg 15720 atgacattcc accattgtga tttgttcctg ccccacccta actgatcaat tgaccttgtg 15780 acaatacacc ttccccaccc ttgagaaggt gctttgtaat attctcccca cccaccccac 15840 gcccgcaccc ccgcaccctt aagaaggtat tttgtaatat tctctccgcc attgagaatg 15900 tgctttgtaa gatccacccc ctgcccacaa aaaattgctc ctaactccac cgcctatccc 15960 aaacctacaa gaactaatga taatcccacc accctttgct gactcttttt ggactcagcc 16020 cacctgcacc caggtgatta aaaagcttta ttgttcacac aaagcctgtt tggtagtctc 16080 ttcacaggga agcatgtgac acccacaatc ccacctagcc caggagagag ctacggcagg 16140 gtgtgtgttt tgacactgag cttggggctt tttccatctt ctccccacag cctctggctc 16200 cacacctcca ccgttcaagc gccagaaaga gctgtctatg cagcctgctc ttgggcctgg 16260 ggatgagaca cacaattcat tggctcctgg attttaagta gacatttgta aatctatagc 16320 taactactgt ccttaaagcc attgtttcca ttacaaaatc caactctctg agagaaaagg 16380 gtgttttaaa tttaaaaaaa taaaaacaaa aaagtttgat tgagacaatg aagagccaat 16440 gtttgaagaa ctggaggcat taatagttta tcccaaaaat atgagtggga atgagtgcct 16500 gataatgaaa aaggcggctt cattcaaaag gaggtcaaga attggcagac tttgggaggc 16560 tgaggtgagt gaatcacttg agcccaggag ttcaagacca gactgggcaa cgtggcaaaa 16620 caccatctct ataaaaaata taaaaatttg gctgggcatg gtggctcaca cctgtaatcc 16680 tagcactttt ggaggctgag gtgggtggat cacttgaggt taggagttcc agaccagcct 16740 aacatggtga aaccacatat ctactaaaaa aaatacaaaa ttagctggat gtggtagtgc 16800 atgcctgtta tcccagctac ttggaaggct gagacaggag aattgcttga acccaggagg 16860 cagaggctgc agtgagccag gattgcgcca ctgcactcca gcctgggcaa caagagcaaa 16920 actctgtcta aaagaaaaaa aaaaaagata caaacattag ctggccatgg tggtgtgtgt 16980 ctgtggtccc agctattcag gaggctgagg tgggaggatc gcctgaatcc aggtcaaggc 17040 tgcagtgagc cgtgatgatg ccactgcact ccaccctggg tgacacagag agatcctatt 17100 tcaaaaacaa ataaataaat aattttaaaa aagagttgac tagctggtag gaggacagaa 17160 ttggcaggaa gtggctacac ttgagactcc tagggtccct gagaagtgtc caacatcctg 17220 cacctctgcc cttaggaagt tgggggaaga gaagtcccag gagggattat agcccaagag 17280 gaggggactt attcttcact ttccagggtg tgggcttcta agagaagcct gcctctcaca 17340 ggacatcaga aagggtgagg acgctagcct gagggatgga gctaggcagc tggcctgagg 17400 tctgcctcac agagtgttcc cagtgtgatg tggttgcaga agagcaacct caacccacag 17460 accgagaggc agtcagaatt ggctctggac aagatggtgc acccgtaaga gcccacggtt 17520 ggggatgaat ctgtgacaga tttagtggat cttgagatga gcagaaatca gatgtgtagc 17580 agtgatggca caagaagatg atgatcaggc aagacgagag acactccaga gagacctgga 17640 tggccaggga atgcttagtc cccagtgcct ctcctcaccg cagggcaatc tgtaaatcat 17700 tcccagggaa aggggaaaga ggattcagtg gttagtcagc caagagaggc tgagtattaa 17760 gaagggaatg agttactccg aaatggaaat aatatgtttt gagggctcat gagccaaaat 17820 gagttcagtt atagagaaat aaagataatt acattatgtc cccctaatga ataactgaaa 17880
ttttaaatcc acgacataat cattattaac atgttattat ctataattcc aaagctttct 17940 gtgcacatat acaaaaacag atacagcaga tctacaaaat ggaaacatcc tatacatgct 18000 gctttgtaac ctgcactgtt catttaacaa taaccttttc aggtcaataa atgtaagtat 18060 ctactattat tgctaattgc cacatgacag ctctgtctcc caggctagag tgcagtggtg 18120 cgatcactct ggggtttaag caatcctcct gcctctgcct ccctaagtgc tgggactaca 18180 ggcgtcagat accatggctg gcccttactt tttttttttt ttttttttga gacaagagtc 18240 tcactctgtc atccaggctg gattgcagtg gcgtgatctc agctcactgc aagctccgca 18300 tcccaggttc atgccattct cctgcctcag cctcctgagt agctgggact acaggagccc 18360 gccaccacgc ccagctaatt tttttttttt tttttttttg tatttttact agacacgggg 18420 tttca 18425 SEQ ID NO: 57 Exemplified Human SFTPB polypeptide MAESHLLQWLLLLLPTLCGPGTAAWTTSSLACAQGPEFWCQSLEQALQCRALGHCLQEVWGHVGADDLCQECE DIVHILNKMAKEAIFQDTMRKFLEQECNVLPLKLLMPQCNQVLDDYFPLVIDYFQNQTDSNGICMHLGLCKSRQPE PEQEPGMSDPLPKPLRDPLPDPLLDKLVLPVLPGALQARPGPHTQDLSEQQFPIPLPYCWLCRALIKRIQAMIPKGA LAVAVAQVCRVVPLVAGGICQCLAERYSVILLDTLLGRMLPQLVCRLVLRCSMDDSAGPRSPTGEWLPRDSECHLC MSVTTQAGNSSEQAIPQAMLQACVGSWLDREKCKQFVEQHTPQLLTLVPRGWDAHTTCQALGVCGTMSSPLQ CIHSPDL* SEQ ID NO: 58 Exemplified ABACA3 (ABCA3) transgene (NM_000542.5) (5115bp) ATGGCTGTGCTCAGGCAGCTGGCGCTCCTCCTCTGGAAGAACTACACCCTGCAGAAGCGGAAGGTCCT GGTGACGGTCCTGGAACTCTTCCTGCCATTGCTGTTTTCTGGGATCCTCATCTGGCTCCGCTTGAAGA TTCAGTCGGAAAATGTGCCCAACGCCACCATCTACCCGGGCCAGTCCATCCAGGAGCTGCCTCTGTTC TTCACCTTCCCTCCGCCAGGAGACACCTGGGAGCTTGCCTACATCCCTTCTCACAGTGACGCTGCCAA GACCGTCACTGAGACAGTGCGCAGGGCACTTGTGATCAACATGCGAGTGCGCGGCTTTCCCTCCGAGA AGGACTTTGAGGACTACATTAGGTACGACAACTGCTCGTCCAGCGTGCTGGCCGCCGTGGTCTTCGAG CACCCCTTCAACCACAGCAAGGAGCCCCTGCCGCTGGCGGTGAAATATCACCTACGGTTCAGTTACAC ACGGAGAAATTACATGTGGACCCAAACAGGCTCCTTTTTCCTGAAAGAGACAGAAGGCTGGCACACTA CTTCCCTTTTCCCGCTTTTCCCAAACCCAGGACCAAGGGAACCTACATCCCCTGATGGCGGAGAACCT GGGTACATCCGGGAAGGCTTCCTGGCCGTGCAGCATGCTGTGGACCGGGCCATCATGGAGTACCATGC CGATGCCGCCACACGCCAGCTGTTCCAGAGACTGACGGTGACCATCAAGAGGTTCCCGTACCCGCCGT TCATCGCAGACCCCTTCCTCGTGGCCATCCAGTACCAGCTGCCCCTGCTGCTGCTGCTCAGCTTCACC TACACCGCGCTCACCATTGCCCGTGCTGTCGTGCAGGAGAAGGAAAGGAGGCTGAAGGAGTACATGCG CATGATGGGGCTCAGCAGCTGGCTGCACTGGAGTGCCTGGTTCCTCTTGTTCTTCCTCTTCCTCCTCA TCGCCGCCTCCTTCATGACCCTGCTCTTCTGTGTCAAGGTGAAGCCAAATGTAGCCGTGCTGTCCCGC AGCGACCCCTCCCTGGTGCTCGCCTTCCTGCTGTGCTTCGCCATCTCTACCATCTCCTTCAGCTTCAT GGTCAGCACCTTCTTCAGCAAAGCCAACATGGCAGCAGCCTTCGGAGGCTTCCTCTACTTCTTCACCT
ACATCCCCTACTTCTTCGTGGCCCCTCGGTACAACTGGATGACTCTGAGCCAGAAGCTCTGCTCCTGC CTCCTGTCTAATGTCGCCATGGCAATGGGAGCCCAGCTCATTGGGAAATTTGAGGCGAAAGGCATGGG CATCCAGTGGCGAGACCTCCTGAGTCCCGTCAACGTGGACGACGACTTCTGCTTCGGGCAGGTGCTGG GGATGCTGCTGCTGGACTCTGTGCTCTATGGCCTGGTGACCTGGTACATGGAGGCCGTCTTCCCAGGG CAGTTCGGCGTGCCTCAGCCCTGGTACTTCTTCATCATGCCCTCCTATTGGTGTGGGAAGCCAAGGGC GGTTGCAGGGAAGGAGGAAGAAGACAGTGACCCCGAGAAAGCACTCAGAAACGAGTACTTTGAAGCCG AGCCAGAGGACCTGGTGGCGGGGATCAAGATCAAGCACCTGTCCAAGGTGTTCAGGGTGGGAAATAAG GACAGGGCGGCCGTCAGAGACCTGAACCTCAACCTGTACGAGGGACAGATCACCGTCCTGCTGGGCCA CAACGGTGCCGGGAAGACCACCACCCTCTCCATGCTCACAGGTCTCTTTCCCCCCACCAGTGGACGGG CATACATCAGCGGGTATGAAATTTCCCAGGACATGGTTCAGATCCGGAAGAGCCTGGGCCTGTGCCCG CAGCACGACATCCTGTTTGACAACTTGACAGTCGCAGAGCACCTTTATTTCTACGCCCAGCTGAAGGG CCTGTCACGTCAGAAGTGCCCTGAAGAAGTCAAGCAGATGCTGCACATCATCGGCCTGGAGGACAAGT GGAACTCACGGAGCCGCTTCCTGAGCGGGGGCATGAGGCGCAAGCTCTCCATCGGCATCGCCCTCATC GCAGGCTCCAAGGTGCTGATACTGGACGAGCCCACCTCGGGCATGGACGCCATCTCCAGGAGGGCCAT CTGGGATCTTCTTCAGCGGCAGAAAAGTGACCGCACCATCGTGCTGACCACCCACTTCATGGACGAGG CTGACCTGCTGGGAGACCGCATCGCCATCATGGCCAAGGGGGAGCTGCAGTGCTGCGGGTCCTCGCTG TTCCTCAAGCAGAAATACGGTGCCGGCTATCACATGACGCTGGTGAAGGAGCCGCACTGCAACCCGGA AGACATCTCCCAGCTGGTCCACCACCACGTGCCCAACGCCACGCTGGAGAGCAGCGCTGGGGCCGAGC TGTCTTTCATCCTTCCCAGAGAGAGCACGCACAGGTTTGAAGGTCTCTTTGCTAAACTGGAGAAGAAG CAGAAAGAGCTGGGCATTGCCAGCTTTGGGGCATCCATCACCACCATGGAGGAAGTCTTCCTTCGGGT CGGGAAGCTGGTGGACAGCAGTATGGACATCCAGGCCATCCAGCTCCCTGCCCTGCAGTACCAGCACG AGAGGCGCGCCAGCGACTGGGCTGTGGACAGCAACCTCTGTGGGGCCATGGACCCCTCCGACGGCATT GGAGCCCTCATCGAGGAGGAGCGCACCGCTGTCAAGCTCAACACTGGGCTCGCCCTGCACTGCCAGCA ATTCTGGGCCATGTTCCTGAAGAAGGCCGCATACAGCTGGCGCGAGTGGAAAATGGTGGCGGCACAGG TCCTGGTGCCTCTGACCTGCGTCACCCTGGCCCTCCTGGCCATCAACTACTCCTCGGAGCTCTTCGAC GACCCCATGCTGAGGCTGACCTTGGGCGAGTACGGCAGAACCGTCGTGCCCTTCTCAGTTCCCGGGAC CTCCCAGCTGGGTCAGCAGCTGTCAGAGCATCTGAAAGACGCACTGCAGGCTGAGGGACAGGAGCCCC GCGAGGTGCTCGGTGACCTGGAGGAGTTCTTGATCTTCAGGGCTTCTGTGGAGGGGGGCGGCTTTAAT GAGCGGTGCCTTGTGGCAGCGTCCTTCAGAGATGTGGGAGAGCGCACGGTCGTCAACGCCTTGTTCAA CAACCAGGCGTACCACTCTCCAGCCACTGCCCTGGCCGTCGTGGACAACCTTCTGTTCAAGCTGCTGT GCGGGCCTCACGCCTCCATTGTGGTCTCCAACTTCCCCCAGCCCCGGAGCGCCCTGCAGGCTGCCAAG GACCAGTTTAACGAGGGCCGGAAGGGATTCGACATTGCCCTCAACCTGCTCTTCGCCATGGCATTCTT GGCCAGCACGTTCTCCATCCTGGCGGTCAGCGAGAGGGCCGTGCAGGCCAAGCATGTGCAGTTTGTGA GTGGAGTCCACGTGGCCAGTTTCTGGCTCTCTGCTCTGCTGTGGGACCTCATCTCCTTCCTCATCCCC AGTCTGCTGCTGCTGGTGGTGTTTAAGGCCTTCGACGTGCGTGCCTTCACGCGGGACGGCCACATGGC TGACACCCTGCTGCTGCTCCTGCTCTACGGCTGGGCCATCATCCCCCTCATGTACCTGATGAACTTCT TCTTCTTGGGGGCGGCCACTGCCTACACGAGGCTGACCATCTTCAACATCCTGTCAGGCATCGCCACC
TTCCTGATGGTCACCATCATGCGCATCCCAGCTGTAAAACTGGAAGAACTTTCCAAAACCCTGGATCA CGTGTTCCTGGTGCTGCCCAACCACTGTCTGGGGATGGCAGTCAGCAGTTTCTACGAGAACTACGAGA CGCGGAGGTACTGCACCTCCTCCGAGGTCGCCGCCCACTACTGCAAGAAATATAACATCCAGTACCAG GAGAACTTCTATGCCTGGAGCGCCCCGGGGGTCGGCCGGTTTGTGGCCTCCATGGCCGCCTCAGGGTG CGCCTACCTCATCCTGCTCTTCCTCATCGAGACCAACCTGCTTCAGAGACTCAGGGGCATCCTCTGCG CCCTCCGGAGGAGGCGGACACTGACAGAATTATACACCCGGATGCCTGTGCTTCCTGAGGACCAAGAT GTAGCGGACGAGAGGACCCGCATCCTGGCCCCCAGTCCGGACTCCCTGCTCCACACACCTCTGATTAT CAAGGAGCTCTCCAAGGTGTACGAGCAGCGGGTGCCCCTCCTGGCCGTGGACAGGCTCTCCCTCGCGG TGCAGAAAGGGGAGTGCTTCGGCCTGCTGGGCTTCAATGGAGCCGGGAAGACCACGACTTTCAAAATG CTGACCGGGGAGGAGAGCCTCACTTCTGGGGATGCCTTTGTCGGGGGTCACAGAATCAGCTCTGATGT CGGAAAGGTGCGGCAGCGGATCGGCTACTGCCCGCAGTTTGATGCCTTGCTGGACCACATGACAGGCC GGGAGATGCTGGTCATGTACGCTCGGCTCCGGGGCATCCCTGAGCGCCACATCGGGGCCTGCGTGGAG AACACTCTGCGGGGCCTGCTGCTGGAGCCACATGCCAACAAGCTGGTCAGGACGTACAGTGGTGGTAA CAAGCGGAAGCTGAGCACCGGCATCGCCCTGATCGGAGAGCCTGCTGTCATCTTCCTGGACGAGCCGT CCACTGGCATGGACCCCGTGGCCCGGCGCCTGCTTTGGGACACCGTGGCACGAGCCCGAGAGTCTGGC AAGGCCATCATCATCACCTCCCACAGCATGGAGGAGTGTGAGGCCCTGTGCACCCGGCTGGCCATCAT GGTGCAGGGGCAGTTCAAGTGCCTGGGCAGCCCCCAGCACCTCAAGAGCAAGTTCGGCAGCGGCTACT CCCTGCGGGCCAAGGTGCAGAGTGAAGGGCAACAGGAGGCGCTGGAGGAGTTCAAGGCCTTCGTGGAC CTGACCTTTCCAGGCAGCGTCCTGGAAGATGAGCACCAAGGCATGGTCCATTACCACCTGCCGGGCCG TGACCTCAGCTGGGCGAAGGTTTTCGGTATTCTGGAGAAAGCCAAGGAAAAGTACGGCGTGGACGACT ACTCCGTGAGCCAGATCTCGCTGGAACAGGTCTTCCTGAGCTTCGCCCACCTGCAGCCGCCCACCGCA GAGGAGGGGCGATGA SEQ ID NO: 59 Exemplified Human ABACA3 (hABCA3) transgene (5139bp) GCTAGCCACCATGGCTGTGCTCAGGCAGCTGGCGCTCCTCCTCTGGAAGAACTACACCCTGCAGAAGCGGAA GGTCCTGGTGACGGTCCTGGAACTCTTCCTGCCATTGCTGTTTTCTGGGATCCTCATCTGGCTCCGCTTGAAGA TTCAGTCGGAAAATGTGCCCAACGCCACCATCTACCCGGGCCAGTCCATCCAGGAGCTGCCTCTGTTCTTCACC TTCCCTCCGCCAGGAGACACCTGGGAGCTTGCCTACATCCCTTCTCACAGTGACGCTGCCAAGACCGTCACTGA GACAGTGCGCAGGGCACTTGTGATCAACATGCGAGTGCGCGGCTTTCCCTCCGAGAAGGACTTTGAGGACTA CATTAGGTACGACAACTGCTCGTCCAGCGTGCTGGCCGCCGTGGTCTTCGAGCACCCCTTCAACCACAGCAAG GAGCCCCTGCCGCTGGCGGTGAAATATCACCTACGGTTCAGTTACACACGGAGAAATTACATGTGGACCCAAA CAGGCTCCTTTTTCCTGAAAGAGACAGAAGGCTGGCACACTACTTCCCTTTTCCCGCTTTTCCCAAACCCAGGA CCAAGGGAACCTACATCCCCTGATGGCGGAGAACCTGGGTACATCCGGGAAGGCTTCCTGGCCGTGCAGCAT GCTGTGGACCGGGCCATCATGGAGTACCATGCCGATGCCGCCACACGCCAGCTGTTCCAGAGACTGACGGTG ACCATCAAGAGGTTCCCGTACCCGCCGTTCATCGCAGACCCCTTCCTCGTGGCCATCCAGTACCAGCTGCCCCT GCTGCTGCTGCTCAGCTTCACCTACACCGCGCTCACCATTGCCCGTGCTGTCGTGCAGGAGAAGGAAAGGAG
GCTGAAGGAGTACATGCGCATGATGGGGCTCAGCAGCTGGCTGCACTGGAGTGCCTGGTTCCTCTTGTTCTTC CTCTTCCTCCTCATCGCCGCCTCCTTCATGACCCTGCTCTTCTGTGTCAAGGTGAAGCCAAATGTAGCCGTGCTG TCCCGCAGCGACCCCTCCCTGGTGCTCGCCTTCCTGCTGTGCTTCGCCATCTCTACCATCTCCTTCAGCTTCATG GTCAGCACCTTCTTCAGCAAAGCCAACATGGCAGCAGCCTTCGGAGGCTTCCTCTACTTCTTCACCTACATCCC CTACTTCTTCGTGGCCCCTCGGTACAACTGGATGACTCTGAGCCAGAAGCTCTGCTCCTGCCTCCTGTCTAATG TCGCCATGGCAATGGGAGCCCAGCTCATTGGGAAATTTGAGGCGAAAGGCATGGGCATCCAGTGGCGAGAC CTCCTGAGTCCCGTCAACGTGGACGACGACTTCTGCTTCGGGCAGGTGCTGGGGATGCTGCTGCTGGACTCTG TGCTCTATGGCCTGGTGACCTGGTACATGGAGGCCGTCTTCCCAGGGCAGTTCGGCGTGCCTCAGCCCTGGTA CTTCTTCATCATGCCCTCCTATTGGTGTGGGAAGCCAAGGGCGGTTGCAGGGAAGGAGGAAGAAGACAGTGA CCCCGAGAAAGCACTCAGAAACGAGTACTTTGAAGCCGAGCCAGAGGACCTGGTGGCGGGGATCAAGATCA AGCACCTGTCCAAGGTGTTCAGGGTGGGAAATAAGGACAGGGCGGCCGTCAGAGACCTGAACCTCAACCTGT ACGAGGGACAGATCACCGTCCTGCTGGGCCACAACGGTGCCGGGAAGACCACCACCCTCTCCATGCTCACAG GTCTCTTTCCCCCCACCAGTGGACGGGCATACATCAGCGGGTATGAAATTTCCCAGGACATGGTTCAGATCCG GAAGAGCCTGGGCCTGTGCCCGCAGCACGACATCCTGTTTGACAACTTGACAGTCGCAGAGCACCTTTATTTC TACGCCCAGCTGAAGGGCCTGTCACGTCAGAAGTGCCCTGAAGAAGTCAAGCAGATGCTGCACATCATCGGC CTGGAGGACAAGTGGAACTCACGGAGCCGCTTCCTGAGCGGGGGCATGAGGCGCAAGCTCTCCATCGGCATC GCCCTCATCGCAGGCTCCAAGGTGCTGATACTGGACGAGCCCACCTCGGGCATGGACGCCATCTCCAGGAGG GCCATCTGGGATCTTCTTCAGCGGCAGAAAAGTGACCGCACCATCGTGCTGACCACCCACTTCATGGACGAGG CTGACCTGCTGGGAGACCGCATCGCCATCATGGCCAAGGGGGAGCTGCAGTGCTGCGGGTCCTCGCTGTTCC TCAAGCAGAAATACGGTGCCGGCTATCACATGACGCTGGTGAAGGAGCCGCACTGCAACCCGGAAGACATCT CCCAGCTGGTCCACCACCACGTGCCCAACGCCACGCTGGAGAGCAGCGCTGGGGCCGAGCTGTCTTTCATCCT TCCCAGAGAGAGCACGCACAGGTTTGAAGGTCTCTTTGCTAAACTGGAGAAGAAGCAGAAAGAGCTGGGCAT TGCCAGCTTTGGGGCATCCATCACCACCATGGAGGAAGTCTTCCTTCGGGTCGGGAAGCTGGTGGACAGCAG TATGGACATCCAGGCCATCCAGCTCCCTGCCCTGCAGTACCAGCACGAGAGGCGCGCCAGCGACTGGGCTGT GGACAGCAACCTCTGTGGGGCCATGGACCCCTCCGACGGCATTGGAGCCCTCATCGAGGAGGAGCGCACCGC TGTCAAGCTCAACACTGGGCTCGCCCTGCACTGCCAGCAATTCTGGGCCATGTTCCTGAAGAAGGCCGCATAC AGCTGGCGCGAGTGGAAAATGGTGGCGGCACAGGTCCTGGTGCCTCTGACCTGCGTCACCCTGGCCCTCCTG GCCATCAACTACTCCTCGGAGCTCTTCGACGACCCCATGCTGAGGCTGACCTTGGGCGAGTACGGCAGAACCG TCGTGCCCTTCTCAGTTCCCGGGACCTCCCAGCTGGGTCAGCAGCTGTCAGAGCATCTGAAAGACGCACTGCA GGCTGAGGGACAGGAGCCCCGCGAGGTGCTCGGTGACCTGGAGGAGTTCTTGATCTTCAGGGCTTCTGTGGA GGGGGGCGGCTTTAATGAGCGGTGCCTTGTGGCAGCGTCCTTCAGAGATGTGGGAGAGCGCACGGTCGTCA ACGCCTTGTTCAACAACCAGGCGTACCACTCTCCAGCCACTGCCCTGGCCGTCGTGGACAACCTTCTGTTCAAG CTGCTGTGCGGGCCTCACGCCTCCATTGTGGTCTCCAACTTCCCCCAGCCCCGGAGCGCCCTGCAGGCTGCCA AGGACCAGTTTAACGAGGGCCGGAAGGGATTCGACATTGCCCTCAACCTGCTCTTCGCCATGGCATTCTTGGC
CAGCACGTTCTCCATCCTGGCGGTCAGCGAGAGGGCCGTGCAGGCCAAGCATGTGCAGTTTGTGAGTGGAGT CCACGTGGCCAGTTTCTGGCTCTCTGCTCTGCTGTGGGACCTCATCTCCTTCCTCATCCCCAGTCTGCTGCTGCT GGTGGTGTTTAAGGCCTTCGACGTGCGTGCCTTCACGCGGGACGGCCACATGGCTGACACCCTGCTGCTGCTC CTGCTCTACGGCTGGGCCATCATCCCCCTCATGTACCTGATGAACTTCTTCTTCTTGGGGGCGGCCACTGCCTA CACGAGGCTGACCATCTTCAACATCCTGTCAGGCATCGCCACCTTCCTGATGGTCACCATCATGCGCATCCCAG CTGTAAAACTGGAAGAACTTTCCAAAACCCTGGATCACGTGTTCCTGGTGCTGCCCAACCACTGTCTGGGGAT GGCAGTCAGCAGTTTCTACGAGAACTACGAGACGCGGAGGTACTGCACCTCCTCCGAGGTCGCCGCCCACTA CTGCAAGAAATATAACATCCAGTACCAGGAGAACTTCTATGCCTGGAGCGCCCCGGGGGTCGGCCGGTTTGT GGCCTCCATGGCCGCCTCAGGGTGCGCCTACCTCATCCTGCTCTTCCTCATCGAGACCAACCTGCTTCAGAGAC TCAGGGGCATCCTCTGCGCCCTCCGGAGGAGGCGGACACTGACAGAATTATACACCCGGATGCCTGTGCTTCC TGAGGACCAAGATGTAGCGGACGAGAGGACCCGCATCCTGGCCCCCAGCCCGGACTCCCTGCTCCACACACC TCTGATTATCAAGGAGCTCTCCAAGGTGTACGAGCAGCGGGTGCCCCTCCTGGCCGTGGACAGGCTCTCCCTC GCGGTGCAGAAAGGGGAGTGCTTCGGCCTGCTGGGCTTCAATGGAGCCGGGAAGACCACGACTTTCAAAAT GCTGACCGGGGAGGAGAGCCTCACTTCTGGGGATGCCTTTGTCGGGGGTCACAGAATCAGCTCTGATGTCGG AAAGGTGCGGCAGCGGATCGGCTACTGCCCGCAGTTTGATGCCTTGCTGGACCACATGACAGGCCGGGAGAT GCTGGTCATGTACGCTCGGCTCCGGGGCATCCCTGAGCGCCACATCGGGGCCTGCGTGGAGAACACTCTGCG GGGCCTGCTGCTGGAGCCACATGCCAACAAGCTGGTCAGGACGTACAGTGGTGGTAACAAGCGGAAGCTGA GCACCGGCATCGCCCTGATCGGAGAGCCTGCTGTCATCTTCCTGGACGAGCCGTCCACTGGCATGGACCCCGT GGCCCGGCGCCTGCTTTGGGACACCGTGGCACGAGCCCGAGAGTCTGGCAAGGCCATCATCATCACCTCCCA CAGCATGGAGGAGTGTGAGGCCCTGTGCACCCGGCTGGCCATCATGGTGCAGGGGCAGTTCAAGTGCCTGG GCAGCCCCCAGCACCTCAAGAGCAAGTTCGGCAGCGGCTACTCCCTGCGGGCCAAGGTGCAGAGTGAAGGG CAACAGGAGGCGCTGGAGGAGTTCAAGGCCTTCGTGGACCTGACCTTTCCAGGCAGCGTCCTGGAAGATGAG CACCAAGGCATGGTCCATTACCACCTGCCGGGCCGTGACCTCAGCTGGGCGAAGGTTTTCGGTATTCTGGAGA AAGCCAAGGAAAAGTACGGCGTGGACGACTACTCCGTGAGCCAGATCTCGCTGGAACAGGTCTTCCTGAGCT TCGCCCACCTGCAGCCGCCCACCGCAGAGGAGGGGCGATGAGCGGCCGCGGGCCC SEQ ID NO: 60 Exemplified codon‐optimised human ABACA3 (cohABCA3) transgene (5131bp) GCTAGCCACCATGGCCGTGCTGCGCCAGCTGGCCCTGCTGCTGTGGAAGAACTACACCCTGCAGAAGCGCAA GGTGCTGGTGACCGTGCTGGAGCTGTTCCTGCCCCTGCTGTTCAGCGGCATCCTGATCTGGCTGCGCCTGAAG ATCCAGAGCGAGAACGTGCCCAACGCCACCATCTACCCCGGCCAGAGCATCCAGGAGCTGCCCCTGTTCTTCA CCTTCCCCCCCCCCGGCGACACCTGGGAGCTGGCCTACATCCCCAGCCACAGCGACGCCGCCAAGACCGTGAC CGAGACCGTGCGCCGCGCCCTGGTGATCAACATGCGCGTGCGCGGCTTCCCCAGCGAGAAGGACTTCGAGGA CTACATCCGCTACGACAACTGCAGCAGCAGCGTGCTGGCCGCCGTGGTGTTCGAGCACCCCTTCAACCACAGC AAGGAGCCCCTGCCCCTGGCCGTGAAGTACCACCTGCGCTTCAGCTACACCCGCCGCAACTACATGTGGACCC
AGACCGGCAGCTTCTTCCTGAAGGAGACCGAGGGCTGGCACACCACCAGCCTGTTCCCCCTGTTCCCCAACCC CGGCCCCCGCGAGCCCACCAGCCCCGACGGCGGCGAGCCCGGCTACATCCGCGAGGGCTTCCTGGCCGTGCA GCACGCCGTGGACCGCGCCATCATGGAGTACCACGCCGACGCCGCCACCCGCCAGCTGTTCCAGCGCCTGAC CGTGACCATCAAGCGCTTCCCCTACCCCCCCTTCATCGCCGACCCCTTCCTGGTGGCCATCCAGTACCAGCTGC CCCTGCTGCTGCTGCTGAGCTTCACCTACACCGCCCTGACCATCGCCCGCGCCGTGGTGCAGGAGAAGGAGCG CCGCCTGAAGGAGTACATGCGCATGATGGGCCTGAGCAGCTGGCTGCACTGGAGCGCCTGGTTCCTGCTGTT CTTCCTGTTCCTGCTGATCGCCGCCAGCTTCATGACCCTGCTGTTCTGCGTGAAGGTGAAGCCCAACGTGGCCG TGCTGAGCCGCAGCGACCCCAGCCTGGTGCTGGCCTTCCTGCTGTGCTTCGCCATCAGCACCATCAGCTTCAG CTTCATGGTGAGCACCTTCTTCAGCAAGGCCAACATGGCCGCCGCCTTCGGCGGCTTCCTGTACTTCTTCACCT ACATCCCCTACTTCTTCGTGGCCCCCCGCTACAACTGGATGACCCTGAGCCAGAAGCTGTGCAGCTGCCTGCTG AGCAACGTGGCCATGGCCATGGGCGCCCAGCTGATCGGCAAGTTCGAGGCCAAGGGCATGGGCATCCAGTG GCGCGACCTGCTGAGCCCCGTGAACGTGGACGACGACTTCTGCTTCGGCCAGGTGCTGGGCATGCTGCTGCT GGACAGCGTGCTGTACGGCCTGGTGACCTGGTACATGGAGGCCGTGTTCCCCGGCCAGTTCGGCGTGCCCCA GCCCTGGTACTTCTTCATCATGCCCAGCTACTGGTGCGGCAAGCCCCGCGCCGTGGCCGGCAAGGAGGAGGA GGACAGCGACCCCGAGAAGGCCCTGCGCAACGAGTACTTCGAGGCCGAGCCCGAGGACCTGGTGGCCGGCA TCAAGATCAAGCACCTGAGCAAGGTGTTCCGCGTGGGCAACAAGGACCGCGCCGCCGTGCGCGACCTGAACC TGAACCTGTACGAGGGCCAGATCACCGTGCTGCTGGGCCACAACGGCGCCGGCAAGACCACCACCCTGAGCA TGCTGACCGGCCTGTTCCCCCCCACCAGCGGCAGGGCCTACATCAGCGGCTACGAGATCAGCCAGGACATGG TGCAGATCCGCAAGAGCCTGGGCCTGTGCCCCCAGCACGACATCCTGTTCGACAACCTGACCGTGGCCGAGC ACCTGTACTTCTACGCCCAGCTGAAGGGCCTGAGCCGCCAGAAGTGCCCCGAGGAGGTGAAGCAGATGCTGC ACATCATCGGCCTGGAGGACAAGTGGAACAGCCGCAGCCGCTTCCTGAGCGGCGGCATGCGCCGCAAGCTG AGCATCGGCATCGCCCTGATCGCCGGCAGCAAGGTGCTGATCCTGGACGAGCCCACCAGCGGCATGGACGCC ATCAGCCGCCGCGCCATCTGGGACCTGCTGCAGCGCCAGAAGAGCGACCGCACCATCGTGCTGACCACCCAC TTCATGGACGAGGCCGACCTGCTGGGCGACCGCATCGCCATCATGGCCAAGGGCGAGCTGCAGTGCTGCGGC AGCAGCCTGTTCCTGAAGCAGAAGTACGGCGCCGGCTACCACATGACCCTGGTGAAGGAGCCCCACTGCAAC CCCGAGGACATCAGCCAGCTGGTGCACCACCACGTGCCCAACGCCACCCTGGAGAGCAGCGCCGGCGCCGAG CTGAGCTTCATCCTGCCCCGCGAGAGCACCCACCGCTTCGAGGGCCTGTTCGCCAAGCTGGAGAAGAAGCAG AAGGAGCTGGGCATCGCCAGCTTCGGCGCCAGCATCACCACCATGGAGGAGGTGTTCCTGCGCGTGGGCAA GCTGGTGGACAGCAGCATGGACATCCAGGCCATCCAGCTGCCCGCCCTGCAGTACCAGCACGAGCGCCGCGC CAGCGACTGGGCCGTGGACAGCAACCTGTGCGGCGCCATGGACCCCAGCGACGGCATCGGCGCCCTGATCG AGGAGGAGCGCACCGCCGTGAAGCTGAACACCGGCCTGGCCCTGCACTGCCAGCAGTTCTGGGCCATGTTCC TGAAGAAGGCCGCCTACAGCTGGCGCGAGTGGAAGATGGTGGCCGCCCAGGTGCTGGTGCCCCTGACCTGC GTGACCCTGGCCCTGCTGGCCATCAACTACAGCAGCGAGCTGTTCGACGACCCCATGCTGCGCCTGACCCTGG GCGAGTACGGCCGCACCGTGGTGCCCTTCAGCGTGCCCGGCACCAGCCAGCTGGGCCAGCAGCTGAGCGAG
CACCTGAAGGACGCCCTGCAGGCCGAGGGCCAGGAGCCCCGCGAGGTGCTGGGCGACCTGGAGGAGTTCCT GATCTTCCGCGCCAGCGTGGAGGGCGGCGGCTTCAACGAGCGCTGCCTGGTGGCCGCCAGCTTCCGCGACGT GGGCGAGCGCACCGTGGTGAACGCCCTGTTCAACAACCAGGCCTACCACAGCCCCGCCACCGCCCTGGCCGT GGTGGACAACCTGCTGTTCAAGCTGCTGTGCGGCCCCCACGCCAGCATCGTGGTGAGCAACTTCCCCCAGCCC CGCAGCGCCCTGCAGGCCGCCAAGGACCAGTTCAACGAGGGCCGCAAGGGCTTCGACATCGCCCTGAACCTG CTGTTCGCCATGGCCTTCCTGGCCAGCACCTTCAGCATCCTGGCCGTGAGCGAGAGGGCCGTGCAGGCCAAG CACGTGCAGTTCGTGAGCGGCGTGCACGTGGCCAGCTTCTGGCTGAGCGCCCTGCTGTGGGACCTGATCAGC TTCCTGATCCCCAGCCTGCTGCTGCTGGTGGTGTTCAAGGCCTTCGACGTGAGGGCCTTCACCCGCGACGGCC ACATGGCCGACACCCTGCTGCTGCTGCTGCTGTACGGCTGGGCCATCATCCCCCTGATGTACCTGATGAACTTC TTCTTCCTGGGCGCCGCCACCGCCTACACCCGCCTGACCATCTTCAACATCCTGAGCGGCATCGCCACCTTCCT GATGGTGACCATCATGCGCATCCCCGCCGTGAAGCTGGAGGAGCTGAGCAAGACCCTGGACCACGTGTTCCT GGTGCTGCCCAACCACTGCCTGGGCATGGCCGTGAGCAGCTTCTACGAGAACTACGAGACCCGCCGCTACTG CACCAGCAGCGAGGTGGCCGCCCACTACTGCAAGAAGTACAACATCCAGTACCAGGAGAACTTCTACGCCTG GAGCGCCCCCGGCGTGGGCCGCTTCGTGGCCAGCATGGCCGCCAGCGGCTGCGCCTACCTGATCCTGCTGTT CCTGATCGAGACCAACCTGCTGCAGCGCCTGCGCGGCATCCTGTGCGCCCTGCGCCGCCGCCGCACCCTGACC GAGCTGTACACCCGCATGCCCGTGCTGCCCGAGGACCAGGACGTGGCCGACGAGCGCACCCGCATCCTGGCC CCCAGCCCCGACAGCCTGCTGCACACCCCCCTGATCATCAAGGAGCTGAGCAAGGTGTACGAGCAGCGCGTG CCCCTGCTGGCCGTGGACCGCCTGAGCCTGGCCGTGCAGAAGGGCGAGTGCTTCGGCCTGCTGGGCTTCAAC GGCGCCGGCAAGACCACCACCTTCAAGATGCTGACCGGCGAGGAGAGCCTGACCAGCGGCGACGCCTTCGT GGGCGGCCACCGCATCAGCAGCGACGTGGGCAAGGTGCGCCAGCGCATCGGCTACTGCCCCCAGTTCGACGC CCTGCTGGACCACATGACCGGCCGCGAGATGCTGGTGATGTACGCCCGCCTGCGCGGCATCCCCGAGCGCCA CATCGGCGCCTGCGTGGAGAACACCCTGCGCGGCCTGCTGCTGGAGCCCCACGCCAACAAGCTGGTGCGCAC CTACAGCGGCGGCAACAAGCGCAAGCTGAGCACCGGCATCGCCCTGATCGGCGAGCCCGCCGTGATCTTCCT GGACGAGCCCAGCACCGGCATGGACCCCGTGGCCCGCCGCCTGCTGTGGGACACCGTGGCCCGCGCCCGCG AGAGCGGCAAGGCCATCATCATCACCAGCCACAGCATGGAGGAGTGCGAGGCCCTGTGCACCCGCCTGGCCA TCATGGTGCAGGGCCAGTTCAAGTGCCTGGGCAGCCCCCAGCACCTGAAGAGCAAGTTCGGCAGCGGCTACA GCCTGAGGGCCAAGGTGCAGAGCGAGGGCCAGCAGGAGGCCCTGGAGGAGTTCAAGGCCTTCGTGGACCTG ACCTTCCCCGGCAGCGTGCTGGAGGACGAGCACCAGGGCATGGTGCACTACCACCTGCCCGGCCGCGACCTG AGCTGGGCCAAGGTGTTCGGCATCCTGGAGAAGGCCAAGGAGAAGTACGGCGTGGACGACTACAGCGTGAG CCAGATCAGCCTGGAGCAGGTGTTCCTGAGCTTCGCCCACCTGCAGCCCCCCACCGCCGAGGAGGGCCGCTA AGGGCCC SEQ ID NO: 61 Exemplified Human ABCA3 polypeptide
MAVLRQLALLLWKNYTLQKRKVLVTVLELFLPLLFSGILIWLRLKIQSENVPNATIYPGQSIQELPLFFTFPPPGDTWE LAYIPSHSDAAKTVTETVRRALVINMRVRGFPSEKDFEDYIRYDNCSSSVLAAVVFEHPFNHSKEPLPLAVKYHLRFSY TRRNYMWTQTGSFFLKETEGWHTTSLFPLFPNPGPREPTSPDGGEPGYIREGFLAVQHAVDRAIMEYHADAATR QLFQRLTVTIKRFPYPPFIADPFLVAIQYQLPLLLLLSFTYTALTIARAVVQEKERRLKEYMRMMGLSSWLHWSAWFL LFFLFLLIAASFMTLLFCVKVKPNVAVLSRSDPSLVLAFLLCFAISTISFSFMVSTFFSKANMAAAFGGFLYFFTYIPYFF VAPRYNWMTLSQKLCSCLLSNVAMAMGAQLIGKFEAKGMGIQWRDLLSPVNVDDDFCFGQVLGMLLLDSVLYG LVTWYMEAVFPGQFGVPQPWYFFIMPSYWCGKPRAVAGKEEEDSDPEKALRNEYFEAEPEDLVAGIKIKHLSKVF RVGNKDRAAVRDLNLNLYEGQITVLLGHNGAGKTTTLSMLTGLFPPTSGRAYISGYEISQDMVQIRKSLGLCPQHDI LFDNLTVAEHLYFYAQLKGLSRQKCPEEVKQMLHIIGLEDKWNSRSRFLSGGMRRKLSIGIALIAGSKVLILDEPTSG MDAISRRAIWDLLQRQKSDRTIVLTTHFMDEADLLGDRIAIMAKGELQCCGSSLFLKQKYGAGYHMTLVKEPHCN PEDISQLVHHHVPNATLESSAGAELSFILPRESTHRFEGLFAKLEKKQKELGIASFGASITTMEEVFLRVGKLVDSSMD IQAIQLPALQYQHERRASDWAVDSNLCGAMDPSDGIGALIEEERTAVKLNTGLALHCQQFWAMFLKKAAYSWRE WKMVAAQVLVPLTCVTLALLAINYSSELFDDPMLRLTLGEYGRTVVPFSVPGTSQLGQQLSEHLKDALQAEGQEPR EVLGDLEEFLIFRASVEGGGFNERCLVAASFRDVGERTVVNALFNNQAYHSPATALAVVDNLLFKLLCGPHASIVVS NFPQPRSALQAAKDQFNEGRKGFDIALNLLFAMAFLASTFSILAVSERAVQAKHVQFVSGVHVASFWLSALLWDLI SFLIPSLLLLVVFKAFDVRAFTRDGHMADTLLLLLLYGWAIIPLMYLMNFFFLGAATAYTRLTIFNILSGIATFLMVTIM RIPAVKLEELSKTLDHVFLVLPNHCLGMAVSSFYENYETRRYCTSSEVAAHYCKKYNIQYQENFYAWSAPGVGRFVA SMAASGCAYLILLFLIETNLLQRLRGILCALRRRRTLTELYTRMPVLPEDQDVADERTRILAPSPDSLLHTPLIIKELSKV YEQRVPLLAVDRLSLAVQKGECFGLLGFNGAGKTTTFKMLTGEESLTSGDAFVGGHRISSDVGKVRQRIGYCPQFD ALLDHMTGREMLVMYARLRGIPERHIGACVENTLRGLLLEPHANKLVRTYSGGNKRKLSTGIALIGEPAVIFLDEPST GMDPVARRLLWDTVARARESGKAIIITSHSMEECEALCTRLAIMVQGQFKCLGSPQHLKSKFGSGYSLRAKVQSEG QQEALEEFKAFVDLTFPGSVLEDEHQGMVHYHLPGRDLSWAKVFGILEKAKEKYGVDDYSVSQISLEQVFLSFAHL QPPTAEEGR* SEQ ID NO: 62 Exemplified Human SFTPC transgene GCTAGCCACCATGGATGTGGGCAGCAAAGAGGTCCTGATGGAGAGCCCGCCGGACTACTCCGCAGCTCCCCG GGGCCGATTTGGCATTCCCTGCTGCCCAGTGCACCTGAAACGCCTTCTTATCGTGGTGGTGGTGGTGGTCCTC ATCGTCGTGGTGATTGTGGGAGCCCTGCTCATGGGTCTCCACATGAGCCAGAAACACACGGAGATGGTTCTG GAGATGAGCATTGGGGCGCCGGAAGCCCAGCAACGCCTGGCCCTGAGTGAGCACCTGGTTACCACTGCCACC TTCTCCATCGGCTCCACTGGCCTCGTGGTGTATGACTACCAGCAGCTGCTGATCGCCTACAAGCCAGCCCCTGG CACCTGCTGCTACATCATGAAGATAGCTCCAGAGAGCATCCCCAGTCTTGAGGCTCTCACTAGAAAAGTCCAC AACTTCCAGGCCAAGCCCGCAGTGCCTACGTCTAAGCTGGGCCAGGCAGAGGGGCGAGATGCAGGCTCAGC ACCCTCCGGAGGGGACCCGGCCTTCCTGGGCATGGCCGTGAGCACCCTGTGTGGCGAGGTGCCGCTCTACTA CATCTAGGGGCCC
Underlined sequence = combined NheI and Kozak sequence Bold and dashed‐underlined sequence = ApaI site Either or both of the underlined and/or bold+dashed‐underlined sequences may be omitted and such a sequence still expressly falls within the scope of the invention (SEQ ID NOs: 84‐86). SEQ ID NO: 63 Exemplified Human SFTPC polypeptide MDVGSKEVLMESPPDYSAAPRGRFGIPCCPVHLKRLLIVVVVVVLIVVVIVGALLMGLHMSQKHTEMVLEMSIGAP EAQQRLALSEHLVTTATFSIGSTGLVVYDYQQLLIAYKPAPGTCCYIMKIAPESIPSLEALTRKVHNFQAKPAVPTSKL GQAEGRDAGSAPSGGDPAFLGMAVSTLCGEVPLYYI* SEQ ID NO: 64 Exemplified Human GM‐CSF (CSF2) transgene GCTAGCCACCATGTGGCTGCAGAGCCTGCTGCTCTTGGGCACTGTGGCCTGCAGCATCTCTGCACCCGCCCGC TCGCCCAGCCCCAGCACGCAGCCCTGGGAGCATGTGAATGCCATCCAGGAGGCCCGGCGTCTCCTGAACCTG AGTAGAGACACTGCTGCTGAGATGAATGAAACAGTAGAAGTCATCTCAGAAATGTTTGACCTCCAGGAGCCG ACCTGCCTACAGACCCGCCTGGAGCTGTACAAGCAGGGCCTGCGGGGCAGCCTCACCAAGCTCAAGGGCCCC TTGACCATGATGGCCAGCCACTACAAGCAGCACTGCCCTCCAACCCCGGAAACTTCCTGTGCAACCCAGATTAT CACCTTTGAAAGTTTCAAAGAGAACCTGAAGGACTTTCTGCTTGTCATCCCCTTTGACTGCTGGGAGCCAGTCC AGGAGTGAGGGCCC SEQ ID NO: 65 Exemplified Human GM‐CSF polypeptide MWLQSLLLLGTVACSISAPARSPSPSTQPWEHVNAIQEARRLLNLSRDTAAEMNETVEVISEMFDLQEPTCLQTRL ELYKQGLRGSLTKLKGPLTMMASHYKQHCPPTPETSCATQIITFESFKENLKDFLLVIPFDCWEPVQE* SEQ ID NO: 66 Exemplified Human DCN (Decorin) transgene GCTAGCCACCATGAAGGCCACTATCATCCTCCTTCTGCTTGCACAAGTTTCCTGGGCTGGACCGTTTCAACAGA GAGGCTTATTTGACTTTATGCTAGAAGATGAGGCTTCTGGGATAGGCCCAGAAGTTCCTGATGACCGCGACTT CGAGCCCTCCCTAGGCCCAGTGTGCCCCTTCCGCTGTCAATGCCATCTTCGAGTGGTCCAGTGTTCTGATTTGG GTCTGGACAAAGTGCCAAAGGATCTTCCCCCTGACACAACTCTGCTAGACCTGCAAAACAACAAAATAACCGA AATCAAAGATGGAGACTTTAAGAACCTGAAGAACCTTCACGCATTGATTCTTGTCAACAATAAAATTAGCAAA GTTAGTCCTGGAGCATTTACACCTTTGGTGAAGTTGGAACGACTTTATCTGTCCAAGAATCAGCTGAAGGAAT TGCCAGAAAAAATGCCCAAAACTCTTCAGGAGCTGCGTGCCCATGAGAATGAGATCACCAAAGTGCGAAAAG TTACTTTCAATGGACTGAACCAGATGATTGTCATAGAACTGGGCACCAATCCGCTGAAGAGCTCAGGAATTGA AAATGGGGCTTTCCAGGGAATGAAGAAGCTCTCCTACATCCGCATTGCTGATACCAATATCACCAGCATTCCTC
AAGGTCTTCCTCCTTCCCTTACGGAATTACATCTTGATGGCAACAAAATCAGCAGAGTTGATGCAGCAAGCCTG AAAGGACTGAATAATTTGGCTAAGTTGGGATTGAGTTTCAACAGCATCTCTGCTGTTGACAATGGCTCTCTGG CCAACACGCCTCATCTGAGGGAGCTTCACTTGGACAACAACAAGCTTACCAGAGTACCTGGTGGGCTGGCAG AGCATAAGTACATCCAGGTTGTCTACCTTCATAACAACAATATCTCTGTAGTTGGATCAAGTGACTTCTGCCCA CCTGGACACAACACCAAAAAGGCTTCTTATTCGGGTGTGAGTCTTTTCAGCAACCCGGTCCAGTACTGGGAGA TACAGCCATCCACCTTCAGATGTGTCTACGTGCGCTCTGCCATTCAACTCGGAAACTATAAGTAAGGGCCCA SEQ ID NO: 67 Exemplified Human Decorin polypeptide MKATIILLLLAQVSWAGPFQQRGLFDFMLEDEASGIGPEVPDDRDFEPSLGPVCPFRCQCHLRVVQCSDLGLDKVP KDLPPDTTLLDLQNNKITEIKDGDFKNLKNLHALILVNNKISKVSPGAFTPLVKLERLYLSKNQLKELPEKMPKTLQEL RAHENEITKVRKVTFNGLNQMIVIELGTNPLKSSGIENGAFQGMKKLSYIRIADTNITSIPQGLPPSLTELHLDGNKIS RVDAASLKGLNNLAKLGLSFNSISAVDNGSLANTPHLRELHLDNNKLTRVPGGLAEHKYIQVVYLHNNNISVVGSSD FCPPGHNTKKASYSGVSLFSNPVQYWEIQPSTFRCVYVRSAIQLGNYK* SEQ ID NO: 68 Exemplified Human TRIM72 transgene GCTAGCCACCATGTCGGCTGCGCCCGGCCTCCTGCACCAGGAGCTGTCCTGCCCGCTGTGCCTGCAGCTGTTC GACGCGCCCGTGACAGCCGAGTGCGGCCACAGTTTCTGCCGCGCCTGCCTAGGCCGCGTGGCCGGGGAGCC GGCGGCGGATGGCACCGTTCTCTGCCCCTGCTGCCAGGCCCCCACGCGGCCGCAGGCACTCAGCACCAACCT GCAGCTGGCGCGCCTGGTGGAGGGGCTGGCCCAGGTGCCGCAGGGCCACTGCGAGGAGCACCTGGACCCGC TGAGCATCTACTGCGAGCAGGACCGCGCGCTGGTGTGCGGAGTGTGCGCCTCACTCGGCTCGCACCGCGGTC ATCGCCTCCTGCCTGCCGCCGAGGCCCACGCACGCCTCAAGACACAGCTGCCACAGCAGAAACTGCAGCTGCA GGAGGCATGCATGCGCAAGGAGAAGAGTGTGGCTGTGCTGGAGCATCAGCTGGTGGAGGTGGAGGAGACA GTGCGTCAGTTCCGGGGGGCCGTGGGGGAGCAGCTGGGCAAGATGCGGGTGTTCCTGGCTGCACTGGAGG GCTCCTTGGACCGCGAGGCAGAGCGTGTACGGGGTGAGGCAGGGGTCGCCTTGCGCCGGGAGCTGGGGAG CCTGAACTCTTACCTGGAGCAGCTGCGGCAGATGGAGAAGGTCCTGGAGGAGGTGGCGGACAAGCCGCAGA CTGAGTTCCTCATGAAATACTGCCTGGTGACCAGCAGGCTGCAGAAGATCCTGGCAGAGTCTCCCCCACCCGC CCGTCTGGACATCCAGCTGCCAATTATCTCAGATGACTTCAAATTCCAGGTGTGGAGGAAGATGTTCCGGGCT CTGATGCCAGCGCTGGAGGAGCTGACCTTTGACCCGAGCTCTGCGCACCCGAGCCTGGTGGTGTCTTCCTCTG GCCGCCGCGTGGAGTGCTCGGAGCAGAAGGCGCCGCCGGCCGGGGAGGACCCGCGCCAGTTCGACAAGGC GGTGGCGGTGGTGGCGCACCAGCAGCTCTCCGAGGGCGAGCACTACTGGGAGGTGGATGTTGGCGACAAGC CGCGCTGGGCGCTGGGCGTGATCGCGGCCGAGGCCCCCCGCCGCGGGCGCCTGCACGCGGTGCCCTCGCAG GGCCTGTGGCTGCTGGGGCTGCGCGAGGGCAAGATCCTGGAGGCACACGTGGAGGCCAAGGAGCCGCGCG CTCTGCGCAGCCCCGAGAGGCGGCCCACGCGCATTGGCCTTTACCTGAGCTTCGGCGACGGCGTCCTCTCCTT CTACGATGCCAGCGACGCCGACGCGCTCGTGCCGCTTTTTGCCTTCCACGAGCGCCTGCCCAGGCCCGTGTAC
CCCTTCTTCGACGTGTGCTGGCACGACAAGGGCAAGAATGCCCAGCCGCTGCTGCTCGTGGGTCCCGAAGGC GCCGAGGCCTGAGGGCCC SEQ ID NO: 69 Exemplified Human TRIM72 polypeptide MSAAPGLLHQELSCPLCLQLFDAPVTAECGHSFCRACLGRVAGEPAADGTVLCPCCQAPTRPQALSTNLQLARLVE GLAQVPQGHCEEHLDPLSIYCEQDRALVCGVCASLGSHRGHRLLPAAEAHARLKTQLPQQKLQLQEACMRKEKSV AVLEHQLVEVEETVRQFRGAVGEQLGKMRVFLAALEGSLDREAERVRGEAGVALRRELGSLNSYLEQLRQMEKVL EEVADKPQTEFLMKYCLVTSRLQKILAESPPPARLDIQLPIISDDFKFQVWRKMFRALMPALEELTFDPSSAHPSLVV SSSGRRVECSEQKAPPAGEDPRQFDKAVAVVAHQQLSEGEHYWEVDVGDKPRWALGVIAAEAPRRGRLHAVPS QGLWLLGLREGKILEAHVEAKEPRALRSPERRPTRIGLYLSFGDGVLSFYDASDADALVPLFAFHERLPRPVYPFFDV CWHDKGKNAQPLLLVGPEGAEA SEQ ID NO: 70 Exemplified SERPINA1 (AAT) transgene ATGCCCAGCTCTGTGTCCTGGGGCATTCTGCTGCTGGCTGGCCTGTGCTGTCTGGTGCCTGTGTCCCTGG CTGAGGACCCTCAGGGGGATGCTGCCCAGAAAACAGACACCTCCCACCATGACCAGGACCACCCCACCTT CAACAAGATCACCCCCAACCTGGCAGAGTTTGCCTTCAGCCTGTACAGACAGCTGGCCCACCAGAGCAAC AGCACCAACATCTTTTTCAGCCCTGTGTCCATTGCCACAGCCTTTGCCATGCTGAGCCTGGGCACCAAGG CTGACACCCATGATGAGATCCTGGAAGGCCTGAACTTCAACCTGACAGAGATCCCTGAGGCCCAGATCCA TGAGGGCTTCCAGGAACTGCTGAGAACCCTGAACCAGCCAGACAGCCAGCTGCAGCTGACAACAGGCAAT GGGCTGTTCCTGTCTGAGGGCCTGAAGCTGGTGGACAAGTTTCTGGAAGATGTGAAGAAGCTGTACCACT CTGAGGCCTTCACAGTGAACTTTGGGGACACAGAAGAGGCCAAGAAACAGATCAATGACTATGTGGAAAA GGGCACCCAGGGCAAGATTGTGGACCTTGTGAAAGAGCTGGACAGGGACACTGTGTTTGCCCTTGTGAAC TACATCTTCTTCAAGGGCAAGTGGGAGAGGCCCTTTGAAGTGAAGGACACTGAGGAAGAGGACTTCCATG TGGACCAAGTGACCACAGTGAAGGTGCCAATGATGAAGAGACTGGGGATGTTCAATATCCAGCACTGCAA GAAACTGAGCAGCTGGGTGCTGCTGATGAAGTACCTGGGCAATGCTACAGCCATATTCTTTCTGCCTGAT GAGGGCAAGCTGCAGCACCTGGAAAATGAGCTGACCCATGACATCATCACCAAATTTCTGGAAAATGAGG ACAGAAGATCTGCCAGCCTGCATCTGCCCAAGCTGAGCATCACAGGCACATATGACCTGAAGTCTGTGCT GGGACAGCTGGGAATCACCAAGGTGTTCAGCAATGGGGCAGACCTGAGTGGAGTGACAGAGGAAGCCCCT CTGAAGCTGTCCAAGGCTGTGCACAAGGCAGTGCTGACCATTGATGAGAAGGGCACAGAGGCTGCTGGGG CCATGTTTCTGGAAGCCATCCCCATGTCCATCCCCCCAGAAGTGAAGTTCAACAAGCCCTTTGTGTTCCT GATGATTGAGCAGAACACCAAGAGCCCCCTGTTCATGGGCAAGGTTGTGAACCCCACCCAGAAATGA SEQ ID NO: 71 Exemplified AAT polypeptide MPSSVSWGILLLAGLCCLVPVSLAEDPQGDAAQKTDTSHHDQDHPTFAEDPQGDAAQKTDTSHHDQDH PTFNKITPNLAEFAFSLYRQLAHQSNSTNIFFSPVSIATAFAMLSLGTKADTHDEILEGLNFNLTEIP EAQIHEGFQELLRTLNQPDSQLQLTTGNGLFLSEGLKLVDKFLEDVKKLYHSEAFTVNFGDTEEAKKQ INDYVEKGTQGKIVDLVKELDRDTVFALVNYIFFKGKWERPFEVKDTEEEDFHVDQVTTVKVPMMKRL GMFNIQHCKKLSSWVLLMKYLGNATAIFFLPDEGKLQHLENELTHDIITKFLENEDRRSASLHLPKLS
ITGTYDLKSVLGQLGITKVFSNGADLSGVTEEAPLKLSKAVHKAVLTIDEKGTEAAGAMFLEAIPMSI PPEVKFNKPFVFLMIEQNTKSPLFMGKVVNPTQK SEQ ID NO: 72 Exemplified FVIII transgene (N6) ATGCAGATTGAGCTGAGCACCTGCTTCTTCCTGTGCCTGCTGAGGTTCTGCTTCTCTGCCACCAGGAGAT ACTACCTGGGGGCTGTGGAGCTGAGCTGGGACTACATGCAGTCTGACCTGGGGGAGCTGCCTGTGGATGC CAGGTTCCCCCCCAGAGTGCCCAAGAGCTTCCCCTTCAACACCTCTGTGGTGTACAAGAAGACCCTGTTT GTGGAGTTCACTGACCACCTGTTCAACATTGCCAAGCCCAGGCCCCCCTGGATGGGCCTGCTGGGCCCCA CCATCCAGGCTGAGGTGTATGACACTGTGGTGATCACCCTGAAGAACATGGCCAGCCACCCTGTGAGCCT GCATGCTGTGGGGGTGAGCTACTGGAAGGCCTCTGAGGGGGCTGAGTATGATGACCAGACCAGCCAGAGG GAGAAGGAGGATGACAAGGTGTTCCCTGGGGGCAGCCACACCTATGTGTGGCAGGTGCTGAAGGAGAATG GCCCCATGGCCTCTGACCCCCTGTGCCTGACCTACAGCTACCTGAGCCATGTGGACCTGGTGAAGGACCT GAACTCTGGCCTGATTGGGGCCCTGCTGGTGTGCAGGGAGGGCAGCCTGGCCAAGGAGAAGACCCAGACC CTGCACAAGTTCATCCTGCTGTTTGCTGTGTTTGATGAGGGCAAGAGCTGGCACTCTGAAACCAAGAACA GCCTGATGCAGGACAGGGATGCTGCCTCTGCCAGGGCCTGGCCCAAGATGCACACTGTGAATGGCTATGT GAACAGGAGCCTGCCTGGCCTGATTGGCTGCCACAGGAAGTCTGTGTACTGGCATGTGATTGGCATGGGC ACCACCCCTGAGGTGCACAGCATCTTCCTGGAGGGCCACACCTTCCTGGTCAGGAACCACAGGCAGGCCA GCCTGGAGATCAGCCCCATCACCTTCCTGACTGCCCAGACCCTGCTGATGGACCTGGGCCAGTTCCTGCT GTTCTGCCACATCAGCAGCCACCAGCATGATGGCATGGAGGCCTATGTGAAGGTGGACAGCTGCCCTGAG GAGCCCCAGCTGAGGATGAAGAACAATGAGGAGGCTGAGGACTATGATGATGACCTGACTGACTCTGAGA TGGATGTGGTGAGGTTTGATGATGACAACAGCCCCAGCTTCATCCAGATCAGGTCTGTGGCCAAGAAGCA CCCCAAGACCTGGGTGCACTACATTGCTGCTGAGGAGGAGGACTGGGACTATGCCCCCCTGGTGCTGGCC CCTGATGACAGGAGCTACAAGAGCCAGTACCTGAACAATGGCCCCCAGAGGATTGGCAGGAAGTACAAGA AGGTCAGGTTCATGGCCTACACTGATGAAACCTTCAAGACCAGGGAGGCCATCCAGCATGAGTCTGGCAT CCTGGGCCCCCTGCTGTATGGGGAGGTGGGGGACACCCTGCTGATCATCTTCAAGAACCAGGCCAGCAGG CCCTACAACATCTACCCCCATGGCATCACTGATGTGAGGCCCCTGTACAGCAGGAGGCTGCCCAAGGGGG TGAAGCACCTGAAGGACTTCCCCATCCTGCCTGGGGAGATCTTCAAGTACAAGTGGACTGTGACTGTGGA GGATGGCCCCACCAAGTCTGACCCCAGGTGCCTGACCAGATACTACAGCAGCTTTGTGAACATGGAGAGG GACCTGGCCTCTGGCCTGATTGGCCCCCTGCTGATCTGCTACAAGGAGTCTGTGGACCAGAGGGGCAACC AGATCATGTCTGACAAGAGGAATGTGATCCTGTTCTCTGTGTTTGATGAGAACAGGAGCTGGTACCTGAC TGAGAACATCCAGAGGTTCCTGCCCAACCCTGCTGGGGTGCAGCTGGAGGACCCTGAGTTCCAGGCCAGC AACATCATGCACAGCATCAATGGCTATGTGTTTGACAGCCTGCAGCTGTCTGTGTGCCTGCATGAGGTGG CCTACTGGTACATCCTGAGCATTGGGGCCCAGACTGACTTCCTGTCTGTGTTCTTCTCTGGCTACACCTT CAAGCACAAGATGGTGTATGAGGACACCCTGACCCTGTTCCCCTTCTCTGGGGAGACTGTGTTCATGAGC ATGGAGAACCCTGGCCTGTGGATTCTGGGCTGCCACAACTCTGACTTCAGGAACAGGGGCATGACTGCCC TGCTGAAAGTCTCCAGCTGTGACAAGAACACTGGGGACTACTATGAGGACAGCTATGAGGACATCTCTGC CTACCTGCTGAGCAAGAACAATGCCATTGAGCCCAGGAGCTTCAGCCAGAACAGCAGGCACCCCAGCACC AGGCAGAAGCAGTTCAATGCCACCACCATCCCTGAGAATGACATAGAGAAGACAGACCCATGGTTTGCCC ACCGGACCCCCATGCCCAAGATCCAGAATGTGAGCAGCTCTGACCTGCTGATGCTGCTGAGGCAGAGCCC CACCCCCCATGGCCTGAGCCTGTCTGACCTGCAGGAGGCCAAGTATGAAACCTTCTCTGATGACCCCAGC
CCTGGGGCCATTGACAGCAACAACAGCCTGTCTGAGATGACCCACTTCAGGCCCCAGCTGCACCACTCTG GGGACATGGTGTTCACCCCTGAGTCTGGCCTGCAGCTGAGGCTGAATGAGAAGCTGGGCACCACTGCTGC CACTGAGCTGAAGAAGCTGGACTTCAAAGTCTCCAGCACCAGCAACAACCTGATCAGCACCATCCCCTCT GACAACCTGGCTGCTGGCACTGACAACACCAGCAGCCTGGGCCCCCCCAGCATGCCTGTGCACTATGACA GCCAGCTGGACACCACCCTGTTTGGCAAGAAGAGCAGCCCCCTGACTGAGTCTGGGGGCCCCCTGAGCCT GTCTGAGGAGAACAATGACAGCAAGCTGCTGGAGTCTGGCCTGATGAACAGCCAGGAGAGCAGCTGGGGC AAGAATGTGAGCAGCAGGGAGATCACCAGGACCACCCTGCAGTCTGACCAGGAGGAGATTGACTATGATG ACACCATCTCTGTGGAGATGAAGAAGGAGGACTTTGACATCTACGACGAGGACGAGAACCAGAGCCCCAG GAGCTTCCAGAAGAAGACCAGGCACTACTTCATTGCTGCTGTGGAGAGGCTGTGGGACTATGGCATGAGC AGCAGCCCCCATGTGCTGAGGAACAGGGCCCAGTCTGGCTCTGTGCCCCAGTTCAAGAAGGTGGTGTTCC AGGAGTTCACTGATGGCAGCTTCACCCAGCCCCTGTACAGAGGGGAGCTGAATGAGCACCTGGGCCTGCT GGGCCCCTACATCAGGGCTGAGGTGGAGGACAACATCATGGTGACCTTCAGGAACCAGGCCAGCAGGCCC TACAGCTTCTACAGCAGCCTGATCAGCTATGAGGAGGACCAGAGGCAGGGGGCTGAGCCCAGGAAGAACT TTGTGAAGCCCAATGAAACCAAGACCTACTTCTGGAAGGTGCAGCACCACATGGCCCCCACCAAGGATGA GTTTGACTGCAAGGCCTGGGCCTACTTCTCTGATGTGGACCTGGAGAAGGATGTGCACTCTGGCCTGATT GGCCCCCTGCTGGTGTGCCACACCAACACCCTGAACCCTGCCCATGGCAGGCAGGTGACTGTGCAGGAGT TTGCCCTGTTCTTCACCATCTTTGATGAAACCAAGAGCTGGTACTTCACTGAGAACATGGAGAGGAACTG CAGGGCCCCCTGCAACATCCAGATGGAGGACCCCACCTTCAAGGAGAACTACAGGTTCCATGCCATCAAT GGCTACATCATGGACACCCTGCCTGGCCTGGTGATGGCCCAGGACCAGAGGATCAGGTGGTACCTGCTGA GCATGGGCAGCAATGAGAACATCCACAGCATCCACTTCTCTGGCCATGTGTTCACTGTGAGGAAGAAGGA GGAGTACAAGATGGCCCTGTACAACCTGTACCCTGGGGTGTTTGAGACTGTGGAGATGCTGCCCAGCAAG GCTGGCATCTGGAGGGTGGAGTGCCTGATTGGGGAGCACCTGCATGCTGGCATGAGCACCCTGTTCCTGG TGTACAGCAACAAGTGCCAGACCCCCCTGGGCATGGCCTCTGGCCACATCAGGGACTTCCAGATCACTGC CTCTGGCCAGTATGGCCAGTGGGCCCCCAAGCTGGCCAGGCTGCACTACTCTGGCAGCATCAATGCCTGG AGCACCAAGGAGCCCTTCAGCTGGATCAAGGTGGACCTGCTGGCCCCCATGATCATCCATGGCATCAAGA CCCAGGGGGCCAGGCAGAAGTTCAGCAGCCTGTACATCAGCCAGTTCATCATCATGTACAGCCTGGATGG CAAGAAGTGGCAGACCTACAGGGGCAACAGCACTGGCACCCTGATGGTGTTCTTTGGCAATGTGGACAGC TCTGGCATCAAGCACAACATCTTCAACCCCCCCATCATTGCCAGATACATCAGGCTGCACCCCACCCACT ACAGCATCAGGAGCACCCTGAGGATGGAGCTGATGGGCTGTGACCTGAACAGCTGCAGCATGCCCCTGGG CATGGAGAGCAAGGCCATCTCTGATGCCCAGATCACTGCCAGCAGCTACTTCACCAACATGTTTGCCACC TGGAGCCCCAGCAAGGCCAGGCTGCACCTGCAGGGCAGGAGCAATGCCTGGAGGCCCCAGGTCAACAACC CCAAGGAGTGGCTGCAGGTGGACTTCCAGAAGACCATGAAGGTGACTGGGGTGACCACCCAGGGGGTGAA GAGCCTGCTGACCAGCATGTATGTGAAGGAGTTCCTGATCAGCAGCAGCCAGGATGGCCACCAGTGGACC CTGTTCTTCCAGAATGGCAAGGTGAAGGTGTTCCAGGGCAACCAGGACAGCTTCACCCCTGTGGTGAACA GCCTGGACCCCCCCCTGCTGACCAGATACCTGAGGATTCACCCCCAGAGCTGGGTGCACCAGATTGCCCT GAGGATGGAGGTGCTGGGCTGTGAGGCCCAGGACCTGTACTGA SEQ ID NO: 73 Exemplified FVIII transgene (V3) ATGCAGATTGAGCTGAGCACCTGCTTCTTCCTGTGCCTGCTGAGGTTCTGCTTCTCTGCCACCAGGAGAT ACTACCTGGGGGCTGTGGAGCTGAGCTGGGACTACATGCAGTCTGACCTGGGGGAGCTGCCTGTGGATGC CAGGTTCCCCCCCAGAGTGCCCAAGAGCTTCCCCTTCAACACCTCTGTGGTGTACAAGAAGACCCTGTTT
GTGGAGTTCACTGACCACCTGTTCAACATTGCCAAGCCCAGGCCCCCCTGGATGGGCCTGCTGGGCCCCA CCATCCAGGCTGAGGTGTATGACACTGTGGTGATCACCCTGAAGAACATGGCCAGCCACCCTGTGAGCCT GCATGCTGTGGGGGTGAGCTACTGGAAGGCCTCTGAGGGGGCTGAGTATGATGACCAGACCAGCCAGAGG GAGAAGGAGGATGACAAGGTGTTCCCTGGGGGCAGCCACACCTATGTGTGGCAGGTGCTGAAGGAGAATG GCCCCATGGCCTCTGACCCCCTGTGCCTGACCTACAGCTACCTGAGCCATGTGGACCTGGTGAAGGACCT GAACTCTGGCCTGATTGGGGCCCTGCTGGTGTGCAGGGAGGGCAGCCTGGCCAAGGAGAAGACCCAGACC CTGCACAAGTTCATCCTGCTGTTTGCTGTGTTTGATGAGGGCAAGAGCTGGCACTCTGAAACCAAGAACA GCCTGATGCAGGACAGGGATGCTGCCTCTGCCAGGGCCTGGCCCAAGATGCACACTGTGAATGGCTATGT GAACAGGAGCCTGCCTGGCCTGATTGGCTGCCACAGGAAGTCTGTGTACTGGCATGTGATTGGCATGGGC ACCACCCCTGAGGTGCACAGCATCTTCCTGGAGGGCCACACCTTCCTGGTCAGGAACCACAGGCAGGCCA GCCTGGAGATCAGCCCCATCACCTTCCTGACTGCCCAGACCCTGCTGATGGACCTGGGCCAGTTCCTGCT GTTCTGCCACATCAGCAGCCACCAGCATGATGGCATGGAGGCCTATGTGAAGGTGGACAGCTGCCCTGAG GAGCCCCAGCTGAGGATGAAGAACAATGAGGAGGCTGAGGACTATGATGATGACCTGACTGACTCTGAGA TGGATGTGGTGAGGTTTGATGATGACAACAGCCCCAGCTTCATCCAGATCAGGTCTGTGGCCAAGAAGCA CCCCAAGACCTGGGTGCACTACATTGCTGCTGAGGAGGAGGACTGGGACTATGCCCCCCTGGTGCTGGCC CCTGATGACAGGAGCTACAAGAGCCAGTACCTGAACAATGGCCCCCAGAGGATTGGCAGGAAGTACAAGA AGGTCAGGTTCATGGCCTACACTGATGAAACCTTCAAGACCAGGGAGGCCATCCAGCATGAGTCTGGCAT CCTGGGCCCCCTGCTGTATGGGGAGGTGGGGGACACCCTGCTGATCATCTTCAAGAACCAGGCCAGCAGG CCCTACAACATCTACCCCCATGGCATCACTGATGTGAGGCCCCTGTACAGCAGGAGGCTGCCCAAGGGGG TGAAGCACCTGAAGGACTTCCCCATCCTGCCTGGGGAGATCTTCAAGTACAAGTGGACTGTGACTGTGGA GGATGGCCCCACCAAGTCTGACCCCAGGTGCCTGACCAGATACTACAGCAGCTTTGTGAACATGGAGAGG GACCTGGCCTCTGGCCTGATTGGCCCCCTGCTGATCTGCTACAAGGAGTCTGTGGACCAGAGGGGCAACC AGATCATGTCTGACAAGAGGAATGTGATCCTGTTCTCTGTGTTTGATGAGAACAGGAGCTGGTACCTGAC TGAGAACATCCAGAGGTTCCTGCCCAACCCTGCTGGGGTGCAGCTGGAGGACCCTGAGTTCCAGGCCAGC AACATCATGCACAGCATCAATGGCTATGTGTTTGACAGCCTGCAGCTGTCTGTGTGCCTGCATGAGGTGG CCTACTGGTACATCCTGAGCATTGGGGCCCAGACTGACTTCCTGTCTGTGTTCTTCTCTGGCTACACCTT CAAGCACAAGATGGTGTATGAGGACACCCTGACCCTGTTCCCCTTCTCTGGGGAGACTGTGTTCATGAGC ATGGAGAACCCTGGCCTGTGGATTCTGGGCTGCCACAACTCTGACTTCAGGAACAGGGGCATGACTGCCC TGCTGAAAGTCTCCAGCTGTGACAAGAACACTGGGGACTACTATGAGGACAGCTATGAGGACATCTCTGC CTACCTGCTGAGCAAGAACAATGCCATTGAGCCCAGGAGCTTCAGCCAGAATGCCACTAATGTGTCTAAC AACAGCAACACCAGCAATGACAGCAATGTGTCTCCCCCAGTGCTGAAGAGGCACCAGAGGGAGATCACCA GGACCACCCTGCAGTCTGACCAGGAGGAGATTGACTATGATGACACCATCTCTGTGGAGATGAAGAAGGA GGACTTTGACATCTACGACGAGGACGAGAACCAGAGCCCCAGGAGCTTCCAGAAGAAGACCAGGCACTAC TTCATTGCTGCTGTGGAGAGGCTGTGGGACTATGGCATGAGCAGCAGCCCCCATGTGCTGAGGAACAGGG CCCAGTCTGGCTCTGTGCCCCAGTTCAAGAAGGTGGTGTTCCAGGAGTTCACTGATGGCAGCTTCACCCA GCCCCTGTACAGAGGGGAGCTGAATGAGCACCTGGGCCTGCTGGGCCCCTACATCAGGGCTGAGGTGGAG GACAACATCATGGTGACCTTCAGGAACCAGGCCAGCAGGCCCTACAGCTTCTACAGCAGCCTGATCAGCT ATGAGGAGGACCAGAGGCAGGGGGCTGAGCCCAGGAAGAACTTTGTGAAGCCCAATGAAACCAAGACCTA CTTCTGGAAGGTGCAGCACCACATGGCCCCCACCAAGGATGAGTTTGACTGCAAGGCCTGGGCCTACTTC TCTGATGTGGACCTGGAGAAGGATGTGCACTCTGGCCTGATTGGCCCCCTGCTGGTGTGCCACACCAACA CCCTGAACCCTGCCCATGGCAGGCAGGTGACTGTGCAGGAGTTTGCCCTGTTCTTCACCATCTTTGATGA
AACCAAGAGCTGGTACTTCACTGAGAACATGGAGAGGAACTGCAGGGCCCCCTGCAACATCCAGATGGAG GACCCCACCTTCAAGGAGAACTACAGGTTCCATGCCATCAATGGCTACATCATGGACACCCTGCCTGGCC TGGTGATGGCCCAGGACCAGAGGATCAGGTGGTACCTGCTGAGCATGGGCAGCAATGAGAACATCCACAG CATCCACTTCTCTGGCCATGTGTTCACTGTGAGGAAGAAGGAGGAGTACAAGATGGCCCTGTACAACCTG TACCCTGGGGTGTTTGAGACTGTGGAGATGCTGCCCAGCAAGGCTGGCATCTGGAGGGTGGAGTGCCTGA TTGGGGAGCACCTGCATGCTGGCATGAGCACCCTGTTCCTGGTGTACAGCAACAAGTGCCAGACCCCCCT GGGCATGGCCTCTGGCCACATCAGGGACTTCCAGATCACTGCCTCTGGCCAGTATGGCCAGTGGGCCCCC AAGCTGGCCAGGCTGCACTACTCTGGCAGCATCAATGCCTGGAGCACCAAGGAGCCCTTCAGCTGGATCA AGGTGGACCTGCTGGCCCCCATGATCATCCATGGCATCAAGACCCAGGGGGCCAGGCAGAAGTTCAGCAG CCTGTACATCAGCCAGTTCATCATCATGTACAGCCTGGATGGCAAGAAGTGGCAGACCTACAGGGGCAAC AGCACTGGCACCCTGATGGTGTTCTTTGGCAATGTGGACAGCTCTGGCATCAAGCACAACATCTTCAACC CCCCCATCATTGCCAGATACATCAGGCTGCACCCCACCCACTACAGCATCAGGAGCACCCTGAGGATGGA GCTGATGGGCTGTGACCTGAACAGCTGCAGCATGCCCCTGGGCATGGAGAGCAAGGCCATCTCTGATGCC CAGATCACTGCCAGCAGCTACTTCACCAACATGTTTGCCACCTGGAGCCCCAGCAAGGCCAGGCTGCACC TGCAGGGCAGGAGCAATGCCTGGAGGCCCCAGGTCAACAACCCCAAGGAGTGGCTGCAGGTGGACTTCCA GAAGACCATGAAGGTGACTGGGGTGACCACCCAGGGGGTGAAGAGCCTGCTGACCAGCATGTATGTGAAG GAGTTCCTGATCAGCAGCAGCCAGGATGGCCACCAGTGGACCCTGTTCTTCCAGAATGGCAAGGTGAAGG TGTTCCAGGGCAACCAGGACAGCTTCACCCCTGTGGTGAACAGCCTGGACCCCCCCCTGCTGACCAGATA CCTGAGGATTCACCCCCAGAGCTGGGTGCACCAGATTGCCCTGAGGATGGAGGTGCTGGGCTGTGAGGCC CAGGACCTGTACTGA SEQ ID NO: 74 Exemplified FVIII polypeptide (N6) MQIELSTCFFLCLLRFCFSATRRYYLGAVELSWDYMQSDLGELPVDARFPPRVPKSFPFNTSVVYKKT LFVEFTDHLFNIAKPRPPWMGLLGPTIQAEVYDTVVITLKNMASHPVSLHAVGVSYWKASEGAEYDDQ TSQREKEDDKVFPGGSHTYVWQVLKENGPMASDPLCLTYSYLSHVDLVKDLNSGLIGALLVCREGSLA KEKTQTLHKFILLFAVFDEGKSWHSETKNSLMQDRDAASARAWPKMHTVNGYVNRSLPGLIGCHRKSV YWHVIGMGTTPEVHSIFLEGHTFLVRNHRQASLEISPITFLTAQTLLMDLGQFLLFCHISSHQHDGME AYVKVDSCPEEPQLRMKNNEEAEDYDDDLTDSEMDVVRFDDDNSPSFIQIRSVAKKHPKTWVHYIAAE EEDWDYAPLVLAPDDRSYKSQYLNNGPQRIGRKYKKVRFMAYTDETFKTREAIQHESGILGPLLYGEV GDTLLIIFKNQASRPYNIYPHGITDVRPLYSRRLPKGVKHLKDFPILPGEIFKYKWTVTVEDGPTKSD PRCLTRYYSSFVNMERDLASGLIGPLLICYKESVDQRGNQIMSDKRNVILFSVFDENRSWYLTENIQR FLPNPAGVQLEDPEFQASNIMHSINGYVFDSLQLSVCLHEVAYWYILSIGAQTDFLSVFFSGYTFKHK MVYEDTLTLFPFSGETVFMSMENPGLWILGCHNSDFRNRGMTALLKVSSCDKNTGDYYEDSYEDISAY LLSKNNAIEPRSFSQNSRHPSTRQKQFNATTIPENDIEKTDPWFAHRTPMPKIQNVSSSDLLMLLRQS PTPHGLSLSDLQEAKYETFSDDPSPGAIDSNNSLSEMTHFRPQLHHSGDMVFTPESGLQLRLNEKLGT TAATELKKLDFKVSSTSNNLISTIPSDNLAAGTDNTSSLGPPSMPVHYDSQLDTTLFGKKSSPLTESG GPLSLSEENNDSKLLESGLMNSQESSWGKNVSSREITRTTLQSDQEEIDYDDTISVEMKKEDFDIYDE DENQSPRSFQKKTRHYFIAAVERLWDYGMSSSPHVLRNRAQSGSVPQFKKVVFQEFTDGSFTQPLYRG ELNEHLGLLGPYIRAEVEDNIMVTFRNQASRPYSFYSSLISYEEDQRQGAEPRKNFVKPNETKTYFWK
VQHHMAPTKDEFDCKAWAYFSDVDLEKDVHSGLIGPLLVCHTNTLNPAHGRQVTVQEFALFFTIFDET KSWYFTENMERNCRAPCNIQMEDPTFKENYRFHAINGYIMDTLPGLVMAQDQRIRWYLLSMGSNENIH SIHFSGHVFTVRKKEEYKMALYNLYPGVFETVEMLPSKAGIWRVECLIGEHLHAGMSTLFLVYSNKCQ TPLGMASGHIRDFQITASGQYGQWAPKLARLHYSGSINAWSTKEPFSWIKVDLLAPMIIHGIKTQGAR QKFSSLYISQFIIMYSLDGKKWQTYRGNSTGTLMVFFGNVDSSGIKHNIFNPPIIARYIRLHPTHYSI RSTLRMELMGCDLNSCSMPLGMESKAISDAQITASSYFTNMFATWSPSKARLHLQGRSNAWRPQVNNP KEWLQVDFQKTMKVTGVTTQGVKSLLTSMYVKEFLISSSQDGHQWTLFFQNGKVKVFQGNQDSFTPVV NSLDPPLLTRYLRIHPQSWVHQIALRMEVLGCEAQDLY SEQ ID NO: 75 Exemplified FVIII polypeptide (V3) MQIELSTCFFLCLLRFCFSATRRYYLGAVELSWDYMQSDLGELPVDARFPPRVPKSFPFNTSVVYKKTLF VEFTDHLFNIAKPRPPWMGLLGPTIQAEVYDTVVITLKNMASHPVSLHAVGVSYWKASEGAEYDDQTSQR EKEDDKVFPGGSHTYVWQVLKENGPMASDPLCLTYSYLSHVDLVKDLNSGLIGALLVCREGSLAKEKTQT LHKFILLFAVFDEGKSWHSETKNSLMQDRDAASARAWPKMHTVNGYVNRSLPGLIGCHRKSVYWHVIGMG TTPEVHSIFLEGHTFLVRNHRQASLEISPITFLTAQTLLMDLGQFLLFCHISSHQHDGMEAYVKVDSCPE EPQLRMKNNEEAEDYDDDLTDSEMDVVRFDDDNSPSFIQIRSVAKKHPKTWVHYIAAEEEDWDYAPLVLA PDDRSYKSQYLNNGPQRIGRKYKKVRFMAYTDETFKTREAIQHESGILGPLLYGEVGDTLLIIFKNQASR PYNIYPHGITDVRPLYSRRLPKGVKHLKDFPILPGEIFKYKWTVTVEDGPTKSDPRCLTRYYSSFVNMER DLASGLIGPLLICYKESVDQRGNQIMSDKRNVILFSVFDENRSWYLTENIQRFLPNPAGVQLEDPEFQAS NIMHSINGYVFDSLQLSVCLHEVAYWYILSIGAQTDFLSVFFSGYTFKHKMVYEDTLTLFPFSGETVFMS MENPGLWILGCHNSDFRNRGMTALLKVSSCDKNTGDYYEDSYEDISAYLLSKNNAIEPRSFSQNATNVSN NSNTSNDSNVSPPVLKRHQREITRTTLQSDQEEIDYDDTISVEMKKEDFDIYDEDENQSPRSFQKKTRHY FIAAVERLWDYGMSSSPHVLRNRAQSGSVPQFKKVVFQEFTDGSFTQPLYRGELNEHLGLLGPYIRAEVE DNIMVTFRNQASRPYSFYSSLISYEEDQRQGAEPRKNFVKPNETKTYFWKVQHHMAPTKDEFDCKAWAYF SDVDLEKDVHSGLIGPLLVCHTNTLNPAHGRQVTVQEFALFFTIFDETKSWYFTENMERNCRAPCNIQME DPTFKENYRFHAINGYIMDTLPGLVMAQDQRIRWYLLSMGSNENIHSIHFSGHVFTVRKKEEYKMALYNL YPGVFETVEMLPSKAGIWRVECLIGEHLHAGMSTLFLVYSNKCQTPLGMASGHIRDFQITASGQYGQWAP KLARLHYSGSINAWSTKEPFSWIKVDLLAPMIIHGIKTQGARQKFSSLYISQFIIMYSLDGKKWQTYRGN STGTLMVFFGNVDSSGIKHNIFNPPIIARYIRLHPTHYSIRSTLRMELMGCDLNSCSMPLGMESKAISDA QITASSYFTNMFATWSPSKARLHLQGRSNAWRPQVNNPKEWLQVDFQKTMKVTGVTTQGVKSLLTSMYVK EFLISSSQDGHQWTLFFQNGKVKVFQGNQDSFTPVVNSLDPPLLTRYLRIHPQSWVHQIALRMEVLGCEA QDLY EXAMPLES The invention is now described with reference to the Examples below. These are not limiting on the scope of the invention, and a person skilled in the art would be appreciate that suitable equivalents could be used within the scope of the present invention. Thus, the Examples may be considered component parts of the invention, and the individual aspects described therein may be considered as disclosed independently, or in any combination.
Example 1 – Design and in vitro validation of a cell‐specific core SFPB (mSPB) promoter sequence The SFTPB promoter fragment of SEQ ID NO: 1 was designed following careful analysis of the SFTPB genomic sequence. The expression driven by the core SFTPB promoter was compared to the full length SFTPB promoter (the 972bp fragment of SEQ ID NO: 2, fSPB) using a human surfactant air‐ liquid interface (SALI) culture model; an in vitro cell culture model that robustly recapitulates human ATII cells in primary cell culture (Munis et al. (2021) Molecular Therapy: Methods & Clinical Development 20: 237‐246). H441 cells, when grown under SALI culture conditions, successfully mimic key characteristics of primary ATII cells. Briefly, SALI cultures were established by culturing cells in 12‐well Transwell inserts. Approximately 1 × 105 cells/well in base media (5 × 104 each of H441 and A549 cells in the case of co‐ culture) were seeded into Transwells and allowed to attach and proliferate for 48 h. On day 3, the medium on the apical side of the Transwell chamber was removed to air‐lift the cells, and the medium on the basolateral side of the chamber was replaced with either “base” medium or “polarization” medium. The polarization medium comprised either RPMI‐1640 (H441s cells and co‐culture) or F12‐K (A549s) supplemented with 2 mM l‐glutamine, 50 U/mL penicillin, 50 mg/mL streptomycin, 1% insulin‐ transferrin‐selenium (GIBCO), 4% FCS, and 1 μM dexamethasone (Sigma). Media were subsequently changed three times per week throughout the described experiments. Cells were grown under SALI conditions for 14 days after air‐lift prior to experimentation unless otherwise stated. SALI cultures were transduced with titre‐matched (1x106 TU) lentiviral vector (LV) expressing EGFP from range of promoters (CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), EF1aS (SEQ ID NO: 7), PGK (SEQ ID NO: 8), the core SFTPB fragment (mSPB) (SEQ ID NO: 1), full‐length SPB (fSPB) (SEQ ID NO: 4), minimal SPC (mSPC) (SEQ ID NO: 9) and full‐length SPC (fSPC) (SEQ ID NO: 10)). After 48 hours cells were subject to FACS analysis to quantify the % of cells expressing EGFP, as indicated in the respective FACS plot shown Figure 1. The negative control sample (Naïve; non‐ transduced cells H441 cells) shows a background level of 0.4%. The % EGFP‐positive cells observed with the widely‐used, non‐specific CMV [SEQ ID NO: 6], EF1aS [SEQ ID NO; 7], and PGK [SEQ ID NO: 8] promoters is relatively high (17‐22%) in line with expectations. The % EGFP‐positive cells observed for the hCEF promoter [SEQ ID NO: 5], which has been used previously in the lungs of patients, is 11%. Promoter sequences taken from lung surfactant genes (mSPB [SEQ ID NO: 1], fSPB [SEQ ID NO: 4], mSPC [SEQ ID NO: 9], fSPC [SEQ ID NO: 10]) led to relatively lower % EGFP‐positive cells (0.7% to 9%) suggesting that expression from these sequences have potential utility in a rationally designed new promoter specific for ATII cells.
In particular, as shown in Figure 1, the mSPB promoter achieves expression levels (9%) similar to hCEF (11%) in hSALI cultures. Notably, the mSPB promoter also achieves higher expression levels compared to the full‐length SPB (fSPB) promoter sequence. By way of comparison, human HEK293T cells were transduced with recombinant SIV lentiviral vectors pseudotyped with VSV‐G and expressing the EGFP transgene from these same promoters (CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), EF1aS (SEQ ID NO: 7), PGK (SEQ ID NO: 8), the core SFTPB fragment (mSPB) (SEQ ID NO: 1), full‐length SPB (fSPB) (SEQ ID NO: 4), minimal SPC (mSPC) (SEQ ID NO: 9) and full‐length SPC (fSPC) (SEQ ID NO: 10)) at a multiplicity of infection (moi) of 0.5 or 1.0. After 48 hours, cells were subject to FACS analysis to quantify the % of cells expressing EGFP. As shown in Figure 2, the % EGFP‐positive cells is low for mSPB, fSPB, mSPC and fSPC promoter sequences in generic HEK293T cells, indicating that the activity of these may be cell‐specific. In particular, the mSPB promoter drives very low level expression in HEK293T cells (see, Figure 2). The expression driven by the mSPB promoter is likely to be specific to lung parenchymal cells, particularly ATII cells. Example 2 – The core SFPB (mSPB) promoter sequence drives cell‐specific gene expression in vivo The specificity of the mSPB promoter was further investigated using an in vivo mouse model. As shown in Figure 3A, on day 0 mice were dosed by nasal instillation with SIV vector. 7 days after dosing, lung tissue was harvested, fixed, frozen and cryosectioned. The sections were then analysed for expression of EGFP. As shown in Figure 3B, on day 0 female BALB/c mice (4‐6 weeks; n=3 per group) were dosed with 1x106 TU (in 100µL TSSM buffer) of rSIV.F/HN.mSPB, an SIV vector comprising EGFP under the control of an mSPB promoter (SEQ ID NO: 1). As controls, mice were alternatively dosed with 100 µl of vehicle only (TSSM) or 1x106 TU (in 100µL TSSM buffer) of rSIV.F/HN.fSPB, an SIV vector comprising EGFP under the control of an fSPB promoter (SEQ ID NO: 4). Representative images are shown in Figure 4. As expected, no EGFP expression was seen in the tissue obtained from the native mice. Foci corresponding to EGFP expression were seen in the sections obtained from mice injected with either rSIV.F/HN.mSPB or rSIV.F/HN.fSPB. However, there are very few EGFP positive cells in the airways, indicating that the SFTPB promoters drive specific gene expression in cells of the lung parenchyma, rather than airway cells. Signficantly, approximately 5‐fold more punctate fluorescent signal was observed in lung sections from the mSPB group compared with the fSPB group of mice. These data suggest that the core SFTPB promoter fragment (mSPB) is capable of achieving significantly greater transgene expression than the full‐length SFTPB promoter of fSPB. Example 3 – Design and production of improved mSPB promoters
The inventors then sought to further increase transgene expression by and/or activity of (i.e. the number of lung parenchyma cells in which the promoter is active) the mSPB promoter, whilst retaining specificity for the lung parenchyma. Specifically, the inventors generated a panel of improved mSPB promoters, each comprising mSPB and an enhancer. Different enhancer sequences were selected, including ubiquitously used strong enhancers (i.e., CMV, hB‐actin or SV40), as well as tissue specific enhancers. Lung‐specific enhancers were selected using the ATII cell gene expression database from LungGENS and a tissue expression database. To generate candidate sequences to boost the level of transgene expression from the mSPB (SEQ ID NO: 1) promoter sequence in ATII cells, tissue expression databases including https://research.cchmc.org/pbge/lunggens/default.html and https://tissues.jensenlab.org/Search were interrogated to identify genes with high mean levels of expression in ATII cells. The resulting panel of ATII specific genes (and enhancers) is shown in Figure 5. Of these, SFTPC, SFTPB, SLC34A2, GOLGA8B, LMO7, VEGFA and ELF3 were found to be most highly expressed in ATII cells. SFTPC was found to be highly specific for the lungs. Although SLC34A2 was found to be highly expressed in the lungs, high levels of expression was also found in the gut. Further, GOLGA8B was found to be most highly expressed in the brain. Enhancer regions within the above‐mentioned genes were identified using Genecards: https://www.genecards.org/Guide/GeneCard. Specific sequences in these enhancer regions were selected (and transcription factor binding sites identified) using University of California at Santa Cruz (UCSC) genome browser https://genome.ucsc.edu. The identified enhancer sequences were cloned into a construct comprising mSPB and an operably linked EGFP2ALux transgene. To determine whether expression from the enhancers is affected by orientation, enhancers were added to the constructs in the forward (F) and reverse (R) direction. Schematics of exemplary constructs are shown in Figure 6, with the mSPB promoter sequence of SEQ ID NO: 1 and enhancer sequences as per SEQ ID NOs: 13 to 20, 23 to 33 and 36 to 40. Similar reporter transgene expression constructs were also generated incorporating the commonly used enhancers hB‐Actin (SEQ ID NOs: 21 and 22), SV40 (SEQ ID NOs: 34 and 35) and CMV (SEQ ID NOs: 11 and 12). Using transient transfection with standard lentiviral producer plasmids, the expression constructs were incorporated into recombinant lentiviral vectors. The various lentiviral vectors were prepared for analysis of EGFP or Lux reporter transgene expression from each enhancer/promoter combination.
Human HEK293T cells, murine LA‐4 cells and human SALI cells were transfected with the mSPB promoter constructs to determine whether the addition or the enhancer (or the orientation) increased gene expression. Recombinant SIV lentiviral vectors pseudotyped with VSV‐G and expressing EGFP2ALux transgene from each of the enhancer/promoter combinations, were produced at small scale. These were used to transduce human HEK293T cells in a 24‐well plate (seeded @1x105 cells/well) at a multiplicity of infection (moi) of approximately 0.1. Luciferase levels (Relative Light Units (RLU)) were determined in n=4 independent biological replicates. Naïve (non‐transduced) cells were used as a control to determine background RLU level. The results for human HEK293T cells are shown in Figure 7. In HEK293T cells, mSPB did not significantly increase gene expression relative to the naïve control. Differences in expression driven by mSPB compared with mSPB + the different enhancers was also not significant. In contrast, significant differences in gene expression are seen in murine LA‐4 cells (see, Figure 8) and human SALI cells (see, Figure 9). In particular, in murine LA‐4 cells, the addition of a CMV enhancer (forwards or reverse), ELF3 enhancer (forwards), SV40 enhancer (forwards), and two of the VEGFA enhancers – VEG1 (forwards or reverse) and VEG2 (forwards) to the mSPB promoter significantly increased expression relative to the mSPB promoter alone. Similarly, in human SALI cells, the addition of a CMV enhancer (forwards or reverse), Elf3 (forwards or reverse) enhancer, SLC3 (reverse) enhancer, SV40 (forwards or reverse) enhancer, and two of the VEGFA enhancers – VEG1 (forwards or reverse) and VEG3 (reverse) to the mSPB promoter significantly increased expression relative to the mSPB promoter alone. These data demonstrate that it is possible to further increase transgene expression in a cell‐ specific manner, even beyond the level obtained using the mSPB promoter (which itself provides a significant increase in expression compared with the fSPB promoter), using specific enhancer sequences. Example 4 – Further characterisation of improved mSPB promoters Using the results of Example 3, the following seven candidate enhancers were selected for further screening: SV40, CMV, SLC34A2‐1, SLC34A2‐2, VEGFA‐2, ELF3‐1 and ELF3‐2. Recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing EGFP2ALux reporter transgene, were used to transduce human SALI cultures (n=6 replicates) in a repeat secondary screening experiment, which included the following promoter sequences: hCEF (SEQ ID NO: 5), CMV (SEQ ID NO: 6), and mSPB (SEQ ID NO: 1), as well as the selected enhancer/promoter combinations SV40 (f) (SEQ ID NO: 34), CMVenh (f (SEQ ID NO: 11)), Elf1 (f) (SEQ ID NO: 13), Elf2 (f)
(SEQ ID NO: 15), Slc1 (f) (SEQ ID NO: 29), Slc2 (f) (SEQ ID NO: 31), and Veg2 (f) (SEQ ID NO: 38)) in combination with mSPB (SEQ ID NO: 1). In a secondary screen, expression driven by the mSPB promoter combined with the CMV, SV40 and VEGFA‐2 enhancer in SALI cells was found to be comparable to the level of expression driven by the strong promoter hCEF (Figure 10). Recombinant SIV lentiviral vectors pseudotyped with the F/HN and expressing EGFP2ALux reporter transgene, were then used to transduce human HEK293T cell cultures (n=6 replicates) in a repeat secondary screening experiment, which included the following selected enhancer/promoter sequences: hCEF (SEQ ID NO: 5), CMV (SEQ ID NO: 6), and mSPB (SEQ ID NO: 1), as well as the selected enhancer/promoter combinations SV40 (f) (SEQ ID NO: 34), CMVenh (f (SEQ ID NO: 11)), Elf1 (f) (SEQ ID NO: 13), Elf2 (f) (SEQ ID NO: 15), Slc1 (f) (SEQ ID NO: 29), Slc2 (f) (SEQ ID NO: 31), and Veg2 (f) (SEQ ID NO: 38)) in combination with mSPB (SEQ ID NO: 1). These data demonstrate that transgene expression using the hCEF promoter was not cell‐specific, with significantly greater expression with hCEF seen in HEK293T cells compared with expression using the mSPB/enhancers (Figure 11). The inventors next assessed whether the mSPB/enhancer improved promoters were able to drive expression of luciferase in vivo. Lentiviral vectors comprising different promoter constructs were administered to mice via the nose and the lungs. The mice were then monitored for luciferase expression. Female BALB/c mice (6 weeks; n=5 per group) were dosed with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing EGFP2ALux reporter transgene from the following promoter sequences: hCEF (SEQ ID NO: 5), CMV (SEQ ID NO: 6), and mSPB (SEQ ID NO: 1), as well as the selected enhancer/promoter combinations SV40 (f) (SEQ ID NO: 34), CMVenh (f) (SEQ ID NO: 11), Elf1 (f) (SEQ ID NO: 13), Elf2 (f) (SEQ ID NO: 15), Slc1 (f) (SEQ ID NO: 29), Slc2 (f) (SEQ ID NO: 31), and Veg2 (f) (SEQ ID NO: 38)) in combination with mSPB (SEQ ID NO: 1). Lentivirus was dosed intranasally (1x107 TU per mouse in 100uL TSSM Buffer). On days 2, 7 and 28 post‐dosing, all mice were anaesthetised and imaged for luciferase expression. Figure 12 shows the expression of luciferase as a heat map. Expression driven by hCEF or CMV promoters is shown as a reference. The inventors selected promoter which drive expression in the lungs, but not the nose. Localised expression in the lungs is important, because the nose has cells that are similar to (airway epithelial cells ciliated/non‐ciliated epithelial) to those in the lungs, so it is important to determine that the transgene will not be expressed in the nose. The expression of luciferase is quantified in Figure 13A. The ratio of expression in the lungs and the nose is shown in Figure 13B. The level of luciferase signal (Figure 13A) was greatest with the CMVenh group (SEQ ID NO: 11) (**** compared to mSPB (SEQ ID NO: 1) although the specificity for lung expression (Figure 13B), determined by the ratio of signal in the lung and nose (L‐to‐N), was not
significantly different from CMV (SEQ ID NO: 6)or hCEF (SEQ ID NO: 5) at day 28. The luciferase signal (Figure 13A) in the slc2 (SEQ ID NO: 31) and Veg2 (SEQ ID NO: 38) groups was significantly increased compared with mSPB (SEQ ID NO: 1) (*and ** respectively, compared to hCEF group; # NS compared to mSPB; Kruskal Wallis) and also showed good overall specificity of lung expression (Figure 13A). Thus, the results indicate that the mSPB promoter combined with an SLC2, VEGF2 or CMV enhancer drives high levels of expression in the lungs, but not the nose. Human SALI cultures (n=8) were transfected with plasmid DNA (2µg plasmid per culture) complexed with linear polyethyleneimine (PEIPro; 100µL per culture) and luciferase activity in Relative Light Units (RLU) determined. A naïve group (n=8; non‐transfected) was included as a negative control. Similar results to those shown in Figure 13 using a SIV vector were observed when the promoters were used to drive expression of luciferase in a non‐viral (plasmid) vector, with the mSPB promoter/enhancer combinations driving increased transgene expression in SALI cells (Figure 14A) and in vivo (Figure 14B) compared with expression in HEK293T cells (Figure 14C). In particular, luciferase activity in the veg2 (SEQ ID NO: 38) and CMVenh (SEQ ID NO: 11) groups were significantly different from mSPB (SEQ ID NO: 1). Overall, the levels of transgene expression (luciferase) from these constructs delivered as a non‐viral (plasmid) formulation were similar to those obtained following delivery with viral (recombinant lentiviral) vectors. Example 5 – Improved mSPB promoters drive targeted expression in cells of the lung parenchyma Further analysis was conducted to confirm that the mSPB promoter/enhancer combinations result in transgene expression in the target cells, particularly that expression is focussed in the cells of the lung parenchyma, rather than airway cells. This was investigated using immunohistochemistry. As described in Example 4, Figure 12, female BALB/c mice (6 weeks; n=5 per group) were dosed with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing EGFP2ALux reporter transgene from different enhancer/promoter sequences. Lentivirus was dosed intranasally (1x107 TU per mouse in 100µL TSSM Buffer) and culled on day 28 post‐dosing. The enhancer sequences CMVenh (SEQ ID NO: 11), Slc2 (SEQ ID NO: 31) and VEGFA‐2 (SEQ ID NO: 38) combined with the mSPB (SEQ ID NO: 1) promoter sequence were assigned the nomenclature Alv‐01, Alv‐02 and Alv‐03 respectively, as shown in SEQ ID NOs: 46, 47 and 48 respectively. Mouse lungs were processed for cryosections and imaging. Cryosections (7µM) were subject to immunohistochemistry using primary antibodies to detect colocalization of EGFP and Pro/Mature Surfactant Protein‐B, the latter being a marker for ATII cells. Cryosections were also stained with DAPI for ease of visualisation and then imaged (at least n=3 sections per group) using a confocal microscope. As shown in Figure 15, Alv‐01, Alv‐02 and Alv‐03 were able to drive EGFP expression which colocalises with the ATII cell‐specific
marker SP‐B. These results demonstrate that the newly identified enhancer/mSPB promoter constructs can facilitate transgene expression in ATII cells in the lung parenchyma. In particular, this experiment confirmed that the mSPB promoter/enhancer combinations tested drove transgene expression primarily in cells of the parenchyma, making them attractive candidates for targeted gene therapy of diseases resulting from or associated with deficiency of expression and/or expression of defective proteins in lung parenchymal cells, such as surfactant deficiencies. Repetition of this experiment using a different ATII cell marker, Surfactant Protein‐C yielded the same pattern of expression, with Alv‐01, Alv‐02 and Alv‐03 driving EGFP expression which colocalises with the ATII cell‐specific marker SP‐C (data not shown). Therefore, Alv‐01, Alv‐02 and Alv‐ 03 have been reproducibly shown to drive expression in the lung parenchyma. Example 6 – Improved mSPB promoters drive long‐term expression in vivo The in vivo expression of luciferase driven by the mSPB/enhancer improved promoters Alv‐ 01, Alv‐02 and Alv‐03 as described in Example 5 above was further characterised. Lentiviral vectors comprising mSPB, Alv‐01, Alv‐02 and Alv‐03 were administered to intranasally. The mice were then monitored for luciferase expression. Female BALB/c mice (n=10 per group) were dosed with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing a firefly luciferase reporter transgene from mSPB, Alv‐01, Alv‐02 and Alv‐03. SIV lentiviral vectors pseudotyped with F/HN and expressing a firefly luciferase reporter transgene from a CMV or hCEF promoter were used as a control. Lentivirus was dosed intranasally (2.5x108 TU per mouse in TSSM Buffer). On days 7, 14, 28, 91 and 186 post‐dosing, mice were anaesthetised and imaged for luciferase expression. Figure 16 shows the expression of luciferase as a heat map. Expression driven by hCEF or CMV promoters is shown as a reference. The expression of luciferase in the lungs over the time course of the experiment is quantified in Figure 16B. The area under the curve is shown in Figure 16C and the ratio of expression in the lungs and the nose is shown in Figure 16D. All of hCEF, mSPB, Alv‐01, Alv‐02 and Alv‐03 gave significantly higher expression of the luciferase reporter compared with the CMV promoter (Figure 16A, B and C) The level of luciferase signal (Figure 16A, B and C) was greatest with Alv‐01. The CMV promoter and hCEF promoter also drove significant expression in the nose compared with the mSPB, Alv‐02 and Alv‐ 03 promoters (Figure 16A). Specificity for lung expression (Figure 16D), determined by the ratio of signal in the lung and nose (L‐to‐N), was significantly increased by the mSPB, Alv‐02 and Alv‐03 promoters compared with the CMV or hCEF promoters at day 28. In this experiment, the lung specificity of Alv‐01 was low (although further experiments below demonstrate specificity with Alv‐
01). The specificity of luciferase signal in the lungs driven by the Alv‐02 and Alv‐03 promoters was observably increased compared with the mSPB promoter(Figure 16D). Furthermore, the high level of luciferase expression and high degree of specificity was maintained over the 6‐month time course of the experiment (Figure 16A and B). Thus, the results indicate that the Alv‐01, Alv‐02 and Alv‐03 promoter drives high levels of expression in the lungs over at least a six‐month period. Further, mSPB, Alv‐02 and Alv‐03 drive high levels of expression in the lungs, but not the nose. Example 7 – Improved mSPB promoters drive targeted expression in cells of the lung parenchyma Further analysis was conducted to confirm that Alv‐01, Alv‐02 and Alv‐03 result in transgene expression in the target cells, particularly that expression is focussed in the cells of the lung parenchyma, rather than airway cells. This was investigated using immunohistochemistry. As described in Example 5, female BALB/c mice (n=5‐10 per group) were dosed with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing EGFP from mSPB, Alv‐01, Alv‐02 and Alv‐03. Lentivirus was dosed intranasally and culled on day 14 post‐dosing. Cryosections (7µM) were subject to immunohistochemistry using primary antibodies to detect EGFP in the airway and parenchymal cells. Cryosections were also stained with DAPI for ease of visualisation and then imaged (at least n=3 sections per group) using a confocal microscope. As shown in Figure 17A, the hCEF promoter mainly drove expression in the cells of the airway, whereas as shown in Figure 17B, the mSPB promoter drove expression primarily in the parenchymal cells. The % of EGFP expression in the parenchymal cells for hCEF, mSPB, Alv‐01, Alv‐02 and Alv‐03 was quantified. As shown in Figure 17C , Alv‐01, Alv‐02 and Alv‐03, particularly Alv‐02 and Alv‐03 were able to drive EGFP expression in the parenchyma compared with the hCEF promoter. These results demonstrate that Alv‐01, Alv‐02 and Alv‐03, particularly Alv‐02 and Alv‐03 can facilitate transgene expression in the lung parenchyma. In particular, this experiment confirmed that the mSPB promoter, Alv‐01, Alv‐02 and Alv‐03, particularly Alv‐02 and Alv‐03, drove transgene expression primarily in cells of the parenchyma, making them attractive candidates for targeted gene therapy of diseases resulting from or associated with deficiency of expression and/or expression of defective proteins in lung parenchymal cells, such as surfactant deficiencies. Example 8 – Expression of SP‐B restores transepithelial electrical resistance (TEER) in SFTPB knock out lung cells To investigate whether the mSPB, Alv‐01, Alv‐02 and Alv‐03 promoters could perform well in the context of a human lung cell model, a Surfactant Air Liquid Interface (SALI) model was used with
the human H441 lung cell line and the H441 SP‐B KO cell line where the SFTPB gene had been ablated (Munis et al. (2021) Mo. Ther. Methods Clin. Dev 20:20:237‐246, herein incorporated by reference). The ablation of the gene encoding SP‐B leads to an observed phenotype – a reduction in transepithelial electrical resistance (TEER) ‐ that can be corrected by expression of human SP‐B. SALI cultures were generated following airlift and transduced with rSIV.F/HN vectors expressing human SP‐B under the control of one of the following promoters: hCEF, mSP‐B, Alv‐1, Alv‐ 2, or Alv‐3. Transduction with a rSIV.F/HN vector expressing EGFP under the control of hCEF was used as a negative control (mock). Before transduction, and up to 30 days thereafter, TEER was measured. As previously published, (Figure 18 A and B) TEER values from SALI cultures generated from the H441 SP‐B KO cells were lower than the parental H441 cell line at 14 days post airlift (i.e. 4 days prior to transductions with viral vectors) constituting a phenotypic defect. At 5 days post transduction, rSIV.F/HN expressing EGFP control vector did not increase the TEER in H441 SP‐B KO cells SALI cultures (Figure 18C), however transduction with any/all of the vectors expressing SP‐B showed correction of the TEER towards normal, calculated as a percentage of the mock‐transduced parental H441 cells (Figure 18D). These data show that expression of human SP‐B under control of mSP‐B, or Alv‐1‐3 promoters can also phenotypically correct the TEER defect offering alternative sequences for gene expression in ATII cells. These data show that switching from the strong hCEF promoter to the lung parenchymal specific mSP‐B, Alv‐1, Alv‐2, or Alv‐3 promoters does not have a negative effect on TEER restoration. Thus, these promoters advantageously allow parenchymal‐specific expression of SFTPB whilst achieving the same clinically desirable restoration of TEER. Example 9 – Pulmonary surfactant has no effect on HEK293/T cell transduction with rSIV.F/HN vectors To investigate whether complexing rSIV.F/HN vectors with synthetic surfactant (such as BLES or Curosurf) affects cell transduction, rSIV.F/HN vector encoding EGFP under CMV promoter control, was mixed 1:1 with TSSM (vehicle control) or with BLES or Curosurf and incubated at room temperature for 30mins. The mixtures were then diluted in OptiMEM‐I (supplemented with polybrene for final 8μg/mL working concentration) and used to transduce HEK293T cells. At 72h post‐transduction, cells were analysed by flow cytometry, and transduction efficiencies were corrected by subtracting the background fluorescence from cells treated with the Sham (no vector) control mixed with respective TSSM, BLES, or Curosurf. As shown in Figure 19, there was no statistically significant difference in transduction efficiency between vectors mixed 1:1 with BLES or
Curosurf, compared with control vector mixed with TSSM buffer. These data show that pulmonary surfactants (BLES and Curosurf) can be complexed with rSIV.F/HN vector without compromising transduction of human HEK293T cells. It was concluded that, despite the rSIV.F/HN lentiviral vectors having a lipid envelope, this is not affected by the presence of the pulmonary surfactant, and thus the pulmonary surfactants do not compromise the integrity of rSIV.F/HN lentiviral vectors. Example 10 – Murine lung transduction with rSIV.F/HN vectors is at least as effective in the presence of a pulmonary surfactant Having surprisingly shown that pulmonary surfactants do not compromise the integrity of rSIV.F/HN lentiviral vectors in an in vitro setting, the effect of pulmonary surfactants on rSIV.FHN transduction of the murine lung was then investigated. rSIV.F/HN vector expressing firefly luciferase from the hCEF promoter (rSIV.F/HN hCEF Flux) was mixed 1:1 with TSSM (vehicle control) or BLES or Curosurf (to generate 2E8 TU in 100uL per dose). BALB/c mice (n=5) were dosed by intranasal administration. TSSM diluent served as vehicle control. Mice were subjected to in vivo bioluminescent imaging to measure Firefly luciferase expression in the murine lungs. The time course of luciferase expression in the lungs of each mouse on days 7, 14, 28, 112 days post‐dosing is shown in Figure 20A, from which it can be seen that the luciferase expression was stable across the duration of the experiment. When the total luciferase was quantified as area under the curve (AUC) in Figure 20B, a statistically significant increase in transformation was achieved in the presence of either surfactant compared with the vehicle control. When this was compared with the TSSM vehicle control (Figure Figure 20C) a ∼2‐3‐fold improved expression was observed when the vector was mixed with either pulmonary surfactant BLES and Curosurf.
Claims
CLAIMS 1. An SFTPB promoter fragment which comprises or consists of at least part of exon 1 of the SFTPB gene, but does not comprise the SFTPB gene start codon.
2. The SFTPB promoter fragment of claim 1, which is less than 800 bases in length, preferably less than 700 bases in length.
3. The SFTPB promoter fragment of claim 1 or 2, which comprises or consists of: (a) SEQ ID NO: 1 or a sequence with at least 80% identity to SEQ ID NO: 1; or (b) bases 81‐710 of SEQ ID NO: 2, or a sequence with at least 80% identity to bases 81‐710 of SEQ ID NO: 2; wherein optionally: (i) said SFTPB promoter fragment further comprises up to 20 bases at the 5’ end, which may optionally correspond to up to 20 bases 5’ to base 81 of SEQ ID NO: 2; and/or (ii) said SFTPB promoter fragment further comprises up to 4 bases at the 3’ end, which may optionally correspond to up to 4 bases 3’ to base 710 of SEQ ID NO: 2.
4. The SFTPB promoter fragment of any one of claims 1‐3, which comprises or consists of SEQ ID NO: 1 or a sequence with at least 90% identity to SEQ ID NO: 1.
5. The SFTPB promoter fragment of any one of the preceding claims, which further comprises a 5’ enhancer.
6. The SFTPB promoter fragment of claim 5, wherein the enhancer is in (i) the forward, or (ii) the reverse, orientation; preferably wherein the enhancer is in the forward orientation.
7. The SFTPB promoter fragment of claim 5 or 6, wherein the enhancer is selected from a SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer, an ELF3 enhancer, an actin enhancer, an LMO7 enhancer, a SFTPC enhancer or a SFTPB enhancer,
wherein preferably the enhancer is selected from a SLC34A2 enhancer, a VEGFA enhancer or a CMV enhancer.
8. The SFTPB promoter fragment of claim 7, wherein: (a) the SLC34A2 enhancer comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 29‐33, preferably SEQ ID NO: 31; (b) the VEGFA enhancer comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38; (c) the CMV enhancer comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11; (d) the SV40 enhancer comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or 35; (e) the ELF3 enhancer comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 13‐20; (f) the actin enhancer comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 21 or 22; (g) the LMO7 enhancer comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 23—25; (h) the SFTPC enhancer comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 27 or 28; or (i) the SFTPB enhancer comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 26.
9. The SFTPB promoter fragment of any one of the preceding claims, which comprises or consists of a nucleic acid sequence of any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 49, or a sequence with at least 90% identity to any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 48, preferably wherein said promoter fragment comprises or consists of a nucleic acid sequence of any one of SEQ ID NOs: 46, 47 or 48, or a sequence with at least 90% identity to any one of SEQ ID NOs: 46, 47 or 48.
10. A nucleic acid cassette comprising: (a) an SFTPB promoter fragment as defined in any one of the preceding claims; and
(b) a transgene.
11. The nucleic acid cassette of claim 10, wherein transgene encodes a therapeutic protein selected from: (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP‐B), Surfactant Protein C (SP‐C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte‐Macrophage Colony‐Stimulating Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody, an anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP‐binding cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.
12. The nucleic acid cassette of claim 10 or 11, wherein the SFTPB promoter fragment increases expression of the transgene by lung parenchymal cells, wherein optionally: (a) expression of the transgene by the SFTPB promoter fragment is increased compared with expression of the transgene by the full‐length SFTPB promoter; (b) expression of the transgene by the SFTPB promoter fragment is increased by at least 2‐fold, preferably at least 5‐fold compared expression of the transgene by the full‐ length SFTPB promoter.
13. The nucleic acid cassette of any one of claims 10 to 12, wherein the lung parenchymal cells comprise one or more cell type selected from: alveolar type I epithelial (ATI) cells, alveolar type II epithelial cells (ATII), and/or club cells, preferably ATII and/or ATI cells.
14. The nucleic acid cassette of any one of claims 11 to 13, wherein: (a) expression of the transgene is specific to lung parenchymal cells; and/or (b) the ratio of lung expression: nose expression of the transgene by the SFTPB promoter fragment is at least 2:1.
15. A gene therapy vector, comprising a nucleic acid cassette as defined in any one of the preceding claims.
16. The gene therapy vector of claim 15, which is a non‐viral vector, wherein optionally: (a) the non‐viral vector is a plasmid; and/or (b) the non‐viral vector is comprised in a cationic liposome, which preferably comprises GL67A.
17. The gene therapy vector of claim 15, which is a viral vector, optionally selected from: (a) a lentiviral vector; (b) an AAV vector; and (c) an adenoviral vector.
18. The gene therapy vector of claim 17, which is a lentiviral vector that is pseudotyped with haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, wherein optionally the respiratory paramyxovirus is a Sendai virus.
19. The gene therapy vector of claim 17 or 18, wherein the lentiviral vector is selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector, preferably a SIV vector.
20. A method of expressing a therapeutic protein in a target cell, comprising delivering a nucleic acid cassette as defined in any one of claims 1‐14 or a gene therapy vector as defined in any one of claims 15‐19 into the target cells.
21. The method of claim 20, wherein said delivering comprises integrating said nucleic acid cassette or gene therapy vector into said target cell's genome.
22. A gene therapy vector as defined in any one of claims 15‐19 for use in a method of treating a disease.
23. The gene therapy vector for use of claim 22, wherein the disease is: (a) a genetic disease;
(b) a respiratory disease, particularly a genetic respiratory disease; (c) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder; and/or (d) selected from Surfactant Protein B (SP‐B) Deficiency; Surfactant Protein C (SP‐C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1‐antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID‐19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia.
24. A cell comprising a nucleic acid cassette as defined in any one of claims 1‐14 or a gene therapy vector as defined in any one of claims 15‐19.
25. A composition comprising a nucleic acid cassette as defined in any one of claims 1‐14 or a gene therapy vector as defined in any one of claims 15‐19 and a pharmaceutically acceptable carrier, diluent or excipient.
26. A lentiviral vector pseudotyped with haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered simultaneously or sequentially with a surfactant.
27. The lentiviral vector for use of claim 26, wherein the disease is: (a) a genetic disease; (b) a respiratory disease, particularly a genetic respiratory disease; (c) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder; and/or (d) selected from Surfactant Protein B (SP‐B) Deficiency; Surfactant Protein C (SP‐C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1‐antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory
distress syndrome (ARDS); COVID‐19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia.
28. The lentiviral vector for use of claim 26 or 27, wherein: (a) the respiratory paramyxovirus is a Sendai virus; and/or (b) the lentiviral vector is selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector, preferably a SIV vector 29. The lentiviral vector for use of any one of claims 26 to 28, wherein the transgene encodes a therapeutic protein selected from: (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP‐B), Surfactant Protein C (SP‐C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte‐Macrophage Colony‐Stimulating Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody, an anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP‐binding cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB. 30. The lentiviral vector for use of any one of claims 26 to 29, wherein: (a) the promoter comprises a SFTPB promoter fragment as defined in any one of claims 1 to 9; and/or (b) the lentiviral vector comprises a nucleic acid cassette as defined in any one of claims 10 to 14. 31. The lentiviral vector for use of any one of claims 26 to 30, wherein: (a) the lentiviral vector is administered before the surfactant; (b) the surfactant is administered before the lentiviral vector; or
Ĩc) the lentiviral vector and surfactant are administered simultaneously, optionally wherein the lentiviral vector and surfactant are mixed prior to administration.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GBGB2303328.5A GB202303328D0 (en) | 2023-03-07 | 2023-03-07 | Synthetic promoters |
| PCT/GB2024/050608 WO2024184649A1 (en) | 2023-03-07 | 2024-03-07 | Synthetic promoters |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4676545A1 true EP4676545A1 (en) | 2026-01-14 |
Family
ID=85980181
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24712568.5A Pending EP4676545A1 (en) | 2023-03-07 | 2024-03-07 | Synthetic promoters |
Country Status (6)
| Country | Link |
|---|---|
| EP (1) | EP4676545A1 (en) |
| JP (1) | JP2026509256A (en) |
| CN (1) | CN121219021A (en) |
| AU (1) | AU2024230905A1 (en) |
| GB (1) | GB202303328D0 (en) |
| WO (1) | WO2024184649A1 (en) |
Family Cites Families (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5223409A (en) | 1988-09-02 | 1993-06-29 | Protein Engineering Corp. | Directed evolution of novel binding proteins |
| IL99552A0 (en) | 1990-09-28 | 1992-08-18 | Ixsys Inc | Compositions containing procaryotic cells,a kit for the preparation of vectors useful for the coexpression of two or more dna sequences and methods for the use thereof |
| US5976873A (en) * | 1994-05-18 | 1999-11-02 | Children's Hospital Medical Center | Nucleic acid sequences controlling lung cell-specific gene expression |
| GB201108879D0 (en) | 2011-05-25 | 2011-07-06 | Isis Innovation | Vector |
| GB201118704D0 (en) | 2011-10-28 | 2011-12-14 | Univ Oxford | Cystic fibrosis treatment |
| US20230190871A1 (en) * | 2020-05-20 | 2023-06-22 | Sana Biotechnology, Inc. | Methods and compositions for treatment of viral infections |
| GB202105277D0 (en) | 2021-04-13 | 2021-05-26 | Imperial College Innovations Ltd | Signal peptides |
-
2023
- 2023-03-07 GB GBGB2303328.5A patent/GB202303328D0/en not_active Ceased
-
2024
- 2024-03-07 WO PCT/GB2024/050608 patent/WO2024184649A1/en not_active Ceased
- 2024-03-07 AU AU2024230905A patent/AU2024230905A1/en active Pending
- 2024-03-07 CN CN202480030016.9A patent/CN121219021A/en active Pending
- 2024-03-07 JP JP2025551911A patent/JP2026509256A/en active Pending
- 2024-03-07 EP EP24712568.5A patent/EP4676545A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| AU2024230905A1 (en) | 2025-10-23 |
| CN121219021A (en) | 2025-12-26 |
| JP2026509256A (en) | 2026-03-17 |
| GB202303328D0 (en) | 2023-04-19 |
| WO2024184649A1 (en) | 2024-09-12 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| KR102724029B1 (en) | Modification of mammalian cells using artificial micro-RNAs and compositions of these products for altering the properties of mammalian cells | |
| AU2018338790B2 (en) | Non-human animals comprising a humanized TTR locus and methods of use | |
| ES2640139T3 (en) | Heavy chain mice of restricted immunoglobulin | |
| US20230338477A1 (en) | Anti-tfr:gaa and anti-cd63:gaa insertion for treatment of pompe disease | |
| KR20190057104A (en) | Non-human animal with hexanucleotide repeat extension in C9ORF72 locus | |
| KR20210129108A (en) | Compositions and methods for treating glycogen storage disease type 1A | |
| KR102927708B1 (en) | Non-human animals comprising a humanized TTR locus with a beta-slip mutation and methods of use thereof | |
| KR102915369B1 (en) | CRISPR and AAV Strategies for the Treatment of X-Linked Juvenile Retinoschisis | |
| AU2019403015B2 (en) | Nuclease-mediated repeat expansion | |
| KR20230148824A (en) | Compositions and methods for delivering nucleic acids | |
| KR20230002788A (en) | Artificial expression constructs for selectively modulating gene expression in neocortical layer 5 glutamatergic neurons | |
| US20240197921A1 (en) | Signal peptides | |
| US20240325567A1 (en) | Cell therapy | |
| EP4676545A1 (en) | Synthetic promoters | |
| CN121909216A (en) | Anti-TfR: acid sphingomyelinase for the treatment of acid sphingomyelinase deficiency | |
| CN109337928B (en) | Methods for improving the efficiency of gene therapy by overexpressing adeno-associated virus receptors | |
| US20250304996A1 (en) | Pseudotyped lentiviral vectors | |
| US20250108133A1 (en) | Vectors and compositions for gene augmentation of crumbs complex homologue 1 (crb1) mutations | |
| RU2784927C1 (en) | Animals other than human, including humanized ttr locus, and application methods | |
| RU2833486C1 (en) | Crispr and aav strategies for therapy of x-linked juvenile retinoschisis | |
| CA3208936A1 (en) | Retroviral vectors | |
| CN117836420A (en) | Recombinant TERT-encoding viral genome and vector | |
| KR20260064747A (en) | Anti-TfR:GAA and anti-CD63:GAA insertion for the treatment of Pompe disease | |
| CN114621971A (en) | Genetically modified non-human animal, and construction method and application thereof | |
| CN118679250A (en) | Anti-TfR:GAA and anti-CD63:GAA insertion for the treatment of Pompe disease |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251007 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |