EP4676545A1 - Synthetic promoters - Google Patents

Synthetic promoters

Info

Publication number
EP4676545A1
EP4676545A1 EP24712568.5A EP24712568A EP4676545A1 EP 4676545 A1 EP4676545 A1 EP 4676545A1 EP 24712568 A EP24712568 A EP 24712568A EP 4676545 A1 EP4676545 A1 EP 4676545A1
Authority
EP
European Patent Office
Prior art keywords
seq
sftpb
vector
enhancer
nucleic acid
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP24712568.5A
Other languages
German (de)
French (fr)
Inventor
Deborah Gill
Steve Hyde
Altar MUNIS
Kamran MIAH
Aimee RUFFLE
Rosie MUNDAY
Mariana VIEGAS
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ip2ipo Innovations Ltd
Original Assignee
Imperial College Innovations Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Imperial College Innovations Ltd filed Critical Imperial College Innovations Ltd
Publication of EP4676545A1 publication Critical patent/EP4676545A1/en
Pending legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/63Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
    • C12N15/79Vectors or expression systems specially adapted for eukaryotic hosts
    • C12N15/85Vectors or expression systems specially adapted for eukaryotic hosts for animal cells
    • AHUMAN NECESSITIES
    • A61MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61KPREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
    • A61K48/00Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy
    • A61K48/005Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy characterised by an aspect of the 'active' part of the composition delivered, i.e. the nucleic acid delivered
    • A61K48/0058Nucleic acids adapted for tissue specific expression, e.g. having tissue specific promoters as part of a contruct
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2740/00Reverse transcribing RNA viruses
    • C12N2740/00011Details
    • C12N2740/10011Retroviridae
    • C12N2740/15011Lentivirus, not HIV, e.g. FIV, SIV
    • C12N2740/15032Use of virus as therapeutic agent, other than vaccine, e.g. as cytolytic agent
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2740/00Reverse transcribing RNA viruses
    • C12N2740/00011Details
    • C12N2740/10011Retroviridae
    • C12N2740/15011Lentivirus, not HIV, e.g. FIV, SIV
    • C12N2740/15041Use of virus, viral particle or viral elements as a vector
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2740/00Reverse transcribing RNA viruses
    • C12N2740/00011Details
    • C12N2740/10011Retroviridae
    • C12N2740/15011Lentivirus, not HIV, e.g. FIV, SIV
    • C12N2740/15041Use of virus, viral particle or viral elements as a vector
    • C12N2740/15043Use of virus, viral particle or viral elements as a vector viral genome or elements thereof as genetic vector

Definitions

  • the present invention relates to nucleic acid cassettes for gene therapy, particularly to promoter and promoter/enhancer combinations for improved expression of transgenes in a lung parenchyma ⁇ specific/preferred manner.
  • the invention further relates nucleic acid cassettes comprising said promoters and promoter/enhancer combinations, viral and non ⁇ viral vectors comprising such nucleic acid cassettes, and the use of such nucleic acid cassettes and vectors to increase expression of therapeutic proteins by lung parenchyma cells.
  • SFTPB deficiency is a severe monogenic interstitial lung disorder that leads to loss of life in infants as a result of alveolar collapse and respiratory distress syndrome.
  • the only curative treatment is thought to be lung transplantation; however, the lack of suitable donor organs makes this a non ⁇ viable option in most circumstances.
  • nucleic acids as medicine, or gene therapy, is a promising new treatment modality, both for SFTPB deficiency and other genetic diseases, including genetic diseases of the respiratory tract.
  • gene therapies currently in use or under development are not effective at curing diseases is because it is difficult to make sufficient protein to reach the therapeutic threshold needed to treat or cure the disease.
  • the present inventors have previously developed a lentiviral vector, which has been pseudotyped with hemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, comprising a promoter and a transgene.
  • the backbone of the vector is from a simian immunodeficiency virus (SIV), such as SIV1 or African green monkey SIV (SIV ⁇ AGM).
  • SIV ⁇ AGM African green monkey SIV
  • the backbone of a viral vector of the invention is from SIV ⁇ AGM.
  • the HN and F proteins function, respectively, to attach to sialic acids and mediate cell fusion for vector entry to target cells.
  • the present inventors discovered that this specifically F/HN ⁇ pseudotyped lentiviral vector can efficiently transduce airway epithelium, resulting in transgene expression sustained for periods beyond the proposed lifespan of airway epithelial cells. Importantly, the present inventors also found that re ⁇ administration does not result in a loss of efficacy. These features make the vectors of the present invention attractive candidates for treating diseases via their use in expressing therapeutic proteins: (i) within the cells of the respiratory tract; (ii) secreted into the lumen of the respiratory tract; and (iii) secreted into the circulatory system. However, even using this state ⁇ of ⁇ the ⁇ art platform technology, the levels of transgene expressed are at the lower predicted threshold required for clinical efficacy.
  • exogenous signal peptides can be used to increase expression and secretion of therapeutic proteins by airway cells.
  • exogenous signal peptides it is possible to produce more protein for every copy of a gene therapy vector or transgene that is put into a cell, increasing the dose of therapeutic protein without increasing the amount of gene therapy vector given to a patient.
  • exogenous signal peptides is not appropriate for all therapeutic proteins or all conditions. For example, not all therapeutic proteins are secreted, and for some there may be clinical reasons why manipulating the signal peptides is undesirable.
  • gene therapy vectors comprising promoters providing high and sustained gene expression in a variety of cell types are preferred, especially in a therapeutic context. For this reason, the inventors previously used a hCEF promoter to drive strong and persistent expression in mouse lung with non ⁇ viral formulations. However for some conditions, expression in a specific tissue or cell type is required to (i) achieve desired therapeutic target, (ii) avoid gene expression ⁇ related toxicities, and (iii) circumvent immune responses to the therapeutic agent stemming from gene expression in undesired cell types.
  • Such promoters, cassettes and vectors may be of particular use in the treatment of genetic diseases, particularly genetic respiratory diseases, such as surfactant protein deficiency.
  • genetic diseases particularly genetic respiratory diseases, such as surfactant protein deficiency.
  • SUMMARY OF THE INVENTION At present, there remains a pressing need for technology that enables high levels of cell ⁇ or tissue ⁇ specific expression of transgenes for gene therapy, including from the inventors’ own lentiviral platform.
  • novel promoters that drive cell ⁇ specific expression in the lung parenchyma.
  • the present inventors have developed a panel of novel promoter sequences comprising a functional fragment of the human SFTPB gene promoter.
  • the inventors have surprisingly shown that the SFTPB promoter fragment of the invention drives increased transgene expression compared with the full ⁇ length SFTPB promoter. Furthermore, the inventors have also surprisingly demonstrated that the SFTPB promoter fragment of the invention can be combined with particular enhancers, including some enhancers which are not associated with cell ⁇ specific expression in the lung parenchyma, to further improve transgene expression. Accordingly, the present invention provides an SFTPB promoter fragment which comprises or consists of at least part of exon 1 of the SFTPB gene, but does not comprise the SFTPB gene start codon. Typically said promoter is less than 800 bases in length, preferably less than 700 bases in length.
  • Said promoter may comprise or consist of: (a) SEQ ID NO: 1 or a sequence with at least 80% identity to SEQ ID NO: 1; or (b) bases 81 ⁇ 710 of SEQ ID NO: 2, or a sequence with at least 80% identity to bases 81 ⁇ 710 of SEQ ID NO: 2; wherein optionally (i) said SFTPB promoter fragment further comprises up to 20 bases at the 5’ end, which may optionally correspond to up to 20 bases 5’ to base 81 of SEQ ID NO: 2; and/or (ii) said SFTPB promoter fragment further comprises up to 4 bases at the 3’ end, which may optionally correspond to up to 4 bases 3’ to base 710 of SEQ ID NO: 2.
  • the SFTPB promoter fragment may comprise or consist of SEQ ID NO: 1, or a sequence with at least 90% identity to SEQ ID NO: 1.
  • the SFTPB promoter fragment of the invention may further comprise a 5’ enhancer.
  • Said enhancer may be in (i) the forward, or (ii) the reverse, orientation; preferably wherein the enhancer is in the forward orientation.
  • Said enhancer may be selected from a SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer, an ELF3 enhancer, an actin enhancer, an LMO7 enhancer, a SFTPC enhancer or a SFTPB enhancer.
  • the enhancer may be selected from a SLC34A2 enhancer, a VEGFA enhancer or a CMV enhancer.
  • the enhancer may be selected from (a) an SLC34A2 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 29 ⁇ 33, preferably SEQ ID NO: 31; (b) a VEGFA enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38; (c) a CMV enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11; (d) a SV40 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or 35; (e) an ELF3 enhancer which comprises or consists of a nucleotide sequence
  • the SFTPB promoter fragment of the invention may comprise or consist of a nucleic acid sequence of any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 49, or a sequence with at least 90% identity to any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 48.
  • the SFTPB promoter fragment of the invention comprises or consists of a nucleic acid sequence of any one of SEQ ID NOs: 46, 47 or 48, or a sequence with at least 90% identity to any one of SEQ ID NOs: 46, 47 or 48.
  • the invention also provides a nucleic acid cassette comprising: (a) an SFTPB promoter fragment of the invention; and (b) a transgene.
  • Said transgene may encode a therapeutic protein selected from: (a) secreted therapeutic protein selected from: Surfactant Protein B (SP ⁇ B), Surfactant Protein C (SP ⁇ C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF), decorin, an anti ⁇ inflammatory protein (e.g. IL ⁇ 10 or TGF ⁇ ) or monoclonal antibody, an anti ⁇ inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP ⁇ binding cassette sub ⁇ family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.
  • a therapeutic protein selected from: (a) secreted therapeutic protein selected from: Surfactant Protein B (SP ⁇ B), Surfactant Protein C (SP ⁇ C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor
  • the SFTPB promoter fragment may increase expression of the transgene by lung parenchymal cells, optionally compared with the full ⁇ length SFTPB promoter. Expression of the transgene may be increased by at least 2 ⁇ fold, preferably at least 5 ⁇ fold compared with the full ⁇ length SFTPB promoter.
  • the lung parenchymal cells may comprise one or more cell type selected from: alveolar type I epithelial (ATI) cells, alveolar type II epithelial cells (ATII), and/or club cells, preferably ATII and/or ATI cells.
  • Expression of the transgene by a nucleic acid cassette or promoter of the invention may be specific to lung parenchymal cells; and/or the ratio of lung expression: nose expression of the transgene by the SFTPB promoter fragment is at least 2:1.
  • the invention further provides a gene therapy vector, comprising a nucleic acid cassette of the invention.
  • Said gene therapy vector may be a non ⁇ viral vector, wherein optionally: (a) the non ⁇ viral vector is a plasmid; and/or (b) the non ⁇ viral vector is comprised in a cationic liposome, which preferably comprises GL67A.
  • Said gene therapy vector may be a viral vector, optionally selected from: a lentiviral vector; an AAV vector; and an adenoviral vector.
  • Said lentiviral vector may be pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, optionally from a Sendai virus.
  • Said lentiviral vector may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector.
  • HIV Human immunodeficiency virus
  • SIV Simian immunodeficiency virus
  • FV Feline immunodeficiency virus
  • EIAV Equine infectious anaemia virus
  • Visna/maedi virus vector a Visna/maedi virus vector.
  • said lentiviral vector is a SIV vector.
  • the invention further provides a method of expressing a therapeutic protein in a target cell, comprising delivering a nucleic acid cassette of the invention or a gene therapy vector of the invention into the target cells.
  • Said delivering may comprise integrating said nucleic acid cassette or gene therapy vector into said target cell's genome.
  • the invention also provides a gene therapy vector of the invention for use in a method of treating a disease.
  • Said disease may be a genetic disease.
  • the disease may be: (a) a respiratory disease, particularly a genetic respiratory disease; or (b) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder.
  • the disease may be selected from Surfactant Protein B (SP ⁇ B) Deficiency; Surfactant Protein C (SP ⁇ C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1 ⁇ antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID ⁇ 19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia.
  • SP ⁇ B Surfactant Protein B
  • SP ⁇ C Surfactant Protein C
  • Pulmonary surfactant metabolism dysfunction 2 SMDP2
  • Pulmonary surfactant metabolism dysfunction 3 SMDP3
  • PCD Primary Cilia
  • the invention further provides a cell comprising a nucleic acid cassette of the invention or a gene therapy vector of the invention.
  • the invention also provides a composition comprising a nucleic acid cassette of the invention or a gene therapy vector of the invention and a pharmaceutically acceptable carrier, diluent or excipient.
  • lentiviral vectors such as the lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus as exemplified herein
  • HN haemagglutinin ⁇ neuraminidase
  • F fusion proteins from a respiratory paramyxovirus
  • the invention also provides lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered simultaneously or sequentially with a surfactant.
  • HN haemagglutinin ⁇ neuraminidase
  • F fusion
  • said disease is: (a) a genetic disease; (b) a respiratory disease, particularly a genetic respiratory disease; (c) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder; and/or (d) selected from Surfactant Protein B (SP ⁇ B) Deficiency; Surfactant Protein C (SP ⁇ C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1 ⁇ antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID ⁇ 19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia.
  • SP ⁇ B Surfactant Protein B
  • SP ⁇ C
  • the respiratory paramyxovirus may be a Sendai virus; and/or the lentiviral vector may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector, preferably a SIV vector
  • the transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP ⁇ B), Surfactant Protein C (SP ⁇ C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF), decorin, an anti ⁇ inflammatory protein (e.g.
  • the promoter may comprise a SFTPB promoter fragment as defined herein; and/or the lentiviral vector may comprise a nucleic acid cassette as defined herein.
  • the lentiviral vector may be administered before the surfactant.
  • the surfactant may be administered before the lentiviral vector.
  • the lentiviral vector and surfactant may be administered simultaneously, optionally wherein the lentiviral vector and surfactant are mixed prior to administration.
  • FIG. 2 Graphs quantifying EGFP expression in human HEK293T cells transduced with recombinant SIV lentiviral vectors pseudotyped with VSV ⁇ G and expressing the EGFP transgene from a range of promoters (CMV (SEQ ID NO: 6), hCEF ( (SEQ ID NO: 5), EF1aS ( (SEQ ID NO: 7), PGK ( (SEQ ID NO: 8), mSPB ( (SEQ ID NO: 1), fSPB (SEQ ID NO: 4), mSPC (SEQ ID NO: 9), fSPC(SEQ ID NO: 10)) at a multiplicity of infection (moi) of 0.5 or 1.0.
  • CMV SEQ ID NO: 6
  • hCEF (SEQ ID NO: 5)
  • EF1aS (SEQ ID NO: 7)
  • PGK (SEQ ID NO: 8)
  • mSPB (SEQ ID NO: 1), fSPB (SEQ
  • FIG. 3 Experimental schematics for experiments to test EGFP expression in vivo using lentiviral vectors expressing EGFP from different promoters. A Schematic showing timing of dosing, on Day 0 mice were dosed with lentiviral vectors via nasal instillation and 7 days post ⁇ dosing the mice were culled and lung tissue harvested for cryosections. B Schematic showing treatment groups.
  • FIG. 5 Panel of ATII specific genes (and enhancers) generated by interrogation of LungGENS and a tissue expression database.
  • Figure 6 Schematics of exemplary mSPB/enhancer constructs. Different lengths (indicated in base pairs (bp)) of the newly identified enhancer sequences (boxes) were sub ⁇ cloned in front of the mSPB (SEQ ID NO: 1) promoter sequence (arrow boxes) in both forward (f) and reverse (r) orientations to generate candidate synthetic promoters expressing the EGFP2ALux reporter transgene.
  • FIG. 7 Graph showing expression of EGFP in HEK293T cells by different mSPB/enhancer constructs.
  • Figure 8 Graph showing expression of EGFP in murine LA ⁇ 4 cells by different mSPB/enhancer constructs. P values are given where significant expression was observed.
  • FIG. 9 Graph showing expression of EGFP in human SALI cells by different mSPB/enhancer constructs. P values are given where significant expression was observed.
  • RLU Relative Light Units
  • Figure 10 Graph showing expression of EGFP in human SALI cells by lentiviral vectors with different mSPB/enhancers in a secondary screen. P values are given where significant expression was observed.
  • CMV SEQ ID NO: 6
  • hCEF SEQ ID NO: 5
  • mSPB (SEQ ID NO: 1)
  • Luciferase expression levels (Relative Light Units (RLU)) were determined and compared with na ⁇ ve (non ⁇ transduced) control cells.
  • Figure 11 Graph showing expression of EGFP in HEK293T cells by lentiviral vectors with different mSPB/enhancers in a secondary screen. P values are given where significant expression was observed.
  • CMV SEQ ID NO: 6
  • hCEF SEQ ID NO: 5
  • Luciferase expression levels (Relative Light Units (RLU)) were determined and compared with na ⁇ ve (non ⁇ transduced) control cells.
  • Figure 12 Heat map of luciferase expression in mice treated with lentiviral vectors with different mSPB/enhancers.
  • CMV SEQ ID NO: 6
  • hCEF SEQ ID NO: 5
  • mSPB (SEQ ID NO: 1)
  • Luciferase signal was observed in the nose and lung areas at all timepoints with the non ⁇ specific CMV[SEQ ID NO: 6] and hCEF[ SEQ ID NO: 5] lung promoter in line with expectations. Luciferase signal in the mSPB[SEQ IDN O: 1] group was overall lower and restricted to the lung, also as expected.
  • Figure 13 Graphs showing in vivo luciferase expression by lentiviral vectors with different mSPB/enhancers (A) and the ratio of luciferase expression in the lungs and the nose (B).
  • Lentivirus was dosed intranasally (1x10 7 TU per mouse in 100uL TSSM Buffer). On days 2, 7 and 28 post ⁇ dosing, mice were anaesthetised and imaged for luciferase expression.
  • Figure 14 Graphs showing luciferase expression by plasmids with different mSPB/enhancers in human SALI cells (A), in vivo (B), and in HEK293T cells (C).
  • the enhancer/mSPB promoter constructs expressing Lux reporter transgene were evaluated in the context of a non ⁇ viral formulation, to deliver plasmids containing the CMV[SEQ ID NO: 6], hCEF[SEQ ID NO: 5] and mSPB[SEQ ID NO: 1] promoter sequences as well as the selected enhancer/mSPB promoter constructs: slc2[SEQ ID NO: 31], veg2[SEQ ID NO: 38], sv40[SEQ ID NO: 34], and CMVenh[SEQ ID NO: 11].
  • FIG. 15 Representative immunohistochemistry images showing EGFP expression in lung parenchyma sections taken from mice treated with lentiviral vectors expressing EGFP from different enhancer/mSPB constructs. Representative images from (A) CMVenh (Alv ⁇ 01, SEQ ID NO: 46), (B) Slc2 (Alv ⁇ 02, SEQ ID NO: 47) and (C) Vegf2 (Alv ⁇ 03, SEQ ID NO: 48) groups are shown.
  • mice were administered D ⁇ luciferin and imaged for luciferase activity in the lung and nasal cavity.
  • Signal in regions of interest ROI
  • ROI Signal in regions of interest
  • AUC area under the curve
  • C Average radiance in the lung
  • D Specificity for expression in the lung parenchyma was determined by calculating the ratio of signal in the lung (indicative of alveolar and airway cell transduction) to signal in the nasal cavity (indicative of airway cell transduction) 28 days after dosing.
  • Representative images of EGFP positive cells in the airway (Aw) and parenchyma (P) are shown after administration of (A) rSIV.F/HN hCEF EGFP or (B) rSIV.F/HN mSP ⁇ B EGFP. Images of lung sections were further analysed (using Visiopharm software) which required manual indication of airways and parenchyma. The percentage of EGFP ⁇ positive cells from the total lung and from the airways was used to (C) estimate the percentage of EGFP ⁇ positive cells observed in the parenchyma from each promoter.
  • FIG. 18 The effect of mSPB, Alv ⁇ 1, Alv ⁇ 2 and Alv ⁇ 3 driven SFP ⁇ B expression on transepithelial electrical resistance (TEER) in a Surfactant Air Liquid Interface (SALI) model.
  • TEER values from SALI cultures generated from the H441 SP ⁇ B KO cells were lower than the parental H441 cell line at 14 days post airlift constituting a phenotypic defect.
  • “capable of interacting” also means interacting
  • “capable of cleaving” also means cleaves
  • “capable of binding” also means binds and "capable of specifically targeting" also means specifically targets.
  • Numeric ranges are inclusive of the numbers defining the range. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Each smaller range between any stated value or intervening value in a stated range and any other stated or intervening value in that stated range is encompassed within this disclosure.
  • “About” may generally mean an acceptable degree of error for the quantity measured given the nature or precision of the measurements. Exemplary degrees of error are within 20 percent (%), typically, within 10%, and more typically, within 5% of a given value or range of values. Preferably, the term “about” shall be understood herein as plus or minus ( ⁇ ) 5%, preferably ⁇ 4%, ⁇ 3%, ⁇ 2%, ⁇ 1%, ⁇ 0.5%, ⁇ 0.1%, of the numerical value of the number with which it is being used.
  • the term “consisting essentially of''” refers to those elements required for a given invention. The term permits the presence of elements that do not materially affect the basic and novel or functional characteristic(s) of that invention (i.e. inactive or non ⁇ immunogenic ingredients).
  • Embodiments described herein as “comprising” one or more features may also be considered as disclosure of the corresponding embodiments “consisting of” and/or “consisting essentially of” such features. Concentrations, amounts, volumes, percentages and other numerical values may be presented herein in a range format.
  • a "vector” or “construct” refers to a macromolecule or complex of molecules comprising a polynucleotide to be delivered to a host cell, either in vitro or in vivo.
  • a vector can be a linear or a circular molecule.
  • a vector of the invention may be viral or non ⁇ viral.
  • lentiviral vectors refers to a common type of non ⁇ viral vector.
  • a plasmid is an extra ⁇ chromosomal DNA molecule separate from the chromosomal DNA which is capable of replicating independently of the chromosomal DNA.
  • a plasmid is circular and may be double ⁇ stranded.
  • the terms "nucleic acid cassette”, “nucleic acid construct”, “expression cassette” and “nucleic acid expression cassette” are used interchangeably to mean a nucleic acid molecule that is capable of directing transcription.
  • a nucleic acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter and a nucleic acid sequence to be transcribed.
  • a nucleic acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter and a nucleic acid sequence encoding a protein of interest.
  • a nucleic acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter, and a nucleic acid encoding a therapeutic protein.
  • a nucleic acid cassette may include additional elements, such as an enhancer, and/or a transcription termination signal.
  • the terms “transduced” and “modified” are used interchangeably to describe cells which have been modified to express a transgene of interest. Typically the modification occurs through transduction of the cells.
  • Titre and “yield” are used interchangeably to mean the amount of viral (e.g. lentiviral, particularly SIV) vector produced by a method of the invention.
  • Titre is the primary benchmark characterising manufacturing efficiency, with higher titres generally indicating that more vector is manufactured (e.g. using the same amount of reagents).
  • Titre or yield may relate to the number of vector genomes that have integrated into the genome of a target cell (integration titre), which is a measure of “active” virus particles, i.e. the number of particles capable of transducing a cell.
  • Transducing units (TU/mL also referred to as TTU/mL) is a biological readout of the number of host cells that get transduced under certain tissue culture/virus dilutions conditions, and is a measure of the number of “active” virus particles.
  • the total number of (active+inactive) virus particles may also be determined using any appropriate means, such as by measuring either how much Gag is present in the test solution or how many copies of viral RNA are in the test solution. Assumptions are then made that a viral (e.g. lentivirus, particularly SIV) particle contains either 2000 Gag molecules or 2 viral RNA molecules. Once total particle number and a transducing titre/TU have been measured, a particle:infectivity ratio calculated.
  • amino acids are referred to herein using the name of the amino acid, the three ⁇ letter abbreviation or the single letter abbreviation. Unless otherwise indicated, any nucleic acid sequences are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively.
  • protein and “polypeptide” are used interchangeably herein to designate a series of amino acid residues, connected to each other by peptide bonds between the alpha ⁇ amino and carboxyl groups of adjacent residues.
  • protein refers to a polymer of amino acids, including modified amino acids (e.g., phosphorylated, glycated, glycosylated, etc.) and amino acid analogues, regardless of its size or function.
  • modified amino acids e.g., phosphorylated, glycated, glycosylated, etc.
  • amino acid analogues regardless of its size or function.
  • Protein and polypeptide are often used in reference to relatively large polypeptides, whereas the term “peptide” is often used in reference to small polypeptides, but usage of these terms in the art overlaps.
  • protein and “polypeptide” are used interchangeably herein when referring to a gene product and fragments thereof.
  • polypeptides or proteins include gene products, naturally occurring proteins, homologs, orthologs, paralogs, fragments and other equivalents, variants, fragments, and analogues of the foregoing.
  • polynucleotides refers to any molecule, preferably a polymeric molecule, incorporating units of ribonucleic acid, deoxyribonucleic acid or an analogue thereof.
  • the nucleic acid can be either single ⁇ stranded or double ⁇ stranded.
  • a single ⁇ stranded nucleic acid can be one nucleic acid strand of a denatured double ⁇ stranded DNA Alternatively, it can be a single ⁇ stranded nucleic acid not derived from any double ⁇ stranded DNA.
  • the nucleic acid can be DNA.
  • the nucleic acid can be RNA Suitable nucleic acid molecules are DNA, including genomic DNA or cDNA. Other suitable nucleic acid molecules are RNA, including siRNA, shRNA, and antisense oligonucleotides.
  • transgene and “gene” are also used interchangeably and both terms encompass fragments or variants thereof encoding the target protein.
  • transgenes of the present invention include nucleic acid sequences that have been removed from their naturally occurring environment, recombinant or cloned DNA isolates, and chemically synthesized analogues or analogues biologically synthesized by heterologous systems. Minor variations in the amino acid sequences of the invention are contemplated as being encompassed by the present invention, providing that the variations in the amino acid sequence(s) maintain at least 60%, at least 70%, more preferably at least 80%, at least 85%, at least 90%, at least 95%, and most preferably at least 97% or at least 99% sequence identity to the amino acid sequence of the invention or a fragment thereof as defined anywhere herein.
  • homology is used herein to mean identity.
  • sequence of a variant or analogue sequence of an amino acid sequence of the invention may differ on the basis of substitution (typically conservative substitution) deletion or insertion. Proteins comprising such variations are referred to herein as variants. Proteins of the invention may include variants in which amino acid residues from one species are substituted for the corresponding residue in another species, either at the conserved or non ⁇ conserved positions. Variants of protein molecules disclosed herein may be produced and used in the present invention. Following the lead of computational chemistry in applying multivariate data analysis techniques to the structure/property ⁇ activity relationships [see for example, Wold, et al. Multivariate data analysis in chemistry. Chemometrics ⁇ Mathematics and Statistics in Chemistry (Ed.: B. Kowalski); D.
  • proteins can be derived from empirical and theoretical models (for example, analysis of likely contact residues or calculated physicochemical property) of proteins sequence, functional and three ⁇ dimensional structures and these properties can be considered individually and in combination.
  • Amino acids are referred to herein using the name of the amino acid, the three ⁇ letter abbreviation or the single letter abbreviation.
  • the term “protein”, as used herein, includes proteins, polypeptides, and peptides.
  • amino acid sequence is synonymous with the term “polypeptide” and/or the term “protein”.
  • amino acid sequence is synonymous with the term “peptide”.
  • the terms "protein” and "polypeptide” are used interchangeably herein.
  • the conventional one ⁇ letter and three ⁇ letter codes for amino acid residues may be used.
  • the 3 ⁇ letter code for amino acids as defined in conformity with the IUPACIUB Joint Commission on Biochemical Nomenclature (JCBN). It is also understood that a polypeptide may be coded for by more than one nucleotide sequence due to the degeneracy of the genetic code. Amino acid residues at non ⁇ conserved positions may be substituted with conservative or non ⁇ conservative residues. In particular, conservative amino acid replacements are contemplated.
  • a “conservative amino acid substitution” is one in which the amino acid residue is replaced with an amino acid residue having a similar side chain.
  • conservatively modified variants in a protein of the invention does not exclude other forms of variant, for example polymorphic variants, interspecies homologs, and alleles.
  • Non ⁇ conservative amino acid substitutions include those in which (i) a residue having an electropositive side chain (e.g., Arg, His or Lys) is substituted for, or by, an electronegative residue (e.g., Glu or Asp), (ii) a hydrophilic residue (e.g., Ser or Thr) is substituted for, or by, a hydrophobic residue (e.g., Ala, Leu, Ile, Phe or Val), (iii) a cysteine or proline is substituted for, or by, any other residue, or (iv) a residue having a bulky hydrophobic or aromatic side chain (e.g., Val, His, Ile or Trp) is substituted for, or by, one having a smaller side chain (e.g., Ala or Ser) or no side chain (e.g., Gly).
  • an electropositive side chain e.g., Arg, His or Lys
  • an electronegative residue e.g., Glu or As
  • “Insertions” or “deletions” are typically in the range of about 1, 2, or 3 amino acids. The variation allowed may be experimentally determined by systematically introducing insertions or deletions of amino acids in a protein using recombinant DNA techniques and assaying the resulting recombinant variants for activity. This does not require more than routine experiments for a skilled person.
  • a “fragment” of a polypeptide comprises at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 97% or more of the original polypeptide.
  • a fragment may comprise at least 5, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20 or more amino acids of the protein from which it is derived.
  • a fragment may be continuous or discontinuous, preferably continuous.
  • the polynucleotides of the present invention may be prepared by any means known in the art. For example, large amounts of the polynucleotides may be produced by replication in a suitable host cell.
  • the natural or synthetic DNA fragments coding for a desired fragment will be incorporated into recombinant nucleic acid constructs, typically DNA constructs, capable of introduction into and replication in a prokaryotic or eukaryotic cell.
  • DNA constructs will be suitable for autonomous replication in a unicellular host, such as yeast or bacteria, but may also be intended for introduction to and integration within the genome of a cultured insect, mammalian, plant or other eukaryotic cell lines.
  • the polynucleotides of the present invention may also be produced by chemical synthesis, e.g. by the phosphoramidite method or the tri ⁇ ester method, and may be performed on commercial automated oligonucleotide synthesizers.
  • a double ⁇ stranded fragment may be obtained from the single stranded product of chemical synthesis either by synthesizing the complementary strand and annealing the strand together under appropriate conditions or by adding the complementary strand using DNA polymerase with an appropriate primer sequence.
  • the term “isolated” in the context of the present invention denotes that the polynucleotide sequence has been removed from its natural genetic milieu and is thus free of other extraneous or unwanted coding sequences (but may include naturally occurring 5' and 3' untranslated regions such as promoters and terminators), and is in a form suitable for use within genetically engineered protein production systems. Such isolated molecules are those that are separated from their natural environment. In view of the degeneracy of the genetic code, considerable sequence variation is possible among the polynucleotides of the present invention.
  • Degenerate codons encompassing all possible codons for a given amino acid are set forth below: Amino Acid Codons Degenerate Codon Cys TGC TGT TGY Ser AGC AGT TCA TCC TCG TCT WSN Thr ACA ACC ACG ACT ACN Pro CCA CCC CCG CCT CCN Ala GCA GCC GCG GCT GCN Gly GGA GGC GGG GGT GGN Asn AAC AAT AAY Asp GAC GAT GAY Glu GAA GAG GAR Gln CAA CAG CAR His CAC CAT CAY Arg AGA AGG CGA CGC CGG CGT MGN Lys AAA AAG AAR Met ATG ATG Ile ATA ATC ATT ATH Leu CTA CTC CTG CTT TTA TTG YTN Val GTA GTC GTG GTT GTN Phe TTC TTT TTY Tyr TAC TAT TAY Trp TGG TGG Ter TAA TAG TGA TRR Asn/ Asp RAY Glu
  • variant amino acid sequences may encode variant amino acid sequences, but one of ordinary skill in the art can easily identify such variant sequences by reference to the amino acid sequences of the present invention.
  • a “variant” nucleic acid sequence has substantial homology or substantial similarity to a reference nucleic acid sequence (or a fragment thereof).
  • a nucleic acid sequence or fragment thereof is “substantially homologous” (or “substantially identical”) to a reference sequence if, when optimally aligned (with appropriate nucleotide insertions or deletions) with the other nucleic acid (or its complementary strand), there is nucleotide sequence identity in at least about 70%, 75%, 80%, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99 or more% of the nucleotide bases. Methods for homology determination of nucleic acid sequences are known in the art.
  • a “variant” nucleic acid sequence is substantially homologous with (or substantially identical to) a reference sequence (or a fragment thereof) if the “variant” and the reference sequence they are capable of hybridizing under stringent (e.g. highly stringent) hybridization conditions.
  • Nucleic acid sequence hybridization will be affected by such conditions as salt concentration (e.g. NaCl), temperature, or organic solvents, in addition to the base composition, length of the complementary strands, and the number of nucleotide base mismatches between the hybridizing nucleic acids, as will be readily appreciated by those skilled in the art.
  • Stringent temperature conditions are preferably employed, and generally include temperatures in excess of 30°C, typically in excess of 37°C and preferably in excess of 45°C.
  • Stringent salt conditions will ordinarily be less than 1000 mM, typically less than 500 mM, and preferably less than 200 mM.
  • the pH is typically between 7.0 and 8.3.
  • Methods of determining nucleic acid percentage sequence identity are known in the art. By way of example, when assessing nucleic acid sequence identity, a sequence having a defined number of contiguous nucleotides may be aligned with a nucleic acid sequence (having the same number of contiguous nucleotides) from the corresponding portion of a nucleic acid sequence of the present invention.
  • Tools known in the art for determining nucleic acid percentage sequence identity include Nucleotide BLAST (as described below).
  • preferential codon usage refers to codons that are most frequently used in cells of a certain species, thus favouring one or a few representatives of the possible codons encoding each amino acid.
  • the amino acid threonine (Thr) may be encoded by ACA, ACC, ACG, or ACT, but in mammalian host cells ACC is the most commonly used codon; in other species, different codons may be preferential.
  • Preferential codons for a particular host cell species can be introduced into the polynucleotides of the present invention by a variety of methods known in the art.
  • any nucleic acid sequence may be codon ⁇ optimised for expression in a host or target cell.
  • the vector genome or corresponding plasmid
  • the REV gene or corresponding plasmid
  • the fusion protein (F) gene or correspond plasmid
  • the hemagglutinin ⁇ neuraminidase (HN) gene or corresponding plasmid, or any combination thereof may be codon ⁇ optimised.
  • a “fragment” of a polynucleotide of interest comprises a series of consecutive nucleotides from the sequence of said full ⁇ length polynucleotide.
  • a “fragment” of a polynucleotide of interest may comprise (or consist of) at least 600 consecutive nucleotides from the sequence of said polynucleotide (e.g. at least 600, 650, 700, 750, 800 850, 900, or 950 consecutive nucleic acid residues of said polynucleotide).
  • a fragment as defined herein retains the same function as the full ⁇ length polynucleotide.
  • the terms “decrease”, “reduced”, “reduction”, or “inhibit” are all used herein to mean a decrease by a statistically significant amount.
  • the terms “reduce,” “reduction” or “decrease” or “inhibit” typically means a decrease by at least 10% as compared to a reference level (e.g.
  • the terms “increased”, “increase”, “enhance”, or “activate” are all used herein to mean an increase by a statically significant amount.
  • the terms “increased”, “increase”, “enhance”, or “activate” can mean an increase of at least 25%, at least 50% as compared to a reference level, for example an increase of at least about 50%, or at least about 75%, or at least about 80%, or at least about 90%, at least about 95%, or at least about 98%, or at least about 99%, or at least about 100%, or at least about 250% or more compared with a reference level, or at least about a 1.5 ⁇ fold, or at least about a 2 ⁇ fold, or at least about a 2.5 ⁇ fold, or at least about a 3 ⁇ fold, or at least about a 4 ⁇ fold, or at least about a 5 ⁇ fold or at least about a 10 ⁇ fold increase, or any increase between 1.5 ⁇ fold and 10 ⁇ fold or greater as compared to a reference level.
  • an “increase” is an observable or statistically significant increase in such level.
  • the terms “individual”, “subject”, and “patient”, are used interchangeably herein to refer to a mammalian subject for whom diagnosis, prognosis, disease monitoring, treatment, therapy, and/or therapy optimisation is desired.
  • the mammal can be (without limitation) a human, non ⁇ human primate, mouse, rat, dog, cat, horse, or cow.
  • the individual, subject, or patient is a human.
  • An “individual” may be an adult, juvenile or infant.
  • An “individual” may be male or female.
  • a "subject in need" of treatment for a particular condition can be an individual having that condition, diagnosed as having that condition, or at risk of developing that condition.
  • a subject can be one who has been previously diagnosed with or identified as suffering from or having a condition in need of treatment or one or more complications or symptoms related to such a condition, and optionally, have already undergone treatment for a condition as defined herein or the one or more complications or symptoms related to said condition.
  • a subject can also be one who has not been previously diagnosed as having a condition as defined herein or one or more or symptoms or complications related to said condition.
  • a subject can be one who exhibits one or more risk factors for a condition, or one or more or symptoms or complications related to said condition or a subject who does not exhibit risk factors.
  • the term “healthy individual” refers to an individual or group of individuals who are in a healthy state, e.g. individuals who have not shown any symptoms of the disease, have not been diagnosed with the disease and/or are not likely to develop the disease e.g. cystic fibrosis (CF) or any other disease described herein).
  • CF cystic fibrosis
  • Preferably said healthy individual(s) is not on medication affecting CF and has not been diagnosed with any other disease.
  • the one or more healthy individuals may have a similar sex, age, and/or body mass index (BMI) as compared with the test individual.
  • BMI body mass index
  • Application of standard statistical methods used in medicine permits determination of normal levels of expression in healthy individuals, and significant deviations from such normal levels.
  • control and “reference population” are used interchangeably.
  • SFTPB Surfactant Protein B promoter fragments
  • SFTPB Surfactant Protein B
  • SFTPB Surfactant Protein B
  • the full ⁇ length SFTPB (surfactant protein B; SFTPB) promoter has previously been defined as a sequence that could be amplified from human genomic DNA with the PCR primers ATTTGAGCTCTTCTTTCTGCTGAACCATCG (sense, SEQ ID NO: 49) and TCTTAGATCTGTCAGACAGCTCTGGGTTCC (antisense, SEQ ID NO: 50), wherein the underlines sequences align with GenBank NCBI Reference Sequence: NG_016967.1 (version 1, accessed 18 November 2022) while the additional 5’ sequences provide SacI (sense) and BglII (antisense) restriction enzyme sites.
  • the forward primer binds to the sense strand from bases 4845 to 4864 in NG_016967.1.
  • the reverse primer binds to the reverse complement of bases 5797 to 5816 in NG_016967.1.
  • the primer pair define a 972 bp genomic fragment.
  • the 972 bp genomic fragment includes all of exon 1 of the SFTPB gene (bases 5543 to 5565 of NG_016967.1), wherein the A at base 5543 is reported as the starting nucleotide of the mRNA generated by the SFTPB promoter and the initiating ATG at bases 5559 to 5561 of NG_016967.1 encodes the first methionine of pre ⁇ pro ⁇ SFTPB.
  • This 972 bp genomic fragment further includes a part of intron 1 from bases 5626 to 5816 of NG_016967.1.
  • references herein to a full ⁇ length SFTPB promoter refer specifically to this 972 bp genomic fragment, which is present SEQ ID NO: 2.
  • the present inventors have identified and isolated functional fragments of the SFTPB (surfactant protein B) gene promoter.
  • SFTPB promoter fragments generated by the inventors are able to drive cell ⁇ specific expression of transgenes in the lung parenchyma.
  • the fragments of the SFTPB gene promoter are shorter than the 972 bp genomic fragment previously identified, i.e.
  • the SFTPB promoter fragments of the invention may be also referred to as a core SFTPB promoters.
  • the SFTPB promoter fragments of the invention surprisingly increase transgene expression compared with the full ⁇ length SFTPB promoter, and can do so in a lung parenchymal cell preferred/specific manner.
  • a promoter of the invention is an SFTPB promoter fragment as described herein. Said SFTPB promoter fragment is functional, also as described herein.
  • an SFTPB promoter fragment of the invention comprises or consists of a core SFTPB promoter fragment, as described herein, or a variant thereof.
  • An SFTPB promoter fragment of the invention may comprise or consist of a fragment of SEQ ID NO: 2, which lack all or part of SFTPB intron 1.
  • SFTPB intron 1 begins at base 5626 of NG_016967.1 (corresponding to residue 782 of SEQ ID NO: 2) and corresponds to bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention may comprise part (but not all) of the SFTPB intron 1.
  • An SFTPB promoter fragment of the invention may not comprise one or more bases from bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention may comprise fewer than 180, fewer than 170, fewer than 160, fewer than 150, fewer than 140, fewer than 130, fewer than 120, fewer than 110, fewer than 100, fewer than 100, fewer than 90, fewer than 80, fewer than 70, fewer than 60, fewer than 50, fewer than 40, fewer than 30, fewer than 20, or fewer than 10 contiguous bases from 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention does not comprise bases 5626 to 5816 of NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2).
  • the SFTPB promoter fragments of the invention do not comprise intron 1 of the SFTPB gene (or any portion thereof).
  • an SFTPB promoter fragment of the invention comprises at least part of exon 1 of the SFTPB gene.
  • SFTPB exon 1 begins at base 5543 of NG_016967.1 (corresponding to residue 699 of SEQ ID NO: 2) and corresponds to bases 5543 to 5565 of NG_016967.1 (corresponding to residues 699 to 781 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention may comprise the contiguous nucleotide sequence of bases 5543 to 5565 of NG_016967.1 (corresponding to residues 699 to 781 of SEQ ID NO: 2).
  • An SFTPB promoter fragment of the invention may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 bases, preferably 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 bases, 3’ to base 5542 of NG_016967.1 (corresponding to residue 698 of SEQ ID NO: 2), and wherein the additional bases correspond to a fragment of exon 1 of the SFTPB gene.
  • an SFTPB promoter fragment of the invention comprises 1 base 3’ to base 5542 of NG_016967.1, the additional base corresponds to base 5543 of NG_016967.1 (corresponding to residue 699 of SEQ ID NO: 2)
  • the additional bases correspond to bases 5543 to 5544 of NG_016967.1 (corresponding to residues 699 to 700 of SEQ ID NO: 2)
  • the additional bases correspond to bases 5543 to 5545 of NG_016967.1 (corresponding to residues 699 to 701 of SEQ ID NO: 2)
  • an SFTPB promoter fragment of the invention comprises 4 bases 3’ to base 5542 of NG_016967.1
  • the additional bases correspond to bases 5543 to 5546 of NG_016967.1 (corresponding to residues 699 to
  • an SFTPB promoter fragment of the invention comprises a portion of SFTPB corresponding to SEQ ID NO: 3.
  • An SFTPB promoter fragment of the invention may not comprise the SFTPB gene start codon.
  • an SFTPB promoter fragment of the invention may comprise of at least part of exon 1 of the SFTPB gene, but does not comprise the SFTPB gene start codon.
  • an SFTPB promoter fragment of the invention typically does not comprise the initiating ATG at bases 5559 to 5561 of NG_016967.1 (corresponding to residues 715 to 717 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention also does not comprise any of the SFTPB gene sequence 3’ of this start codon.
  • An SFTPB promoter fragment of the invention may comprise fewer than 900 bases in length, fewer than 850 bases in length, fewer than 800 bases in length, fewer than 750 bases in length, fewer than 700 bases in length, fewer than 650 bases in length, fewer than 640 bases in length, fewer than 630 bases in length, fewer than 620 bases in length, fewer than 610 bases in length or fewer than 600 bases in length.
  • an SFTPB promoter fragment of the invention comprises fewer than 800 bases in length.
  • a particularly preferred SFTPB promoter fragment of the invention comprises fewer than 700 bases in length.
  • An SFTPB promoter fragment of the invention may consist of fewer than 900 bases in length, fewer than 850 bases in length, fewer than 800 bases in length, fewer than 750 bases in length, fewer than 700 bases in length, fewer than 650 bases in length, fewer than 640 bases in length, fewer than 630 bases in length, fewer than 620 bases in length, fewer than 610 bases in length or fewer than 600 bases in length.
  • an SFTPB promoter fragment of the invention consists of fewer than 800 bases in length.
  • a particularly preferred SFTPB promoter fragment of the invention may consist of fewer than 700 bases in length.
  • the exemplified SFTPB promoter fragment of the invention consists of 630 or 635 bases in length.
  • An SFTPB promoter fragment of the invention may comprise from 600 bases to 800 bases in length, from 600 bases to 750 bases in length, from 600 bases to 700 bases in length, or from 600 bases to 650 bases in length.
  • An SFTPB promoter fragment of the invention may consist of from 600 bases to 800 bases in length, from 600 bases to 750 bases in length, from 600 bases to 700 bases in length, or from 600 bases to 650 bases in length.
  • An SFTPB promoter fragment of the invention may comprise from 630 bases to 800 bases in length, from 635 bases to 800 bases in length, from 630 bases to 750 bases in length, from 635 bases to 750 bases in length, from 630 bases to 700 bases in length, from 635 bases to 700 bases in length, from 630 bases to 650 bases in length, or from 635 bases to 650 bases in length.
  • An SFTPB promoter fragment of the invention may consist of from 630 bases to 800 bases in length, from 635 bases to 800 bases in length, from 630 bases to 750 bases in length, from 635 bases to 750 bases in length, from 630 bases to 700 bases in length, from 635 bases to 700 bases in length, from 630 bases to 650 bases in length, or from 635 bases to 650 bases in length.
  • An SFTPB promoter fragment of the invention may be a fragment of a mammalian or avian, preferably a mammalian SFTPB promoter, i.e. an SFTPB promoter fragment of the invention may be a mammalian or avian SFTPB promoter fragment.
  • a mammalian SFTPB promoter fragment may be a human, non ⁇ human primate, mouse, rat, dog, cat, horse, or cow SFTPB promoter fragment.
  • the SFTPB promoter fragment is a fragment of a human SFTPB promoter, i.e. preferably the SFTPB promoter fragment of the invention is a human SFTPB promoter fragment.
  • an SFTPB promoter fragment of the invention may comprise any functional fragment of SEQ ID NO: 2.
  • such an SFTPB promoter fragment of the invention is of a length as described herein.
  • An SFTPB promoter fragment of the invention may comprise or consist of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8%sequence identity to a fragment of SEQ ID NO: 2, wherein the promoter retains the function of the SFTPB gene promoter, as defined herein.
  • an SFTPB promoter fragment of the invention is of a length as described herein.
  • such an SFTPB promoter fragment of the invention may comprise any additional feature (e.g.
  • An SFTPB promoter fragment of the invention may comprise a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention may comprise a sequence having at least 90% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
  • An SFTPB promoter fragment of the invention may comprise a sequence having at least at least 95% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
  • An SFTPB promoter fragment of the invention may comprise the sequence of bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
  • An SFTPB promoter fragment of the invention may consist of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity or more to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
  • An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 90% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
  • An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 95% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
  • An SFTPB promoter fragment of the invention may consist of the sequence of bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention derived from SEQ ID NO: 2, such as those described above, may comprise a substitution at residue 683 of SEQ ID NO: 2 (corresponding to base 5527 of NG_016967.1).
  • an SFTPB promoter fragment of the invention derived from SEQ ID NO: 2, such as those described above, may comprise substitution of an alanine residue by a cytosine at residue 683 of SEQ ID NO: 2 (corresponding to base 5527 of NG_016967.1), in other words, may comprise an A683C substitution.
  • An exemplified SFTPB promoter fragment of the invention is SEQ ID NO: 1.
  • an SFTPB promoter fragment which comprises a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to SEQ ID NO: 1.
  • an SFTPB promoter fragment of the invention may comprise a sequence having at least 90% identity to SEQ ID NO: 1.
  • An SFTPB promoter fragment of the invention may comprise a sequence having at least at least 95% identity to SEQ ID NO: 1.
  • an SFTPB promoter fragment of the invention may comprise the sequence of SEQ ID NO: 1.
  • the present invention provides an SFTPB promoter fragment which consists of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to SEQ ID NO: 1.
  • an SFTPB promoter fragment of the invention may consist of a sequence having at least 90% identity to SEQ ID NO: 1.
  • An SFTPB promoter fragment of the invention may consist of a sequence having at least at least 95% identity to SEQ ID NO: 1.
  • an SFTPB promoter fragment of the invention may consist of the sequence of SEQ ID NO: 1.
  • An SFTPB promoter fragment of the invention may further comprise up to 10 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bases), up to 15 bases, up to 20 bases, up to 25 bases or up to 25 bases at the 5’ end.
  • additional bases may preferably be present when said SFTPB promoter fragment (i) comprises a sequence having at least 70% identity to SEQ ID NO: 1; and/or (ii) comprises a sequence having at least 70% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2); as described herein.
  • additional bases are present at the 5’ end of an SFTPB promoter fragment of the invention
  • said additional bases may be bases which correspond to the corresponding number of bases 5’ to base 4925 of NG_016967.1 (corresponding to residue 81 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention comprises 10 bases 5’ to base 4925 of NG_016967.1
  • the additional base corresponds to bases 4915 to 4924 of NG_016967.1 (corresponding to residues 71 to 80 of SEQ ID NO: 2)
  • the additional bases correspond to bases 4905 to 4924 of NG_016967.1 (corresponding to residues 61 to 80 of SEQ ID NO: 2), and so on.
  • an SFTPB promoter fragment of the invention may further comprise up to 20 bases at the 5’ end, and particularly preferably, the up to 20 additional bases correspond to bases 5’ of residue 4925 of NG_016967.1 (corresponding to residue 81 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention may further comprise up to 10 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bases) bases at the 3’ end.
  • Such additional bases may preferably be present when said SFTPB promoter fragment (i) comprises a sequence having at least 70% identity to SEQ ID NO: 1; and/or (ii) comprises a sequence having at least 70% identity to bases 4925 ⁇ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2); as described herein.
  • said additional bases may be bases which correspond to the corresponding number of bases 3’ to base 5554 of NG_016967.1 (corresponding to residue 710 of SEQ ID NO: 2).
  • an SFTPB promoter fragment of the invention comprises 10 bases 3’ to base 5554 of NG_016967.1
  • the additional base corresponds to bases 5555 to 5564 of NG_016967.1 (corresponding to residues 711 to 720 of SEQ ID NO: 2)
  • the additional bases correspond to bases 5555 to 5559 of NG_016967.1 (corresponding to residues 711 to 715 of SEQ ID NO: 2)
  • the additional bases correspond to bases 5555 to 5558 of NG_016967.1 (corresponding to residues 711 to 714 of SEQ ID NO: 2), and so on.
  • an SFTPB promoter fragment of the invention may further comprise up to 4 bases at the 3’ end, and particularly preferably, the up to 4 additional bases correspond to bases 3’ of residue 5554 of NG_016967.1 (corresponding to residue 710 of SEQ ID NO: 2).
  • An SFTPB promoter fragment of the invention may comprise additional sequences at the 5’ and/or 3’ end to facilitate molecular biology applications of said SFTPB promoter fragment.
  • one or more restriction enzyme site may be added at the 5’ and/or 3’ end of the sequence. Such restriction enzyme sites may be used to facilitate cloning of the SFTPB promoter fragment into a non ⁇ viral vector (e.g.
  • restriction enzyme sites are present at both the 5’ and 3’ end of the SFTPB promoter fragment, each restriction enzyme site may be selected independently, and thus the 5’ and 3’ sites may be the same or different.
  • restriction enzyme sites include NheI and BgIII.
  • a SFTPB promoter fragment of the invention may have a 5’ BgIII restriction enzyme site (5’AGATCT3’) and/or a 3’ NheI restriction enzyme site (5’GCTAGC3’).
  • Enhancer As exemplified herein, the inventors have also modified the SFTPB promoter fragments of the invention to further increase transgene expression levels and/or to increase promoter activity, particularly in the lung parenchyma.
  • the inventors modified the SFTPB promoter fragment to include an enhancer sequence.
  • the invention further provides an SFTPB promoter fragment which further comprises an enhancer. All disclosure herein to SFTPB promoter fragments of the invention applies equally and without reservation to SFTPB promoter fragment which further comprise an enhancer.
  • An enhancer is a cis ⁇ acting DNA sequence which can increase gene transcription.
  • An enhancer of the invention may be from about 150 to about 900 bp in length, such as from about 200 to about 900 bp, from about 200 to about 800 bp, from about 300 to about 700 bp, from about 400 to about 800 bp, or from about 300 to about 600 bp in length.
  • An enhancer of the invention may be linked to an SFTPB promoter fragments of the invention by a linker.
  • Said linker is typically a short DNA sequence, which may be from about 1 to about 50 bp in length, such as from about 1 to about 20 bp, from about 1 to about 10 bp, from about 5 to about 20 bp in length, or from about 5 to about 10 bp in length.
  • a linker may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bp, particularly 6 bp, in length.
  • a non ⁇ limiting example of a linker sequence is given in SEQ ID NO: 51.
  • Said linker may comprise or consist of one or more restriction enzyme site, non ⁇ limiting examples of which are described herein.
  • a promoter/enhancer combination of the invention may be joined by a linker which comprises or consists of a NheI restriction site and/or a BgIII restriction site, preferably a BgIII restriction site.
  • An enhancer of the invention may comprise additional sequences at the 5’ and/or 3’ end.
  • one or more restriction enzyme site may be added at the 5’ and/or 3’ end of the enhancer.
  • each restriction enzyme site may be selected independently, and thus the 5’ and 3’ sites may be the same or different.
  • Said 5’ or 3’ restriction site may comprise all or part of a linker joining the enhancer to the SFTPB promoter fragment of the invention.
  • the enhancer may be (i) 5’ to an SFTPB promoter fragment of the invention, or (ii) 3’ to an SFTPB promoter fragment of the invention.
  • the enhancer is 5’ to an SFTPB promoter fragment of the invention.
  • the (5’ or 3’, preferably 5’) enhancer may be in (i) the forward, or (ii) the reverse orientation.
  • the enhancer is in the forward orientation.
  • an SFTPB promoter fragment further comprising an enhancer sequence as defined herein has the potential to provide an even greater increase in transgene expression compared with the full ⁇ length SFTPB promoter (i.e. transgene expression increases full ⁇ length SFTPB promoter ⁇ SFTPB promoter fragment of the invention ⁇ SFTPB promoter of the invention further comprising an enhancer).
  • Higher expression may be defined as greater than 105%, greater than 110%, greater than 120%, greater than 130%, greater than 140%, greater than 150%, greater than 175%, greater than 200%, greater than 225%, greater than 250%, greater than 300%, greater than 350%, greater than 400% or more of the transgene expression compared with a suitable control, such as the SFTPB promoter fragment alone, or the full ⁇ length SFTPB promoter.
  • a suitable control such as the SFTPB promoter fragment alone, or the full ⁇ length SFTPB promoter.
  • expression of a transgene by an SFTPB promoter of the invention further comprising an enhancer may be greater than 105%, greater than 110%, greater than 120%, greater than 130%, greater than 140%, greater than 150%, greater than 175%, greater than 200%, greater than 225%, or greater than 250% of the transgene expression when using the SFTPB promoter fragment alone.
  • Transgene expression may be quantified at the nucleic acid and/or protein level, and can be quantified by any suitable standard technique known to the person skilled in the art, for example, by real ⁇ time reverse transcription polymerase chain reaction (RT ⁇ qPCR), Western blotting and enzyme ⁇ linked immunosorbent assay or ELISA.
  • the inventors also surprisingly found that the (cell ⁇ specific) expression driven by a SFTPB promoter fragment comprising an enhancer sequence as defined herein is comparable to gene expression levels when using ubiquitously used strong promoters (e.g. CMV, hCEF).
  • Comparable expression may be defined as at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% of gene expression relative to expression of the same transgene using the same strong promoter (e.g. CMV or hCEF promoter [such as SEQ ID NOs: 6 and 5 respectively]).
  • Expression of the transgene using the SFTPB promoter fragment may be higher than expression using a strong promoter (e.g. CMV or hCEF) promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more relative to expression of the same transgene using the same strong promoter (e.g. CMV or hCEF promoter).
  • a strong promoter e.g. CMV or hCEF promoter
  • the SFTPB promoter fragment comprises a SFTPB promoter fragment as defined above operably linked to an enhancer
  • said enhancer may preferably be of viral origin, a lung ⁇ preferred enhancer, a lung ⁇ parenchyma ⁇ preferred enhancer, a pneumocyte ⁇ preferred enhancer, an ATII and club cell ⁇ preferred enhancer, or an ATII cell preferred enhancer, a lung ⁇ specific enhancer, a lung ⁇ parenchyma ⁇ specific enhancer, a pneumocyte ⁇ specific enhancer, an ATII and club cell ⁇ specific enhancer, or an ATII cell specific enhancer.
  • the terms “preferred” and “specific” are defined herein.
  • the term “preferred” may alternatively or additionally be defined as higher expression in the lung relative to expression in the nose (particularly the nasal cavity), and thus give rise to an increased ratio of lung expression: nose expression, as described herein.
  • Expression in the lung and nasal cavity can be determined using an in vivo luciferase reporter assay (e.g., wherein vectors comprising nucleic acid cassettes of the invention are administered to mice, and the relative bioluminescence of the lungs and nasal cavity is quantified).
  • the enhancer may be selected from a hB ⁇ actin enhancer, a SLC34A2 (Sodium ⁇ dependent phosphate transport protein 2B) enhancer, a VEGFA (Vascular Endothelial Growth Factor A) enhancer, a CMV (Cytomegalovirus) enhancer, an SV40 (simian virus 40) enhancer, an ELF3 (E74 like ETS transcription factor 3) enhancer, an SFTPC (surfactant protein C) enhancer, an SFTPB enhancer, or a LMO7 (LIM domain 7) enhancer.
  • a hB ⁇ actin enhancer a SLC34A2 (Sodium ⁇ dependent phosphate transport protein 2B) enhancer, a VEGFA (Vascular Endothelial Growth Factor A) enhancer, a CMV (Cytomegalovirus) enhancer, an SV40 (simian virus 40) enhancer, an ELF3 (E74 like ETS transcription factor 3) enhancer, an SFTPC
  • the enhancer is selected from a SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer.
  • the enhancer is a SLC34A2 enhancer.
  • the enhancer is a VEGFA enhancer.
  • the enhancer is a CMV enhancer.
  • the enhancer is a SV40 enhancer.
  • the enhancer is an ELF3 enhancer.
  • Combinations of enhancers, typically those identified herein, and combinations of one or more preferred enhancer described herein, may be used according to the present invention.
  • the CMV enhancer may be in the forwards orientation.
  • the CMV enhancer is in the reverse orientation.
  • the ELF3 enhancer may be in the forwards orientation.
  • the SV40 enhancer may be in the forwards orientation.
  • the SV40 enhancer may be in the reverse orientation.
  • the VEGFA enhancer may be in the forwards orientation.
  • the VEGFA enhancer may be in the reverse orientation.
  • the SLC34A2 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
  • the SLC34A2 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SQE ID NO: 31.
  • the SLC34A2 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
  • the SLC34A2 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
  • the SLC34A2 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
  • the SLC34A2 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.
  • the VEGFA enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
  • the VEGFA enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
  • the VEGFA enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
  • the VEGFA enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
  • the VEGFA enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
  • the VEGFA enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 36 ⁇ 40, preferably SEQ ID NO: 38.
  • the CMV enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11.
  • the CMV enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11.
  • the CMV enhancer may comprise the nucleotide sequence of SEQ ID NO: 5.
  • the CMV enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11.
  • the CMV enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11.
  • the CMV enhancer may consist of the nucleotide sequence of SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11. This exemplary CMV enhancer sequence is CpG ⁇ free.
  • the SV40 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 34 or 35.
  • the SV40 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 34 or 35.
  • the SV40 enhancer may comprise the nucleotide sequence of SEQ ID NO: 34 or 35.
  • the SV40 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 34 or 35.
  • the SV40 enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or 35.
  • the SV40 enhancer may consist of the nucleotide sequence of SEQ ID NO: 34 or 35.
  • the ELF3 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 13 to 20.
  • the ELF3 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 13 to 20.
  • the ELF3 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 13 to 20.
  • the ELF3 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 13 to 20.
  • the ELF3 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 13 to 20.
  • the ELF3 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 13 to 20.
  • the actin enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 21 or 22.
  • the actin enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 21 or 22.
  • the actin enhancer may comprise the nucleotide sequence of SEQ ID NO: 21 or 22.
  • the actin enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 21 or 22.
  • the actin enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 21 or 22.
  • the actin enhancer may consist of the nucleotide sequence of SEQ ID NO: 21 or 22.
  • the LMO7 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 23 to 25.
  • the LMO7 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to any one of SEQ ID NOs: 23 to 25.
  • the LMO7 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 23 to 25.
  • the LMO7 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID NOs: 23 to 25.
  • the LMO7 enhancer may consist of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 23 to 25.
  • the LMO7 enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 23 to 25.
  • the SFTPC enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 27 or 28.
  • the SFTPC enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 27 or 28.
  • the SFTPC enhancer may comprise the nucleotide sequence of SEQ ID NO: 27 or 28.
  • the SFTPC enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 27 or 28.
  • the SFTPC enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 27 or 28.
  • the SFTPC enhancer may consist of the nucleotide sequence of SEQ ID NO: 27 or 28.
  • the SFTPB enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 26.
  • the SFTPB enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID NO: 26.
  • the SFTPB enhancer may comprise the nucleotide sequence of SEQ ID NO: 26.
  • the SFTPB enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 26.
  • the SFTPB enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 26.
  • the SFTPB enhancer may consist of the nucleotide sequence of SEQ ID NO: 26. Any SFTPB promoter fragment of the invention may be combined with any enhancer of the invention.
  • any preferred SFTPB promoter fragment of the invention may be combined with any preferred enhancer (e.g. an SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer).
  • any preferred enhancer e.g. an SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer.
  • a preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).
  • a particularly preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).
  • a particularly preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 47.
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 47.
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 47.
  • Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 47.
  • a preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).
  • a particularly preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83).
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83).
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83).
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 83).
  • a particularly preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 48.
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 48.
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of SEQ ID NO: 48.
  • Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of SEQ ID NO: 48.
  • a preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).
  • a particularly preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81).
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81).
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81).
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 81).
  • a particularly preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 46.
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 46.
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 46.
  • Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO: 46.
  • a preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).
  • Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).
  • Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).
  • Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).
  • a preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).
  • Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).
  • Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise a nucleic acid sequence of SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).
  • Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).
  • a preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%,a t least 98%, at least 99% identity or more to any one of SEQ ID NOs: 41 to 48.
  • a preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to any one of SEQ ID NOs: 41 to 48.
  • a preferred SFTPB promoter fragment further comprising an enhancer may comprise a nucleic acid sequence of any one of SEQ ID NOs: 41 to 48.
  • a preferred SFTPB promoter fragment further comprising an enhancer may consist of a nucleic acid sequence of any one of SEQ ID NOs: 41 to 48.
  • a particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%,a t least 98%, at least 99% identity or more to any one of SEQ ID NOs: 46 to 48.
  • a particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to any one of SEQ ID NOs: 46 to 48.
  • a particularly preferred SFTPB promoter fragment further comprising an enhancer may comprise a nucleic acid sequence of any one of SEQ ID NOs: 46 to 48.
  • a particularly preferred SFTPB promoter fragment further comprising an enhancer may consist of a nucleic acid sequence of any one of SEQ ID NOs: 46 to 48.
  • Lung parenchyma and specific/preferred expression therein The respiratory system can be divided into airways and lung parenchyma.
  • the airways consist of the bronchus, which bifurcates off the trachea and divides into bronchioles and then further into alveoli.
  • the parenchyma is responsible for gas exchange and includes the alveoli, alveolar ducts, and terminal and respiratory bronchioles.
  • the most prominent structure in the lung parenchyma is the alveolus. Two types of epithelial cell line the alveolus.
  • ATI cells exhibit a broad, flattened morphology and cover around 95% of the surface area, whilst the cuboidal alveolar type II cells (ATII cells) line the remainder of the alveolus.
  • ATI cells provide a gas exchange interface with the underlying endothelium, whereas ATII cells serve as both progenitors of ATI cells and also play a critical role in maintaining the homeostasis of the alveolus. The latter role is fulfilled by the secretion of surfactant proteins from specialised organelles within ATII cells, so ⁇ called ‘lamellar bodies’, into the alveolar space.
  • ATII cells are the only epithelial cell of the lung which synthesise and release all four surfactant proteins A, B, C and D, with surfactant protein C being unique to the ATII cell.
  • ATII cells have the following functions: (1) the transepithelial movement of water and ions regulating the volume of the alveolar surface liquid (ASL) preventing alveoli flooding, (2) the expression of immunomodulatory proteins necessary for host defence and the regulation of innate immunity and (3) the regeneration of alveolar epithelium after injury.
  • ASL alveolar surface liquid
  • surfactant proteins A, B and D are also synthesised by club cells (previously named Clara Cells) founds in the terminal and respiratory bronchioles of humans.
  • Club cells are non ⁇ ciliated epithelial cells found mainly in bronchioles as well as basal cells found in large airways. They have been ascribed several protective roles, including airway repair after injury, secretion of anti ⁇ inflammatory and immunomodulatory proteins, and detoxification.
  • ATI dysfunction, ATII dysfunction and/or club cell dysfunction or dropout is associated with the pathogenesis of various parenchymal lung diseases. Accordingly, the lung parenchyma may be targeted for treating genetic diseases such as surfactant deficiencies and interstitial lung disease.
  • the promoters of the invention drive transgene expression in the lung parenchyma, typically one or more cell type of the lung parenchyma.
  • the promoters of the invention drive transgene expression in (i) ATI cells, (ii) ATII cells, (iii) club cells, (iv) ATI cells and ATII cells, (v) ATI cells and club cells, (vi) ATII cells and club cells, or (vii) ATI cells, ATII cells and club cells.
  • a SFTPB promoter fragment of the invention is functional. As used herein, the term “functional” may mean that an SFTPB promoter fragment of the invention retains the functionality of the full ⁇ length SFTPB promoter.
  • a promoter of the invention may be defined as being capable of expressing a gene of interest (e.g., a transgene) in a cell type which, in a healthy subject, would express SFTPB.
  • a gene of interest e.g., a transgene
  • the an SFTPB promoter fragment may be used to express a transgene in ATII cells, club cells and/or ATI cells.
  • an SFTPB promoter fragment of the invention may be defined as functional as it is capable of preferentially or specifically expressing a gene of interest (e.g., a transgene) in a tissue ⁇ type or cell ⁇ type which, in a healthy subject, would express SP ⁇ B.
  • tan SFTPB promoter fragment of the invention may be a lung ⁇ parenchyma preferred promoter.
  • An SFTPB promoter fragment of the invention may be a lung ⁇ parenchyma specific promoter.
  • An SFTPB promoter fragment of the invention may be a pneumocyte ⁇ preferred promoter, whereby a pneumocyte is defined as any of the specialized cells of the alveoli of the lungs.
  • An SFTPB promoter fragment of the invention may be a pneumocyte ⁇ specific promoter.
  • An SFTPB promoter fragment of the invention may be an ATII cell ⁇ preferred promoter, a club cell ⁇ preferred promoter and/or an ATI cell ⁇ preferred promoter.
  • An SFTPB promoter fragment of the invention may be an ATII cell ⁇ specific promoter, a club cell ⁇ specific expression and/or an ATI cell ⁇ specific promoter.
  • An SFTPB promoter fragment of the invention may preferably be an ATII cell ⁇ preferred promoter.
  • An SFTPB promoter fragment of the invention may preferably be an ATII cell ⁇ specific promoter.
  • Tissue or cell preferred expression may be defined as expression that is higher in said tissue or cell than other tissue or cell types.
  • lung ⁇ parenchyma preferred expression may be defined as expression that is significantly higher in the lung ⁇ parenchyma (or one or more cell type therein, as described above) than expression in one or more of: the brain, the eye, the endocrine tissues, the proximal digestive tract, the gastrointestinal tract, liver and gall bladder, pancreas, the kidney and/or urinary bladder, male tissues (i.e., the testis, epididymis, prostate and/or seminal vesicle) and female tissues (i.e., the vagina, breast, cervix, endometrium, fallopian tube, ovary and placenta), muscle tissues, connective and soft tissues, skin, bone marrow and lymphoid tissue.
  • male tissues i.e., the testis, epididymis, prostate and/or seminal vesicle
  • female tissues i.e., the vagina, breast, cervix, endometrium, fallopian tube, ovary and place
  • Lung ⁇ parenchyma preferred expression may be defined as expression that is at least about 5 times greater, at least about 10 times greater, at least about 20 times greater, at least about 50 times greater or at least about 100 or more times greater in the lung parenchyma than one or more of the above reference tissue types, especially the gastrointestinal tract and/or the brain.
  • Tissue or cell specific expression may be defined as expression that is at least about 10 times greater, at least about 20 times greater, at least about 50 times greater or at least about 100 or more times greater in said tissue or cell than any other tissue or cell types.
  • Preferential and/or specific expression may be assessed at the level of RNA and/or protein expression, preferably protein expression. Expression can be measured by any suitable standard technique known to the person skilled in the art.
  • RNA expression levels can be measured by quantitative real ⁇ time PCR.
  • Protein expression can be measured by western blotting or immunohistochemistry.
  • restricting the expression of transgenes to cells expressing endogenous SFTPB is expected to reduce the effects of off ⁇ target gene expression, overexpression (e.g., toxicity/ER stress/UPR, etc) and/or reduce immune responses.
  • the SFTPB promoters of the invention as a consequence of their preferential and/or specific expression in the lung parenchyma, or one or more cell type thereof, have potential clinical benefits as a result of these advantageous properties.
  • an SFTPB promoter fragment of the invention may reduce the effects of off ⁇ target gene expression, overexpression, and/or reduce immune responses. Additionally, or alternatively, compared to a hCEF promoter of SEQ ID NO: 5and/or a CMV promoter of SEQ ID NO: 6, an SFTPB promoter fragment of the invention may preferentially or specifically express a transgene in the lung parenchyma, pneumocytes, such as ATII cells, ATI cells, and/or club cells.
  • an SFTPB promoter fragment of the invention may increase transgene expression compared with a full ⁇ length SFTPB promoter, such as that described herein.
  • an SFTPB promoter fragment of the invention increases transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)).
  • An SFTPB promoter fragment of the invention increases transgene expression in lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2 ⁇ fold, at least about 2.5 ⁇ fold, at least about 3 ⁇ fold, at least about 4 ⁇ fold, at least about 5 ⁇ fold, at least about 7.5 ⁇ fold or more.
  • the increase in transgene expression by an SFTPB promoter fragment of the invention may be quantified compared with a suitable control, preferably compared with a full ⁇ length SFTPB promoter, such as that described herein.
  • an SFTPB promoter fragment of the invention increases transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with expression of the same transgene in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with a full ⁇ length SFTPB promoter, such as that described herein.
  • An SFTPB promoter fragment of the invention increases transgene expression in lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2 ⁇ fold, at least about 2.5 ⁇ fold, at least about 3 ⁇ fold, at least about 4 ⁇ fold, at least about 5 ⁇ fold, at least about 7.5 ⁇ fold or more compared with expression of the same transgene in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with a full ⁇ length SFTPB promoter, such as that described herein.
  • the SFTPB promoter fragment of the invention increases transgene expression in ATII cells, and optionally one or more additional lung parenchymal cell type as described herein. Again, this increase in expression is preferably compared with expression of the same transgene in the same cell type(s) by a full ⁇ length SFTPB promoter, such as that described herein.
  • an SFTPB promoter fragment of the invention may preferentially drive transgene expression in the lung (particularly the lung parenchyma or one or more cell type thereof) compared with transgene expression in the nose (or cells thereof).
  • an SFTPB promoter fragment of the invention may preferentially increase transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with transgene expression in the nose (or cells thereof).
  • An SFTPB promoter fragment of the invention may preferentially increase transgene expression in the lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2 ⁇ fold, at least about 2.5 ⁇ fold, at least about 3 ⁇ fold, at least about 4 ⁇ fold, at least about 5 ⁇ fold, at least about 7.5 ⁇ fold or more compared with transgene expression in the nose (or cells thereof).
  • the ratio of lung expression: nose expression by an SFTPB promoter fragment of the invention may be at least about 2:1, such as at least about 2.5:1, at least about 3:1, at least about 4:1, at least about 5:1 or more.
  • any disclosure herein in relation to promoters (i.e. SFTPB promoter fragments) of the invention applies equally and without reservation to other aspects of the invention comprising, or relating to, said promoters.
  • any disclosure herein in relation to promoters (i.e. SFTPB promoter fragments) of the invention applies equally and without reservation to promoter/enhancer combinations, viral and non ⁇ viral vectors of the invention.
  • promoter/enhancer combinations, viral and non ⁇ viral vectors of the invention drive transgene expression in the lung parenchyma, typically one or more cell type of the lung parenchyma.
  • the promoters i.e. SFTPB promoter fragments
  • promoter/enhancer combinations, viral and non ⁇ viral vectors of the invention drive transgene expression in (i) ATI cells, (ii) ATII cells, (iii) club cells, (iv) ATI cells and ATII cells, (v) ATI cells and club cells, (vi) ATII cells and club cells, or (vii) ATI cells, ATII cells and club cells.
  • Nucleic Acid Cassettes The invention also provides a nucleic acid cassette.
  • the present invention provides a nucleic acid cassette comprising (a) an SFTPB promoter fragment; and (b) a transgene.
  • a transgene may be defined as anucleic acid sequence encoding a therapeutic protein.Thus, the terms “nucleic acid sequence encoding a therapeutic protein” and the term “transgene” may be used interchangeably.
  • an SFTPB promoter fragment may increase expression of the transgene by the lung parenchyma (e.g. ATII cells), as defined herein.
  • the increase in expression of a transgene by an SFTPB promoter fragment of the invention may be as defined herein, including disclosure of increasing transgene expression using SFTPB promoter fragments and/or SFTPB promoter fragments combined with an enhancer, as described above.
  • any disclosure herein in relation to increasing transgene expression using an SFTPB promoter fragment and/or SFTPB promoter fragment combined with an enhancer of the invention applies equally and without reservation to nucleic acid cassettes of the invention.
  • the increase in expression of a transgene by an SFTPB promoter fragment of the invention is an increase of at least about 50%, at least about 60%, at least about 70%, at least about 80% or more.
  • the increase in expression of a transgene by an SFTPB promoter fragment of the invention is an increase of at least about 50%.
  • an SFTPB promoter fragment of the invention may increase expression of the transgene by the lung parenchyma (e.g.
  • the SFTPB promoter fragment of the invention may increase expression of the transgene by the lung parenchyma (e.g. ATII cells) relative to expression using the full ⁇ length SFTPB gene promoter (e.g., the 972 bp genomic fragment defined above).
  • expression of the transgene by the lung parenchyma (e.g. ATII cells) using the SFTPB promoter fragment in a nucleic acid of the invention may be comparable to the expression of the transgene using a ubiquitously used strong promoter (e.g. CMV or hCEF).
  • the expression of the transgene using the SFTPB promoter fragment may be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% of the expression using a CMV promoter, as described herein.
  • expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may be higher than expression using a CMV promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more relative to expression of the same transgene using a CMV promoter.
  • expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% of the expression using a hCEF promoter.
  • expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may be higher than expression using a hCEF promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more relative to expression of the same transgene using a CMV promoter.
  • a nucleic acid cassette or vector of the invention enables long ⁇ term transgene expression, resulting in long ⁇ term expression of a transgene, which offers clinical benefits for the expression of therapeutic proteins in patients.
  • Long ⁇ term expression means expression of a transgene, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more.
  • long ⁇ term expression means expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
  • the long ⁇ term expression is typically accompanied by long ⁇ term secretion or long ⁇ term membrane insertion of the (therapeutic) protein encoded by the transgene, depending on whether the (therapeutic) protein is a secreted protein (e.g. SFTPB) or a membrane protein.
  • Long ⁇ term secretion means secretion of a (therapeutic) protein encoded by a transgene of the invention, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more.
  • long ⁇ term secretion means secretion for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
  • Long ⁇ term membrane insertion means that a (therapeutic) protein encoded by a transgene of the invention is inserted into and present in the cell membrane, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more.
  • long ⁇ term expression means membrane insertion for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
  • a nucleic acid cassette or vector of the invention may drive (increased) long ⁇ lasting expression of a transgene, and/or (preferably and) expression and secretion/membrane insertion of a (therapeutic) protein encoded by said transgene in one or more cell type of the lung parenchyma, as described herein, in vivo in a patient.
  • a nucleic acid cassette or vector of the invention drives expression of a transgene, and/or (preferably and) expression and secretion/membrane insertion of a (therapeutic) protein encoded by a transgene of the invention in one or more cell / cell type of the lung parenchyma, as described herein, for at least 45 days, more preferably at least 90 days.
  • the nucleic acid of the nucleic acid cassette may be as defined herein.
  • the nucleic acid cassette comprise DNA or RNA.
  • the nucleic acid cassette is DNA.
  • a nucleic acid cassette of the invention may optionally be codon optimised for expression in a particular cell type, for example, eukaryotic cells (e.g.
  • codon optimised refers to the replacement of at least one codon within a base polynucleotide sequence with a codon that is preferentially used by the host organism in which the polynucleotide is to be expressed. Typically, the most frequently used codons in the host organism are used in the codon ⁇ optimised polynucleotide sequence. Methods of codon optimisation are well known in the art. It will be understood by a skilled person that numerous different polynucleotides can encode the same polypeptide as a result of the degeneracy of the genetic code.
  • a nucleic acid cassette that comprises or consists of the SFTPB promoter fragment and transgene (e.g., encoding a therapeutic protein) of the invention includes all polynucleotide sequences that are degenerate versions of each other and that encode the same amino acid sequence.
  • a nucleic acid cassette of the invention preferably comprises an SFTPB promoter fragment comprising an SFTPB promoter fragment operably linked to an enhancer sequence, as described herein.
  • a nucleic acid cassette of the invention typically comprises a SFTPB promoter fragment comprising a SFTPB promoter fragment and an enhancer sequence, wherein the SFTPB promoter fragment is operably linked to a nucleic acid sequence comprising or consisting of a transgene (e.g., encoding a therapeutic protein).
  • a transgene e.g., encoding a therapeutic protein.
  • operably linked it is meant that the SFTPB promoter fragment is configured to express the transgene (e.g., encoding the therapeutic protein).
  • the transgene encoding the therapeutic protein may also be linked to a suitable terminator sequence. Suitable terminator sequences are well known in the art.
  • the promoter included in the nucleic acid cassettes and vectors of the invention may be specifically selected and/or modified to further refine regulation of expression of the therapeutic gene.
  • an SFTPB promoter fragment of the invention may be modified to reduce the number of CpG dinucleotides, or to render the SFTPB promoter fragment CpG ⁇ free.
  • a number of (CpG ⁇ free) promoters, and methods for the generation of CpG ⁇ free promoters which are suitable for use in the present invention are described in Pringle et al. (J. Mol. Med. Berl. 2012, 90(12): 1487 ⁇ 96), which is herein incorporated by reference in its entirety.
  • the nucleic acid cassettes and vectors of the invention comprise an SFTPB promoter fragment having low or no CpG dinucleotide content.
  • Low CpG dinucleotide content may be defined as 10 CpG dinucleotides or less, preferably 5 CpG dinucleotides or less, such as 5, 4, 3, 2 or 1 CpG dinucleotides.
  • An SFTPB promoter fragment may have some or all CG dinucleotides replaced with any one of AG, TG or GT.
  • the absence (or reduction) of CpG dinucleotides further improves the performance of some nucleic acid cassettes and vectors of the invention, particularly lentiviral (e.g. SIV) vectors of the invention and in particular in situations where it is not desired to induce an immune response against an expressed antigen or an inflammatory response against the delivered expression construct.
  • the elimination or reduction of CpG dinucleotides reduces the occurrence of flu ⁇ like symptoms and inflammation which may result from administration of constructs, particularly when administered to the airways.
  • the nucleic acid cassettes and vectors of the invention may be modified to allow shut down of gene expression. Standard techniques for modifying the vector in this way are known in the art. As a non ⁇ limiting example, Tet ⁇ responsive promoters are widely used.
  • the nucleic acid cassette of the invention (or a vector comprising said cassette) may have an intron positioned between the promoter and the transgene.
  • suitable introns are found for example, in UK Application No. 2213936.4, which is herein incorporated by reference in its entirety.
  • nucleic acid of the invention is present in a non ⁇ viral vector (e.g. plasmid), the presence of at least one intron between the SFTPB promoter fragment and the transgene may be preferred, for example an intron as described in UK Application No. 2213936.4.
  • the nucleic acid cassettes and vectors of the invention may include at least one part of a vector, in particular, regulatory elements.
  • the promoter within a nucleic acid cassette of the invention may be used to express more than one polypeptide, including one or more therapeutic protein.
  • the nucleic acid cassette may comprise a nucleic acid sequence which, when transcribed, gives rise to multiple polypeptides, for instance a transcript may contain multiple open reading frames (ORFs) and also one or more Internal Ribosome Entry Sites (IRES) to allow translation of ORFs after the first ORF.
  • a transcript may be polycistronic, i.e. it may be translated to give a polypeptide which is subsequently cleaved to give a plurality of polypeptides.
  • a nucleic acid cassette of the invention may comprise multiple promoters, including multiple SFTPB promoter fragments of the invention, or multiple copies of any specific an SFTPB promoter fragment of the invention, and hence give rise to a plurality of transcripts and hence a plurality of polypeptides, including a plurality of therapeutic proteins.
  • Nucleic acid cassettes may, for instance, express one, two, three, four or more polypeptides via a promoter or promoters, including one or more SFTPB promoter fragment of the invention.
  • a nucleic acid cassette may comprise one or more translation initiation sequence (TIS).
  • Translation initiation plays an important role in mRNA translation, canonically a methionyl tRNA unique for initiation (Met ⁇ tRNAi) identifies the AUG start codon and triggers the downstream translation process.
  • Non ⁇ canonical start codons e.g. CUG for valyl ⁇ tRNA
  • the nucleic acid cassettes of the present invention may comprise at least one termination signal.
  • a “termination signal” or “terminator” is comprised of the DNA sequences involved in specific termination of an RNA transcript by an RNA polymerase. Thus, a termination signal that ends the production of an RNA transcript is contemplated according to the present invention.
  • a terminator may be necessary in vivo to achieve desirable message levels.
  • a terminator region may also comprise specific DNA sequences that permit site ⁇ specific cleavage of the new transcript so as to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch of about 200 A residues (polyA) to the 3’ end of the transcript. RNA molecules modified with this polyA tail appear to more stable and are translated more efficiently.
  • a terminator typically comprises a signal for the cleavage of the RNA, and it is preferred that the terminator signal promotes polyadenylation of the message.
  • the terminator and/or polyadenylation site elements can serve to enhance message levels and to minimize read through from the cassette into other sequences.
  • Terminators contemplated for use in the invention include any known terminator of transcription described herein or known to one of ordinary skill in the art, including but not limited to, for example, the termination sequences of genes, such as for example the bovine growth hormone terminator or viral termination sequences, such as for example the SV40 terminator.
  • the termination signal may be a lack of transcribable or translatable sequence, such as due to a sequence truncation.
  • the invention also provides gene therapy vectors comprising a nucleic acid cassette of the invention.
  • nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in a given host cell.
  • Any appropriate host cell may be used, such as mammalian, bacterial, insect, yeast, and/or plant host cells.
  • cell ⁇ free expression systems may be used. Such expression systems and host cells are standard in the art.
  • nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in the lung.
  • the nucleic acid cassettes and vectors of the invention may be capable of expressing the transgene in the lung parenchyma.
  • the nucleic acid cassettes and vectors of the invention may be capable of expressing the transgene in one or more cell type selected from ATII cells, ATI cells, club cells, and/or bronchioalveolar stem cells.
  • the nucleic acid cassettes and vectors of the invention are typically capable of expressing the transgene in one or more ATII cells, ATI cells, club cells, bronchioalveolar stem cells in the terminal bronchioles.
  • the nucleic acid cassettes and vectors of the invention are capable of expressing the transgene in ATII cells.
  • the nucleic acid cassettes and vectors of the invention are capable of expressing the (therapeutic) protein encoded by the transgene in one or more cell/cell type of the lung parenchyma, particularly ATII cells, ATI cells, club cells and/or bronchioalveolar stem cells in the terminal bronchioles, particularly in ATII cells.
  • the nucleic acid cassettes of the invention may be made using any suitable process known in the art.
  • the nucleic acid cassettes may be made using chemical synthesis techniques.
  • the nucleic acid cassettes of the invention may be made using molecular biology techniques.
  • Non ⁇ Viral Vectors The present invention also provides a vector comprising a nucleic acid cassette of the invention.
  • the vector(s) may be present in the form of a therapeutic composition or formulation.
  • the vector may be a non ⁇ viral vector.
  • the non ⁇ viral vector(s) may be a DNA vector, such as a DNA plasmid.
  • the vector(s) may be an RNA vector, such as a mRNA vector or a self ⁇ amplifying RNA vector.
  • the non ⁇ viral vector may be an exosomes or microvesicle (MV).
  • the non ⁇ viral (e.g. DNA and/or RNA) vector(s) of the invention may be capable of expression in eukaryotic and/or prokaryotic cells.
  • the non ⁇ viral e.g.
  • DNA and/or RNA vector(s) are capable of expression in a cell of a subject, for example, a cell of a mammalian or avian subject to be immunised.
  • the nucleic acid cassettes and vectors of the invention are capable of expressing a transgene in airway cells, preferably lung parenchymal cells (as described herein).
  • a non ⁇ viral vector of the present invention may be a phage vector, such as an AAV/phage hybrid vector as described in Hajitou et al., Cell 2006; 125(2) pp. 385 ⁇ 398; herein incorporated by reference.
  • Vector(s) of the present invention e.g.
  • non ⁇ viral DNA or RNA vectors may be designed in silico, and then synthesised by conventional polynucleotide synthesis techniques.
  • Non ⁇ viral plasmids cannot replicate in the subject to be treated, as they lack the viral genetic material which hijacks the body's normal production machinery. However they are capable of replicating in appropriate host cells, such as yeasts or bacteria including E. coli, and particularly airway cells as defined herein.
  • the term "plasmid” as used herein refers to a construction comprised of genetic material designed to direct transformation of a targeted cell.
  • the plasmid contains a plasmid backbone.
  • a "plasmid backbone” as used herein contains multiple genetic elements positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transfected cells.
  • the plasmid backbone can contain one or more unique restriction sites within the backbone.
  • the plasmid may be capable of autonomous replication in a defined host or organism such that the cloned sequence is reproduced.
  • the plasmid can confer some well ⁇ defined phenotype on the host organism which is either selectable or readily detected.
  • the plasmid or plasmid backbone may have a linear or circular configuration.
  • the components of a plasmid can contain, but is not limited to, a DNA molecule incorporating: (1) the plasmid backbone; (2) a sequence comprising or consisting of an SFTPB promoter fragment; (3) a transgene sequence encoding a (therapeutic) protein; and optionally (4) additional regulatory elements for transcription, translation, RNA stability and replication.
  • the purpose of the plasmid in human gene therapy for the efficient delivery of nucleic acid sequences to, and expression of therapeutic proteins in, a cell or tissue.
  • the purpose of the plasmid is to achieve high copy number, avoid potential causes of plasmid instability and provide a means for plasmid selection.
  • the nucleic acid cassette contains the necessary elements for expression of the nucleic acid within the cassette. Expression includes the efficient transcription of an inserted gene, nucleic acid sequence, or nucleic acid cassette with the plasmid.
  • a DNA plasmid may be CpG ⁇ free, or be optimised to reduce CpG dinucleotides as described herein.
  • a DNA plasmid of the invention may be codon ⁇ optimised as described herein. Methods of preparing plasmid DNA are well known in the art. Typically, they are capable of autonomous replication in an appropriate host or producer cell.
  • the term "exosome” as used herein refers to an extracellular vesicle formed by exocytosis from a cell of origin.
  • An exosome typically comprises a nucleic cassette of the invention. Exosomes may be used to transform a targeted cell.
  • the nucleic acid within an exosome may contain one or more genetic elements, which may be positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transformed cells.
  • the term "microvesicle” (MV) as used herein refers to an extracellular vesicle formed typically between about 30 to about 1,000 nm in diameter.
  • An MV typically comprises a nucleic cassette of the invention. Exosomes may be used to transform a targeted cell.
  • the nucleic acid within an MV may contain one or more genetic elements, which may be positionally and sequentially oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when necessary translated in the transformed cells.
  • Host cells containing (e.g. transformed, transfected, or electroporated with) the plasmid may be prokaryotic or eukaryotic in nature, either stably or transiently transformed, transfected, or electroporated with the plasmid.
  • Suitable host cells include bacterial, yeast, fungal, invertebrate, and mammalian cells.
  • the host cell is bacterial; more preferably E. coli.
  • Host cells can then be used in methods for the large scale production of the plasmid.
  • the cells are grown in a suitable culture medium under favourable conditions, and the desired plasmid isolated from the cells, or from the medium in which the cells are grown, by any purification technique well known to those skilled in the art; e.g. see Sambrook et al, supra.
  • Any appropriate delivery means can be used to deliver a non ⁇ viral vector (e.g. plasmid) of the invention to a target cell or patient.
  • Suitable delivery means are known in the art and within the routine skill of one of ordinary skill in the art.
  • Non ⁇ limiting examples include the use of cationic lipids, polymers (e.g. polyethyleneimine and poly ⁇ L ⁇ lysine) and electroporation.
  • cationic lipids may be used to deliver non ⁇ viral (e.g. plasmid) vectors of the invention to target cells or to a patient.
  • non ⁇ viral e.g. plasmid
  • cationic lipids suitable for use according to the invention are GL67A and lipofectamine.
  • the cationic lipid mixture GL67A is a mixture of three components ⁇ GL67 (Cholest ⁇ 5 ⁇ en ⁇ 3 ⁇ ol (3 ⁇ ) ⁇ ,3 ⁇ [(3 ⁇ aminopropyl)[4 ⁇ [(3 ⁇ aminopropyl)amino]butyl]carbamate], (CAS Number: 179075 ⁇ 30 ⁇ 0)), DOPE (1,2 ⁇ dioleoyl ⁇ sn ⁇ glycero ⁇ 3 ⁇ phosphoethanolamine) and DMPE ⁇ PEG5000 (1,2 ⁇ Dimyristoyl ⁇ sn ⁇ Glycero ⁇ 3 ⁇ Phosphoethanolamine ⁇ N ⁇ [methoxy (Polyethylene glycol)5000]). These components are formulated at a 1:2:0.05 molar ratio to form GL67A.
  • Lipofectamine consists of a 3:1 mixture of DOSPA (2,3 ⁇ dioleoyloxy ⁇ N ⁇ [2(sperminecarboxamido)ethyl] ⁇ N,N ⁇ dimethyl ⁇ 1 ⁇ propaniminium trifluoroacetate) and DOPE.
  • Viral vectors The present invention also provides a vector comprising a nucleic acid cassette of the invention.
  • the vector(s) may be present in the form of a therapeutic composition or formulation.
  • the vector may be a viral vector.
  • a viral vector of the invention may be a lentiviral vector, an adeno ⁇ associated virus (AAV) vector, an adenoviral vector, a poxvirus vector, a herpes simplex virus (HSV) vector. Derivatives of these viral vectors, such as lentivirus ⁇ derived particles are also encompassed within the invention.
  • adenoviral vectors include human serotypes such as AdHu5, simian serotypes such as ChAd63, ChAdOX1 or ChAdOX2, and other forms.
  • Non ⁇ limiting examples of poxvirus vectors include a modified vaccinia Ankara (MVA)).
  • ChAdOX1 and ChAdOX2 are disclosed in WO2012/172277 (herein incorporated by reference in its entirety).
  • ChAdOX2 is a BAC ⁇ derived and E4 modified AdC68 ⁇ based viral vector.
  • Viral vectors are usually non ⁇ replicating or replication impaired vectors, which means that the viral vector cannot replicate to any significant extent in normal cells (e.g. normal human cells), as measured by conventional means – e.g. via measuring DNA synthesis and/or viral titre.
  • Non ⁇ replicating or replication impaired vectors may have become so naturally (i.e. they have been isolated as such from nature) or artificially (e.g.
  • viral vector is incapable of causing a significant infection in an animal subject, typically in a mammalian subject such as a human or other primate.
  • Viral vector(s) of the present invention may be designed in silico, and then synthesised by conventional polynucleotide synthesis techniques.
  • the invention relates to retroviral vectors, particularly lentiviral vectors.
  • lentivirus refers to a family of retroviruses.
  • Retroviral/lentiviral vectors of the invention can integrate into the genome of transduced cells and lead to long ⁇ lasting expression.
  • retroviruses suitable for use in the present invention include gammaretroviruses such as murine leukaemia virus (MLV) and feline leukaemia virus (FLV).
  • lentiviruses suitable for use in the present invention include Simian immunodeficiency virus (SIV), Human immunodeficiency virus (HIV), Feline immunodeficiency virus (FIV), Equine infectious anaemia virus (EIAV), and Visna/maedi virus.
  • a particularly preferred lentiviral vector is an SIV vector (including all strains and subtypes), such as a SIV ⁇ AGM (originally isolated from African green monkeys, Cercopithecus aethiops).
  • the retroviral/lentiviral (e.g. SIV) vectors of the present invention are typically pseudotyped with hemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, or with G glycoprotein from Vesicular Stomatitis Virus (G ⁇ VSV).
  • HN hemagglutinin ⁇ neuraminidase
  • F fusion
  • G ⁇ VSV Vesicular Stomatitis Virus
  • the lentiviral (e.g. SIV) vectors of the present invention are pseudotyped with HN and F from a respiratory paramyxovirus.
  • the respiratory paramyxovirus is a Sendai virus (murine parainfluenza virus type 1).
  • the F protein may be a truncated F protein, typically one in which the cytoplasmic domain is truncated.
  • the truncated F protein is Fct4, in which 38 amino acids have been truncated from the C ⁇ terminus of the F protein, with 4 amino acids of the F protein cytoplasmic domain being retained.
  • the F protein may be a truncated F protein, typically one in which the cytoplasmic domain is truncated.
  • the truncated F protein is Fct4, in which 38 amino acids have been truncated from the C ⁇ terminus of the F protein, with 4 amino acids of the F protein cytoplasmic domain being retained.
  • a retroviral/lentiviral (e.g. SIV) vector for use according to the invention may be integrase ⁇ competent (IC).
  • the lentiviral (e.g. SIV) vector may be integrase ⁇ deficient (ID).
  • Viral vectors of the invention, particularly retroviral/lentiviral (e.g. SIV) vectors as described herein may transduce one or more cells types as described herein to achieve long term transgene expression.
  • the HN protein may be a truncated and/or chimeric HN protein, typically one in which the cytoplasmic domain is truncated or substituted.
  • the HN protein is a chimeric HN protein in which (i) the cytoplasmic domain of the HN is replaced by the cytoplasmic domain of the transmembrane (TMP) protein; or (ii) the cytoplasmic domain of the TMP is added to the cytoplasmic domain of the HN protein.
  • TMP transmembrane
  • the HN protein may be as described in Kobayashi et al. (J. Virol. (2003) 77(4):2607 ⁇ 2614), which is herein incorporated by reference in its entirety.
  • the viral vectors of the invention particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention enable high levels of transgene expression. Together with the increased levels of expression/secretion/membrane insertion of a therapeutic protein resulting from the use of an SFTPB promoter fragment of the invention, these viral vectors typically result in high levels (therapeutic levels) of expression of the transgene, and the (therapeutic) protein encoded by said transgene.
  • retroviral/lentiviral vectors of the present invention enable high levels of transgene expression. Together with the increased levels of expression/secretion/membrane insertion of a therapeutic protein resulting from the use of an SFTPB promoter fragment of the invention, these viral vectors typically result in high levels (therapeutic levels) of expression of the transgene, and the (therapeutic) protein encoded by said transgene.
  • the transgene to be included in a viral vector of the invention may be modified to facilitate expression.
  • the transgene sequence may be in CpG ⁇ depleted /low (or CpG ⁇ fee) and/or codon ⁇ optimised form to facilitate gene expression. Standard techniques for modifying the transgene sequence in this way are known in the art.
  • the viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention exhibit enhanced expression of the transgene. Accordingly, the viral vectors of the invention, particularly the retroviral/lentiviral (e.g.
  • SIV vectors of the invention are capable of producing long ⁇ lasting, repeatable, high ⁇ level transgene expression, particularly in lung parenchyma without inducing side effects (e.g., an undue immune response).
  • the viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention enable long ⁇ term transgene expression, resulting in long ⁇ term expression (and secretion or membrane insertion) of a (therapeutic) protein by cells of the lung parenchyma as described herein.
  • Long ⁇ term expression means expression of a transgene gene and/or encoded (therapeutic) protein, preferably at therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more.
  • long ⁇ term expression means expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days or more.
  • the invention relates to the use of F/HN lentiviral vectors comprising a nucleic acid cassette of the invention, particularly SIV F/HN vectors.
  • the nucleic acid cassette comprised in a viral vector of the invention, particularly a retroviral/lentiviral (e.g. SIV) vector of the invention may have no intron positioned between the promoter and the nucleic acid encoding the signal peptide and/or the nucleic acid encoding the therapeutic protein.
  • the viral vectors of the invention may be made using any suitable process known in the art.
  • retroviral/lentiviral (e.g. SIV) vectors of the invention may be made using the methods disclosed in UK Application No. 2102832.9, which is herein incorporated by reference in its entirety).
  • the viral vectors of the invention may comprise a central polypurine tract (cPPT) and/or the Woodchuck hepatitis virus posttranscriptional regulatory elements (WPRE).
  • cPPT central polypurine tract
  • WPRE Woodchuck hepatitis virus posttranscriptional regulatory elements
  • An exemplary WPRE sequence is provided by SEQ ID NO: 52.
  • Transgenes A nucleic acid cassette of the invention comprises a transgene. Typically, the transgene encodes a therapeutic protein. A therapeutic protein is one which has potential utility in the treatment or prevention of a disease or condition, such as those describe herein.
  • a nucleic acid cassette of the invention may comprise a nucleic acid encoding a protein which has a therapeutic effect on a disease or condition to be treated.
  • a nucleic acid cassette of the invention may comprise a nucleic acid encoding a therapeutic protein which is a functional or wild ⁇ type form of a protein which is present in a patient to be treated in a dysfunctional form (whether the dysfunction is inherent or acquired).
  • the phrase "inherent dysfunction” refers to a protein which is innately dysfunctional due to genetic factors and the phrase “acquired dysfunction” refers to a protein which is dysfunctional due to environmental or other factors after birth.
  • a nucleic acid cassette of the invention may comprise a nucleic acid encoding a protein which is a functional or wild ⁇ type form of a protein which is present in a patient, but which that has become dysfunctional due to a genetic disease, such as a genetic respiratory disease.
  • the nucleic acid cassettes of the present invention are useful in the treatment of diseases via their use in expressing therapeutic proteins in lung parenchyma as described herein (e.g. ATI, ATII cells).
  • the nucleic acid cassettes of the present invention are useful in the treatment of diseases via their use in expressing therapeutic proteins in ATII cells, ATI cells and/or club cells.
  • the transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally Surfactant Protein B (SFTPB), Surfactant Protein C (SFTPC), alpha ⁇ 1 ⁇ antitrypsin (AAT), Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, von Willebrand Factor, Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF), decorin, an anti ⁇ inflammatory protein (e.g.
  • the therapeutic protein is not an antibody, particularly not a monoclonal antibody.
  • the transgene may encode a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally SFTPB, SP ⁇ C, AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, von Willebrand Factor, GM ⁇ CSF, an anti ⁇ inflammatory protein (e.g.
  • the nucleic acid cassettes of the invention are particularly efficient at driving the expression, secretion and/or membrane insertion of proteins (e.g. therapeutic proteins as described herein) by the lung parenchyma. This is particularly the case when such cassettes are comprised within F/HN pseudotyped viral vectors of the invention (as described herein), which are efficient at targeting cells in the lung parenchyma.
  • the nucleic acid cassettes of the invention and vectors comprising said cassettes
  • nucleic acid cassettes of the invention are typically delivered to lung parenchyma as described herein. Accordingly, the nucleic acid cassettes of the invention (and vectors comprising said cassettes) are particularly suited for treatment of diseases or disorders of the lung parenchyma. Typically, the nucleic acid cassettes of the invention (and vectors comprising said cassettes) may be used for the treatment of a genetic respiratory disease.
  • a nucleic acid cassette of the invention (or vector comprising said cassette) may comprise a nucleic acid encoding a polypeptide or protein that is therapeutic for the treatment of such diseases, particularly a disease or disorder of the lung parenchyma.
  • a nucleic acid cassette of the invention may comprise a nucleic acid sequence encoding a therapeutic protein selected from: (a) a secreted therapeutic protein, optionally SFTPB, SFTPC, AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, von Willebrand Factor, GM ⁇ CSF, an anti ⁇ inflammatory protein (e.g.
  • IL ⁇ 10, TGG ⁇ , or TNF ⁇ alpha or monoclonal antibody, an anti ⁇ inflammatory decoy and a monoclonal antibody against an infectious agent; or (b) ABCA3, TRIM72, CSF2RA, CSF2RB or DCN.
  • Other preferred examples of therapeutic proteins that may be encoded by a nucleic acid sequence comprised in a nucleic acid cassette of the invention (or vector comprising said cassette) include genes related to or associated with other surfactant deficiencies.
  • the therapeutic protein encoded by a nucleic acid cassette of the invention (or vector comprising said cassette) may be SFTPB. Examples of an SFTPB therapeutic transgene are provided by SEQ ID NOs: 53, 54 and 56.
  • SFTPB transgene An exemplary codon ⁇ optimised SFTPB transgene is provided by SEQ ID NO: 55.
  • the therapeutic protein encoded by said SFTPB transgene may be exemplified by the polypeptide of SEQ ID NO: 57. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 53 to 57.
  • the transgene may encode ABCA3. Examples of a ABACA3 transgene are provided by SEQ ID NOs: 58 and 59.
  • An exemplary codon ⁇ optimised ABACA3 transgene is provided by SEQ ID NO: 60.
  • the polypeptide encoded by said ABACA3 transgene may be exemplified by the polypeptide of SEQ ID NO: 61. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 58 to 61.
  • the therapeutic protein encoded by a nucleic acid cassette of the invention may be SFTPC.
  • An example of an SFTPC therapeutic transgene is provided by SEQ ID NO: 62.
  • the therapeutic protein encoded by said SFTPC transgene may be exemplified by the polypeptide of SEQ ID NO: 63.
  • variants thereof are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 62 or 63.
  • the therapeutic protein encoded by a nucleic acid cassette of the invention may be an AAT.
  • An example of an AAT therapeutic transgene (SERPINA1) is provided by SEQ ID NO: 70.
  • SEQ ID NO: 70 is a codon ⁇ optimized CpG depleted AAT transgene (SERPINA1) previously designed by the present inventors to enhance translation in human cells. Such optimisation has been shown to enhance gene expression by up to 15 ⁇ fold.
  • variants of same sequence which possess the same technical effect of enhancing translation compared with the unmodified (wild ⁇ type) AAT gene sequence are also encompassed by the present invention.
  • the therapeutic protein encoded by said AAT transgene may be exemplified by the polypeptide of SEQ ID NO: 71. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 70 or 71.
  • the therapeutic protein encoded by a nucleic acid cassette of the invention may be an FVIII.
  • FVIII therapeutic transgene examples are provided by SEQ ID NOs: 72 and 73.
  • the polypeptide encoded by the FVIII transgene may be exemplified by the polypeptide of SEQ ID NO: 74 and 75. Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 72 to 75.
  • the therapeutic protein encoded by a nucleic acid cassette of the invention may be GM ⁇ CSF.
  • a GM ⁇ CSF transgene may comprise or consist of SEQ ID NO: 64 (human).
  • the polypeptide encoded by the GM ⁇ CSF transgene may be exemplified by the polypeptide of SEQ ID NO: 65 (human). Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 64and 65.
  • the transgene may encode decorin.
  • An example of a DCN transgene is provided by SEQ ID NO: 66.
  • the polypeptide encoded by said DCN transgene may be exemplified by the polypeptide of SEQ ID NO: 67.
  • Variants thereof are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to SEQ ID NO: 66or 67.
  • the transgene may encode TRIM72.
  • An example of a TRIM72 transgene is provided by SEQ ID NO: 68.
  • the polypeptide encoded by said TRIM72 transgene may be exemplified by the polypeptide of SEQ ID NO: 69.
  • Variants thereof (as described therein) are also included, particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to SEQ ID NO: 68or 69.
  • the therapeutic protein encoded by a nucleic acid cassette of the invention may be encoded by any one of SFTPB, SFTPC, Factor V, Factor VII, Factor IX, Factor X and/or Factor XI, von Willebrand Factor, GM ⁇ CSF, ABCA3, TRIM72 or DCN, or other known related gene.
  • the transgene may preferably be SFTPB, SFTPC, ABCA3 or GM ⁇ CSF.
  • the therapeutic protein may be a monoclonal antibody (mAb) against an infectious agent (bacterial, fungal or viral, e.g.
  • the therapeutic protein may be anti ⁇ TNF alpha.
  • the therapeutic protein may be one implicated in an inflammatory, immune or metabolic condition.
  • a nucleic acid cassette of the invention (or a vector comprising said cassette) may be delivered to one or more cell/cell type of the lung parenchyma to allow production of proteins to be secreted into circulatory system.
  • the therapeutic protein may be any one of Factor VII, Factor VIII, Factor IX, Factor X, Factor XI and/or von Willebrand’s factor.
  • nucleic acid cassette of the invention may be used in the treatment of diseases, particularly cardiovascular diseases and blood disorders, preferably blood clotting deficiencies such as haemophilia.
  • the therapeutic protein may be an mAb against an infectious agent or a protein implicated in an inflammatory, immune or metabolic condition, such as, lysosomal storage disease.
  • the nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the promoter and the nucleic acid encoding the therapeutic protein.
  • nucleic acid cassette when the nucleic acid cassette is comprised in a viral vector, there may be no intron between the promoter and the transgene in the vector genome (pDNA1) plasmid used to make said viral vector, as described herein.
  • said nucleic acid cassette of the invention (or a vector comprising said cassette) may have an intron positioned between the promoter and the transgene, particularly if the cassette (or vector comprising said cassette) is non ⁇ viral.
  • the nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise a SFTPB promoter fragment and an SFTPB transgene, including those described herein.
  • nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an SFTPC transgene, including those described herein.
  • said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an AAT transgene (SERPINA1), including those described herein.
  • nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FVIII transgene, including those described herein.
  • said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FVII transgene, including those described herein.
  • nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FIX transgene, including those described herein.
  • said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FX transgene, including those described herein.
  • nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and an FXI transgene, including those described herein.
  • said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and a von Willebrand Factor transgene, including those described herein.
  • nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention may comprise a SFTPB promoter fragment and a Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF) transgene, including those described herein.
  • GM ⁇ CSF Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor
  • said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the lentiviral e.g.
  • SIV vector comprises a SFTPB promoter fragment and an DCN transgene, including those described herein.
  • said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the lentiviral (e.g. SIV) vector comprises a SFTPB promoter fragment and a TRIM72 transgene, including those described herein.
  • said nucleic acid cassette of the invention (or a vector comprising said cassette) may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the lentiviral e.g.
  • SIV vector comprises a SFTPB promoter fragment and a ABACA3 transgene, including those described herein.
  • said nucleic acid cassette of the invention may have no intron positioned between the SFTPB promoter fragment and the transgene.
  • the nucleic acid cassette of the invention (or a vector comprising said cassette) comprises a nucleic acid encoding a therapeutic protein (said nucleic acid is referred to interchangeably herein as a transgene).
  • the nucleic acid sequence encodes a gene product, e.g., a protein, particularly a therapeutic protein.
  • the nucleic acid cassette of the invention (or a vector comprising said cassette) comprises a nucleic acid sequence encoding an SFTPB, SFTPC, AAT, GM ⁇ CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 and said nucleic acid sequence comprises (or consists of) a nucleic acid sequence having at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100%) sequence identity to the SFTPB, SFTPC, AAT, GM ⁇ CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 nucleic acid sequence respectively, examples of which are described herein.
  • the nucleic acid sequence encoding SFTPB, SFTPC, AAT, GM ⁇ CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 may preferably comprise (or consist of) a nucleic acid sequence having at least 95% (such as at least 95, 96, 97, 98, 99 or 100%) sequence identity to the SFTPB, SFTPC, AAT, GM ⁇ CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 nucleic acid sequence respectively, examples of which are described herein.
  • the amino acid sequence of the (therapeutic) protein encoded by the transgene may be a functional variant having at least 95% (such as at least 95, 96, 97, 98, 99 or 100%) sequence identity to the functional protein.
  • SFTPB promoter fragment and/or enhancer of the invention may be linked to a transgene by a linker.
  • Said linker is typically a short DNA sequence, as defined herein, and may comprise or consist of one or more restriction enzyme site, examples of which are also described herein.
  • a SFTPB promoter fragment of the invention may be joined to a transgene by a linker which comprises or consists of a NheI restriction site and/or a BgIII restriction site, preferably a NheI restriction site.
  • Signal peptides The transgene encoding for a (therapeutic) protein may further comprise a nucleic acid sequence encoding for a signal peptide.
  • Said signal peptide may be the endogenous signal peptide of the (therapeutic) protein, or a signal peptide exogenous to said (therapeutic) protein.
  • the transgene further comprises a nucleic acid encoding for an exogenous signal peptide
  • the transgene preferably exclude a nucleic acid sequence encoding for the endogenous signal peptide.
  • the exogenous signal peptide is typically the sole signal peptide linked with (and hence driving secretion and/or membrane insertion) of the therapeutic protein. All disclosure herein relates to both transgenes and therapeutic proteins including and excluding endogenous signal peptides unless explicitly stated.
  • sequence identity of variants, and/or lengths of fragments may be based on the sequence with or without a signal peptide.
  • Any signal peptide and therapeutic combination may be used, provided that this combination is effective in increasing the expression, secretion and/or membrane insertion of a (therapeutic) protein as defined herein.
  • Selection of a signal peptide may depend on the specific (therapeutic) protein and/or the specific lung parenchymal cell type by which the (therapeutic) protein is to be expressed/secreted/inserted into the cell membrane.
  • lentiviral vectors such as the lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus
  • HN haemagglutinin ⁇ neuraminidase
  • F fusion
  • the invention therefore provides lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered in combination with a surfactant.
  • HN haemagglutinin ⁇ neuraminidase
  • F fusion
  • the invention therefore provides lentiviral vector pseudotyped with haemagglutinin ⁇ neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus and comprising a transgene operably linked to a promoter for use in a method of treating a disease, wherein the lentiviral vector is administered simultaneously or sequentially with a surfactant.
  • the lentiviral (e.g. SIV) vector and surfactant are administered in combination.
  • Administered "in combination,” encompasses both simultaneous (also referred to as concurrent) administration/delivery and sequential (also referred to as separate) administration/delivery.
  • the delivery of the lentiviral (e.g. SIV) vector may still be occurring when the delivery of the surfactant begins, or the delivery of surfactant may still be occurring when the delivery of the lentiviral (e.g. SIV) vector begins, so that there is overlap in terms of administration.
  • Simultaneous delivery may encompass delivery of the lentiviral (e.g. SIV) vector and surfactant within weeks to months or even years of each other, typically so that the lentiviral (e.g. SIV) vector delivery overlaps with the delivery of the surfactant.
  • the delivery of the lentiviral (e.g. SIV) vector may still be occurring when the delivery of the surfactant begins, or the delivery of surfactant may still be occurring when the delivery of the lentiviral (e.g. SIV) vector begins, so that there is overlap in terms of administration.
  • Simultaneous delivery may encompass delivery of the lentiviral (e.g. SIV) vector and surfactant within weeks to months or even
  • SIV vector may end before the delivery of the surfactant begins, or the delivery of the surfactant may end before delivery of the lentiviral (e.g. SIV) vector begins.
  • Sequential administration may involve the lentiviral (e.g. SIV) vector and surfactant being administered within 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 6 hours, 12 hours or 24 hours or longer of each other.
  • the lentiviral vector may be administered before the surfactant.
  • the surfactant may be administered before the lentiviral vector.
  • the lentiviral vector and surfactant may be administered simultaneously, optionally wherein the lentiviral vector and surfactant are mixed prior to administration.
  • the treatment is more effective because of combined administration.
  • treatment with the lentiviral e.g.
  • SIV vector may be more effective, e.g., an equivalent effect is seen with less of the lentiviral (e.g. SIV) vector, or the lentiviral (e.g. SIV) vector reduces symptoms to a greater extent, than would be seen if the lentiviral (e.g. SIV) vector were administered in the absence of the surfactant.
  • treatment with the surfactant may be more effective, e.g., an equivalent effect is seen with less of the surfactant, or the surfactant reduces symptoms to a greater extent, than would be seen if the surfactant were administered in the absence of the lentiviral (e.g. SIV) vector.
  • a combination therapy of the invention may increase transgene expression by at least 1.2 fold, at least 1.3 fold, at least 1.4 fold, at least 1.5 fold, at least 2 fold, at least 2.5 fold or more compared with treatment with the lentiviral (e.g. SIV) vector alone (i.e. compared with the increase in transgene expression achieved when treating with the lentiviral (e.g. SIV) alone).
  • the surfactant may aid the distribution of the lentiviral (e.g. SIV) vector within the lungs, enabling it to penetrate more deeply into the respiratory tree and thus facilitating transduction of the lung parenchyma.
  • appropriate dosage of the lentiviral (e.g. SIV) vector and/or the surfactant will depend on the specific agent, and can also vary from patient to patient. Any surfactant may be used in a combination therapy according to the present invention. It will be appreciated that it is within the routine practice of one of ordinary skill in the art to select such a surfactant.
  • a disease to be treated with such a lentiviral vector and a surfactant may be a genetic disease.
  • the disease to be treated may be a respiratory disease, particularly a genetic respiratory disease; or a cardiovascular disease or blood disorder, particularly a genetic cardiovascular disease or blood disorder.
  • Non ⁇ limiting examples of diseases which may be treated according to this aspect of the invention include Surfactant Protein B (SP ⁇ B) Deficiency; Surfactant Protein C (SP ⁇ C) deficiency; ABCA3 deficiency; Pulmonary surfactant metabolism dysfunction 2 (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD); Alpha 1 ⁇ antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID ⁇ 19; a pulmonary fibrotic disease; a pulmonary allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in the lungs; and haemophilia.
  • SP ⁇ B Surfactant Protein B
  • SP ⁇ C Surfactant Protein C
  • ABCA3 deficiency Pulmonary surfactant metabolism dysfunction 2
  • SMDP3 Pulmonary surfactant metabolism dysfunction 3
  • a lentiviral vector for use in combination with a surfactant may comprise HN and F proteins from a Sendai virus, as described herein.
  • a lentiviral vector for use in combination with a surfactant may be selected from the group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector.
  • said lentiviral vector may be a SIV vector.
  • the transgene may encode any suitable therapeutic protein as described herein.
  • said transgene may be selected from: (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP ⁇ B), Surfactant Protein C (SP ⁇ C), AAT, Factor VIII, Factor VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte ⁇ Macrophage Colony ⁇ Stimulating Factor (GM ⁇ CSF), decorin, an anti ⁇ inflammatory protein (e.g.
  • a lentiviral vector for use in combination with a surfactant may comprise a SFTPB promoter fragment as defined herein.
  • said lentiviral vector may comprise a nucleic acid cassette as defined herein.
  • Therapeutic Indications The nucleic acid cassettes and vectors of the present invention enable cell ⁇ preferred or cell ⁇ specific expression of a transgene encoding a (therapeutic) protein.
  • nucleic acid cassette or vector facilitating efficient transgene expression.
  • the nucleic acid cassettes and vectors of the invention, and particularly the F/HN ⁇ pseudotyped retroviral/lentiviral (e.g. SIV) vectors of the invention are capable of: (i) transduction of one or more cell/cell type of the lung parenchyma without disruption of epithelial integrity; (ii) persistent gene expression; (iii) lack of chronic toxicity; and/or (iv) efficient repeat administration. Long term/persistent stable gene expression, preferably at a therapeutically ⁇ effective level, may be achieved using repeat doses of a nucleic acid cassette or vector of the present invention.
  • the nucleic acid cassettes and vectors of the present invention and particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention can be used in gene therapy.
  • the present invention provides a nucleic acid cassette or gene therapy vector as defined herein for use in a method of treating or preventing a disease.
  • the disease to be treated may be chronic or acute.
  • the nucleic acid cassettes and vectors (viral and non ⁇ viral) of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention may be used to deliver any transgene useful in gene therapy.
  • the nucleic acid cassettes and vectors (viral and non ⁇ viral) of the invention are for use in gene therapy for the treatment of a disease or disorder of the lung parenchyma.
  • efficient airway cell uptake properties of the nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention make them highly suitable for treating respiratory or lung diseases, particularly genetic respiratory diseases, particularly preferably those of/involving the lung parenchyma.
  • SIV vectors of the invention can also be used in methods of gene therapy to promote secretion of therapeutic proteins.
  • the invention provides secretion of therapeutic proteins into alveoli or lumen of the bronchioles.
  • Administration of a nucleic acid cassettes and vectors of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention and its uptake by airway cells may be used to enable the use of the lungs as a “factory” to produce a therapeutic protein that is then secreted and enters the general circulation at therapeutic levels, where it can travel to cells/tissues of interest to elicit a therapeutic effect.
  • nucleic acid cassettes and vectors of the present invention can also be treated by the nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention.
  • Nucleic acid cassettes and vectors of the present invention, and particularly the retroviral/lentiviral (e.g. SIV) vectors of the invention can effectively treat a disease by providing a transgene for the correction of the disease. For example, resulting in the expression and secretion of SFTPB from cells of the lung parenchyma, to compensate for the pathologically low levels of SFTPB expression in patients with SFTPB deficiency.
  • nucleic acid cassettes and vectors of the present invention may be used to treat alpha ⁇ 1 ⁇ antitrypsin (AAT) deficiency, typically by gene therapy with a AAT transgene (SERPINA1) as described herein.
  • AAT alpha ⁇ 1 ⁇ antitrypsin
  • SERPINA1 AAT transgene
  • AAT is a secreted anti ⁇ protease that is produced mainly in the liver and then trafficked to the lung, with smaller amounts also being produced in the lung itself.
  • the main function of AAT is to bind and neutralise/inhibit neutrophil elastase.
  • Gene therapy with AAT according to the present invention is relevant to AAT deficient patient, as well as in other lung diseases such as CF or chronic obstructive pulmonary disease (COPD), and offers the opportunity to overcome some of the problems encountered by conventional enzyme replacement therapy (in which AAT isolated from human blood and administered intravenously every week), providing stable, long ⁇ lasting expression in the target tissue (lung/nasal epithelium), ease of administration and unlimited availability.
  • Transduction with a nucleic acid cassettes and vectors of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention may lead to secretion of the recombinant protein into the lumen of the lung as well as into the circulation.
  • AAT gene therapy may therefore also be beneficial in other disease indications, non ⁇ limiting examples of which include type 1 and type 2 diabetes, acute myocardial infarction, ischemic heart disease, rheumatoid arthritis, inflammatory bowel disease, transplant rejection, graft versus host (GvH) disease, multiple sclerosis, liver disease, cirrhosis, vasculitides and infections, such as bacterial and/or viral infections.
  • AAT has numerous other anti ⁇ inflammatory and tissue ⁇ protective effects, for example in pre ⁇ clinical models of diabetes, graft versus host disease and inflammatory bowel disease.
  • AAT in the lung and/or nose following transduction according to the present invention may, therefore, be more widely applicable, including to these indications.
  • diseases that may be treated with gene therapy of a secreted protein according to the present invention include cardiovascular diseases and blood disorders, particularly blood clotting deficiencies such as haemophilia (A, B or C), von Willebrand disease and Factor VII deficiency.
  • the disease to be treated is selected from a surfactant protein deficiency, such as Surfactant Protein B (SFTPB) Deficiency, Surfactant Protein C (SFTPC) deficiency, ABCA3 deficiency, Pulmonary surfactant metabolism dysfunction 2 (SMDP2) Pulmonary surfactant metabolism dysfunction 3 (SMDP3), or other surfactant deficiencies; Primary Ciliary Dyskinesia (PCD); Alpha 1 ⁇ antitrypsin Deficiency (A1AD); Pulmonary Alveolar Proteinosis (PAP, hereditary and/or acquired); Chronic obstructive pulmonary disease (COPD); Acute respiratory distress syndrome (ARDS); COVID ⁇ 19; a pulmonary fibrotic disease (including idiopathic pulmonary fibrosis); a pulmonary allergic condition; a pulmonary bacterial infection; asthma; lung cancer; a dysplastic change in the lungs; and haemophilia.
  • SFTPB Surfactant Protein B
  • SFTPC Surfactant Protein C
  • diseases or disorders to be treated include Primary Ciliary Dyskinesia (PCD), acute lung injury, and/or inflammatory, infectious, immune or metabolic conditions, such as lysosomal storage diseases or a pulmonary bacterial infection, or any other lung disease or disorder.
  • PCD Primary Ciliary Dyskinesia
  • the nucleic acid cassettes and vectors (viral and non ⁇ viral) of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention typically provide high expression levels of a transgene of interest, and the (therapeutic) protein encoded thereby, when administered to a patient.
  • high expression and therapeutic expression are used interchangeably herein.
  • Expression may be measured by any appropriate method (qualitative or quantitative, preferably quantitative), and concentrations given in any appropriate unit of measurement, for example ng/ml or ⁇ M. Expression/secretion/membrane insertion of a transgene, or the (therapeutic) protein of interest encoded thereby may be given in absolute terms.
  • expression/secretion/membrane insertion of a therapeutic protein may be given in relative terms, for example relative to the expression/secretion/membrane insertion of the therapeutic protein encoded by a corresponding nucleic acid cassette or vector of the invention without the SFTB promoter fragment of the invention, relative to the expression/secretion/membrane insertion of the same transgene using the full ⁇ length SFTB promoter as described herein, or relative to the expression/secretion/membrane insertion of the corresponding endogenous (defective) gene.
  • Expression may be measured in terms of mRNA or protein expression.
  • the expression of the therapeutic protein of the invention may be quantified relative to the endogenous protein or gene in terms of protein concentration, mRNA copies per cell or any other appropriate unit.
  • Secretion and/or membrane insertion of a therapeutic protein may be quantified relative to secretion/membrane insertion of the corresponding endogenous protein, or relative to the level of secretion/membrane insertion of the therapeutic protein introduced via an expression cassette with the same transgene but without the SFTB promoter fragment of the invention or comprising the full ⁇ length SFTB promoter as described herein.
  • Expression levels of a nucleic acid encoding a (therapeutic) protein and/or the expression/secretion/membrane insertion of the encoded (therapeutic) protein of the invention may be measured ex vivo (e.g. in the conditioned media used to culture the cells or within the cells themselves) or in vivo (e.g. in the lung tissue, epithelial lining fluid and/or serum/plasma) as appropriate.
  • a high and/or therapeutic expression level may therefore refer to the concentration in the lung, epithelial lining fluid and/or serum/plasma.
  • SIV vectors of the present invention may be administered twice ⁇ daily, daily, twice ⁇ weekly, weekly, monthly, every two months, every three months, every four months, every six months, yearly, every two years, or more. Dosing may be continued for as long as required, for example, for at least six months, at least one year, two years, three years, four years, five years, ten years, fifteen years, twenty years, or more, up to for the lifetime of the patient to be treated.
  • the invention also provides nucleic acid cassettes and vectors of the present invention, and particularly retroviral/lentiviral (e.g.
  • SIV vectors of the invention as described herein for use in a method of gene therapy comprising the steps of: (a) transducing cells (e.g. lung parenchyma) ex vivo to produce modified cells expressing a transgene of interest; and (b) administering the resulting modified cells.
  • the invention provides a method of treating a disease, the method comprising administering a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention to a subject. Any disease described herein may be treated according to the invention.
  • the invention provides a method of treating a lung disease using a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention.
  • the disease to be treated may be a chronic disease.
  • the invention also provides a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention as described herein for use in a method of treating a disease. Any disease described herein may be treated according to the invention.
  • the invention provides a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention for use in a method of treating a lung disease.
  • the disease to be treated may be a chronic disease.
  • the invention also provides the use of a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention as described herein in the manufacture of a medicament for use in a method of treating a disease. Any disease described herein may be treated according to the invention.
  • the invention provides the use of a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention for the manufacture of a medicament for use in a method of treating a lung disease.
  • the disease to be treated may be a chronic disease.
  • the invention also provides a cell comprising a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector.
  • Said cell may be a lung parenchyma cell as described herein.
  • the invention further provides a method of expressing a transgene, typically a transgene encoding a (therapeutic) protein in a target cell, comprising delivering a nucleic acid cassette or a vector of the invention into the target cells. Said method may be carried out in vitro, ex vivo, or in vivo, preferably in vitro or ex vivo.
  • the target cell may be any appropriate cell type, such as those described herein.
  • the target cells may be prokaryotic or eukaryotic, preferably eukaryotic. Particularly preferred are mammalian cells, such as human, non ⁇ human primate, mouse, rat, dog, cat, horse, or cow cells.
  • the cells may be primary cells or cell lines. Non ⁇ limiting examples of cells include ATII cells, ATI cells, club cells and/or HEK293T cells.
  • the step of delivering the nucleic acid cassette or vector may comprise integrating said nucleic acid cassette or gene therapy vector into said target cell's genome. Any appropriate technique may be used to deliver the nucleic acid cassette or vector, examples of which are known in the art and within the routine practice of one of ordinary skill in the art.
  • the method may further comprise a step of culturing cells expressing the transgene, and/or isolating or purifying the expressed (therapeutic) protein from said cells.
  • Any and all disclosure herein in relation to nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention applies equally and without reservation to the therapeutic uses and methods described herein. Further, for the avoidance of doubt, any and all disclosure herein in relation to therapeutic uses and methods using nucleic acid cassettes or vectors of the present invention, applies equally and without reservation to the combination therapies of lentiviral (e.g. SIV) vectors and surfactants as described herein.
  • Long term/persistent stable gene expression may be achieved using repeat doses of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention of the present invention.
  • a single dose may be used to achieve the desired long ⁇ term expression.
  • Formulation and administration also provides a composition comprising a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention, and optionally a pharmaceutically acceptable carrier, excipient, buffer or diluent.
  • the nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral e.g.
  • SIV vectors of the invention may be administered in any dosage appropriate for achieving the desired therapeutic effect.
  • Appropriate dosages may be determined by a clinician or other medical practitioner using standard techniques and within the normal course of their work.
  • suitable dosages of viral vectors of the invention include 1x10 8 transduction units (TU), 1x10 9 TU, 1x10 10 TU, 1x10 11 TU or more.
  • Non ⁇ limiting examples of suitable dosages of non ⁇ viral vectors/delivery means of the invention include a maximum of 30 mL per dose, a maximum of 25 mL per dose, a maximum of 20 mL per dose, a maximum of 15 mL per dose, a maximum of 10 mL per dose, or less, preferably a maximum of 20 mL per dose.
  • Non ⁇ limiting examples of pharmaceutically acceptable carriers that may be comprised in a composition of the invention include water, saline, and phosphate ⁇ buffered saline. In some embodiments, however, the composition is in lyophilized form, in which case it may include a stabilizer, such as bovine serum albumin (BSA).
  • BSA bovine serum albumin
  • compositions with a preservative such as thiomersal or sodium azide
  • a preservative such as thiomersal or sodium azide
  • the nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered by any appropriate route. It may be desired to direct the compositions of the present invention (as described above) to the respiratory system of a subject. Efficient transmission of a therapeutic/prophylactic composition or medicament to the site of a disease or disorder in the respiratory tract may be achieved by oral or intra ⁇ nasal administration, for example, as aerosols (e.g. nasal sprays), or by catheters.
  • aerosols e.g. nasal sprays
  • nucleic acid cassettes or vectors of the present invention and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention are stable in clinically relevant nebulisers, inhalers (including metered dose inhalers), catheters and aerosols, etc.
  • Other routes of administration including but not limited to i.v. administration, intranasal administration and intraplural injection are also encompassed by the present invention. Suitable administration routes are known in the art.
  • the nose is a preferred production site for a therapeutic protein using nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g.
  • SIV vectors of the invention for at least one of the following reasons: (i) extracellular barriers such as inflammatory cells and sputum are less pronounced in the nose; (ii) ease of vector administration; (iii) smaller quantities of vector required; and (iv) ethical considerations.
  • nasal administration of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may result in efficient (high ⁇ level) and long ⁇ lasting expression of the therapeutic protein of interest. Accordingly, nasal administration of nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be preferred.
  • Formulations for intra ⁇ nasal administration may be in the form of nasal droplets or a nasal spray.
  • An intra ⁇ nasal formulation may comprise droplets having approximate diameters in the range of 100 ⁇ 5000 ⁇ m, such as 500 ⁇ 4000 ⁇ m, 1000 ⁇ 3000 ⁇ m or 100 ⁇ 1000 ⁇ m.
  • the droplets may be in the range of about 0.001 ⁇ 100 ⁇ l, such as 0.1 ⁇ 50 ⁇ l or 1.0 ⁇ 25 ⁇ l, or such as 0.001 ⁇ 1 ⁇ l.
  • the aerosol formulation may take the form of a powder, suspension or solution. The size of aerosol particles is relevant to the delivery capability of an aerosol. Smaller particles may travel further down the respiratory airway towards the alveoli than would larger particles.
  • the aerosol particles have a diameter distribution to facilitate delivery along the entire length of the bronchi, bronchioles, and alveoli.
  • the particle size distribution may be selected to target a particular section of the respiratory airway, for example the alveoli.
  • the particles may have diameters in the approximate range of 0.1 ⁇ 50 ⁇ m, preferably 1 ⁇ 25 ⁇ m, more preferably 1 ⁇ 5 ⁇ m.
  • Aerosol particles may be for delivery using a nebulizer (e.g. via the mouth) or nasal spray.
  • An aerosol formulation may optionally contain a propellant and/or surfactant. The formulation of pharmaceutical aerosols is routine to those skilled in the art, see for example, Sciarra, J.
  • compositions comprising nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention, in particular where intranasal delivery is to be used, may comprise a humectant.
  • Suitable humectants include, for instance, sorbitol, mineral oil, vegetable oil and glycerol; soothing agents; membrane conditioners; sweeteners; and combinations thereof.
  • the compositions may comprise a surfactant.
  • Suitable surfactants include non ⁇ ionic, anionic and cationic surfactants. Examples of surfactants that may be used include, for example, polyoxyethylene derivatives of fatty acid partial esters of sorbitol anhydrides, such as for example, Tween 80, Polyoxyl 40 Stearate, Polyoxy ethylene 50 Stearate, fusieates, bile salts and Octoxynol.
  • nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be performed after an initial administration.
  • the administration may, for instance, be at least a week, two weeks, a month, two months, six months, a year or more after the initial administration.
  • nucleic acid cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered at least once a week, once a fortnight, once a month, every two months, every six months, annually or at longer intervals.
  • administration is every six months, more preferably annually.
  • nucleic acid cassettes or vectors of the present invention and particularly retroviral/lentiviral (e.g. SIV) vectors may, for instance, be administered at intervals dictated by when the effects of the previous administration are decreasing.
  • retroviral/lentiviral vectors e.g. SIV
  • any and all disclosure herein in relation formulations of nucleic acid cassettes or vectors of the present invention applies equally and without reservation to the formulations of lentiviral (e.g. SIV) vectors for combination therapies of lentiviral (e.g. SIV) vectors and surfactants as described herein.
  • SEQUENCE HOMOLOGY Any of a variety of sequence alignment methods can be used to determine percent identity, including, without limitation, global methods, local methods and hybrid methods, such as, e.g., segment approach methods. Protocols to determine percent identity are routine procedures within the scope of one skilled in the art. Global methods align sequences from the beginning to the end of the molecule and determine the best alignment by adding up scores of individual residue pairs and by imposing gap penalties. Non ⁇ limiting methods include, e.g., CLUSTAL W, see, e.g., Julie D.
  • Non ⁇ limiting methods include, e.g., Match ⁇ box, see, e.g., Eric Depiereux and Ernest Feytmans, Match ⁇ Box: A Fundamentally New Algorithm for the Simultaneous Alignment of Several Protein Sequences, 8(5) CABIOS 501 ⁇ 509 (1992); Gibbs sampling, see, e.g., C. E.
  • % sequence identity between two or more nucleic acid or amino acid sequences is a function of the number of identical positions shared by the sequences. Thus, % identity may be calculated as the number of identical nucleotides / amino acids divided by the total number of nucleotides / amino acids, multiplied by 100. Calculations of % sequence identity may also take into account the number of gaps, and the length of each gap that needs to be introduced to optimize alignment of two or more sequences.
  • a limited number of non ⁇ conservative amino acids, amino acids that are not encoded by the genetic code, and unnatural amino acids may be substituted for polypeptide amino acid residues.
  • the polypeptides of the present invention can also comprise non ⁇ naturally occurring amino acid residues.
  • Non ⁇ naturally occurring amino acids include, without limitation, trans ⁇ 3 ⁇ methylproline, 2,4 ⁇ methano ⁇ proline, cis ⁇ 4 ⁇ hydroxyproline, trans ⁇ 4 ⁇ hydroxy ⁇ proline, N ⁇ methylglycine, allo ⁇ threonine, methyl ⁇ threonine, hydroxy ⁇ ethylcysteine, hydroxyethylhomo ⁇ cysteine, nitro ⁇ glutamine, homoglutamine, pipecolic acid, tert ⁇ leucine, norvaline, 2 ⁇ azaphenylalanine, 3 ⁇ azaphenyl ⁇ alanine, 4 ⁇ azaphenyl ⁇ alanine, and 4 ⁇ fluorophenylalanine.
  • coli cells are cultured in the absence of a natural amino acid that is to be replaced (e.g., phenylalanine) and in the presence of the desired non ⁇ naturally occurring amino acid(s) (e.g., 2 ⁇ azaphenylalanine, 3 ⁇ azaphenylalanine, 4 ⁇ azaphenylalanine, or 4 ⁇ fluorophenylalanine).
  • a natural amino acid that is to be replaced e.g., phenylalanine
  • non ⁇ naturally occurring amino acid(s) e.g., 2 ⁇ azaphenylalanine, 3 ⁇ azaphenylalanine, 4 ⁇ azaphenylalanine, or 4 ⁇ fluorophenylalanine.
  • the non ⁇ naturally occurring amino acid is incorporated into the polypeptide in place of its natural counterpart. See, Koide et al., Biochem. 33:7470 ⁇ 6, 1994.
  • Naturally occurring amino acid residues can be converted to non ⁇ naturally occurring species by in vitro chemical modification.
  • Chemical modification can be combined with site ⁇ directed mutagenesis to further expand the range of substitutions (Wynn and Richards, Protein Sci. 2:395 ⁇ 403, 1993).
  • a limited number of non ⁇ conservative amino acids, amino acids that are not encoded by the genetic code, non ⁇ naturally occurring amino acids, and unnatural amino acids may be substituted for amino acid residues of polypeptides of the present invention.
  • Essential amino acids in the polypeptides of the present invention can be identified according to procedures known in the art, such as site ⁇ directed mutagenesis or alanine ⁇ scanning mutagenesis (Cunningham and Wells, Science 244: 1081 ⁇ 5, 1989).
  • Sites of biological interaction can also be determined by physical analysis of structure, as determined by such techniques as nuclear magnetic resonance, crystallography, electron diffraction or photoaffinity labeling, in conjunction with mutation of putative contact site amino acids. See, for example, de Vos et al., Science 255:306 ⁇ 12, 1992; Smith et al., J. Mol. Biol. 224:899 ⁇ 904, 1992; Wlodaver et al., FEBS Lett. 309:59 ⁇ 64, 1992.
  • the identities of essential amino acids can also be inferred from analysis of homologies with related components (e.g. the translocation or protease components) of the polypeptides of the present invention.
  • phage display e.g., Lowman et al., Biochem. 30:10832 ⁇ 7, 1991; Ladner et al., U.S. Patent No. 5,223,409; Huse, WIPO Publication WO 92/06204
  • region ⁇ directed mutagenesis e.g., region ⁇ directed mutagenesis
  • Exon 1 of the SFTPB gene (as described by NG_016967.1) is dash ⁇ underlined (corresponding to bases 5543 ⁇ 5565 of NG_016967.1).
  • the first A of this exon is the starting nucleotide of the mRNA generated by the SFTPB promoter.
  • the bold and italicised ATG (corresponding to bases 5559 ⁇ 5561 of NG_016967.1) encodes the first methionine of pre ⁇ pro ⁇ SFTPB.
  • 3’ of the double ⁇ underlined exon 1 is a partial portion of intron 1 (corresponding to bases 5626 ⁇ 5816 of NG_016967.1).
  • the first base (base 81 of SEQ ID NO: 2) and last base (base 710 of SEQ ID NO: 2) of the core SFTP promoter fragment of the invention are double ⁇ underlined.
  • the wavy ⁇ underlined sequence is the predicted TATA box of the SFTPB promoter.
  • the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 76).
  • SEQ ID NO: 42 VEGFA ⁇ 2 Enhancer & core SFTPB Promoter (1316bp)* gaggcacaaa gcgatcccca tcactgctcc acaatcattc attagctaac aagacagagc 60 agctcataaa aaaaagccg ttaaaaaat tccggggaaaaaagcag gaggtgatgc 120 aagccctggt taacaaggc tgagggttgg ggggaggcat gagagggtgt gagtggaata 180 acccaagcct gataagccac aaagcagcgc ctgctgctcccccccccc attag
  • the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 77).
  • SEQ ID NO: 43 CpG ⁇ Free CMV Enhancer & core SFTPB Promoter (938bp)* gttacataac ttatggtaaa tggcctgcct ggctgactgc ccaatgaccc ctgcccaatg 60 atgtcaataa tgatgtatgt tcccatgtaa tgccaatagg gactttccat tgatgtcaat 120 gggtggagta tttatggtaa ctgcccactt ggcagtacat caagtgtatc atatgccaag 180 tatgcccct attgatgtca atgatggtaa atggcctgcc tggcat
  • the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 78).
  • SEQ ID NO: 44 ELF3 Enhancer & core SFTPB Promoter (1290bp)* tagggatggg ccgaggctgg cactgatgct agacttccgt gcacagggca agtatggaca 60 agccccaagt ggctttgtga ggcccacaca gtgaagcttg ggaaatggga agtggggctg 120 cgcccagatt ctggtatcta tgacaactaa ggccgctgca catcctcatg gctctcccag 180 agacctcagg tgaggccctt ctgtgtgtcct caagcaccca
  • the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 79).
  • SEQ ID NO: 45 SV40 Enhancer & mSPB Promoter (873bp)* cgatggagcg gagaatgggc ggaactgggc ggagttaggg gcgggatggg cggagttagg 60 ggcgggacta tggttgctga ctaattgaga tgcatgctttt gcatacttct gcctgctggg 120 gagcctgggg actttccaca cctggttgct gactaattga gatgcatgct tgcatactt 180 ctgcctgctg gggagcctgg ggactttcca caccctaact gacacacatt ccacagc
  • the invention also expressly encompasses a version of this sequence in which the underlined BglII restriction site is omitted (SEQ ID NO: 80).
  • SEQ ID NO: 46 Alv ⁇ 01 (CMV forward enhancer + mSFPB promoter) (955bp) AGATCTGTTACATAACTTATGGTAAATGGCCTGCCTGGCTGACTGCCCAATGACCCCTGCCCAATGATGTCAAT AATGATGTATGTTCCCATGTAATGCCAATAGGGACTTTCCATTGATGTCAATGGGTGGAGTATTTATGGTAACT GCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTATGCCCCCTATTGATGTCAATGATGGTAAATGGCC TGCCTGGCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTATGTATTAGTCATTGC TATTACCATGGATCTTATAGGAGCCACTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGG GT
  • the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 81).
  • the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 82).
  • SEQ ID NO: 48 Alv ⁇ 03 (VEGFA ⁇ 2 forward enhancer + mSFPB promoter) (1333bp) AGATCTGAGGCACAAAGCGATCCCCATCACTGCTCCACAATCATTCATTAGCTAACAAGACAGAGCAGCTCAT AAAAAAAAAGCCGTTAAAAAAATTCCGGGGAAATGGAAAGCAGGAGGTGATGCAAGCCCTGGTTAACAAAG GCTGAGGGTTGGGGGGAGGCATGAGAGGGTGTGAGTGGAATAACCCAAGCCTGATAAGCCACAAAGCAGC GCCTGCTGCTCCCTCCTCCCTCTGCCGCCTGAGTCAGAAGCCGGGATGTGTTCAAAATCAAGCAATGTAATT AGCAGGTTGCATCATCATGCCTCTCGATTATAAATTACAATGGCCACAAAGAGGGTGGATGGAGGAGCAGGGATT AGATGGGGTAACCCAA
  • the invention also expressly encompasses a version of this sequence in which this Nhe1 restriction site is omitted (SEQ ID NO: 83).
  • SEQ ID NO: 49 ATTTGAGCTCTTCTTTCTGCTGAACCATCG Underlined sequence is an SacI restriction enzyme site
  • SEQ ID NO: 50 SFTPB promoter reverse primer TCTTAGATCTGTCAGACAGCTCTGGGTTCC Underlined sequence is a BgIII restriction enzyme site
  • SEQ ID NO: 51 Exemplary linker between SFTPB promoter fragment and enhancer agatct
  • SEQ ID NO: 52 Exemplified WPRE component (mWPRE) 1 GGGCCCAATC AACCTCTGGA TTACAAAATT TGTGAAAGAT TGACTGGTAT TCTTAACTAT 61 GTTGCTCCTT TTACGCTATG TGGATACGCT GCTTTAATGC CTTTGTATCA TGCTATTGCT 121 TCCCGTATGG CTTTCATTTT CTCCTCCTTG TATAAATCCT
  • Example 1 Design and in vitro validation of a cell ⁇ specific core SFPB (mSPB) promoter sequence
  • mSPB cell ⁇ specific core SFPB
  • SFTPB promoter The expression driven by the core SFTPB promoter was compared to the full length SFTPB promoter (the 972bp fragment of SEQ ID NO: 2, fSPB) using a human surfactant air ⁇ liquid interface (SALI) culture model; an in vitro cell culture model that robustly recapitulates human ATII cells in primary cell culture (Munis et al. (2021) Molecular Therapy: Methods & Clinical Development 20: 237 ⁇ 246). H441 cells, when grown under SALI culture conditions, successfully mimic key characteristics of primary ATII cells. Briefly, SALI cultures were established by culturing cells in 12 ⁇ well Transwell inserts.
  • SALI human surfactant air ⁇ liquid interface
  • the polarization medium comprised either RPMI ⁇ 1640 (H441s cells and co ⁇ culture) or F12 ⁇ K (A549s) supplemented with 2 mM l ⁇ glutamine, 50 U/mL penicillin, 50 mg/mL streptomycin, 1% insulin ⁇ transferrin ⁇ selenium (GIBCO), 4% FCS, and 1 ⁇ M dexamethasone (Sigma). Media were subsequently changed three times per week throughout the described experiments. Cells were grown under SALI conditions for 14 days after air ⁇ lift prior to experimentation unless otherwise stated.
  • the negative control sample (Na ⁇ ve; non ⁇ transduced cells H441 cells) shows a background level of 0.4%.
  • the % EGFP ⁇ positive cells observed with the widely ⁇ used, non ⁇ specific CMV [SEQ ID NO: 6], EF1aS [SEQ ID NO; 7], and PGK [SEQ ID NO: 8] promoters is relatively high (17 ⁇ 22%) in line with expectations.
  • the % EGFP ⁇ positive cells observed for the hCEF promoter [SEQ ID NO: 5], which has been used previously in the lungs of patients, is 11%.
  • mSPB SEQ ID NO: 1
  • fSPB SEQ ID NO: 4
  • mSPC SEQ ID NO: 9
  • fSPC fSPC
  • human HEK293T cells were transduced with recombinant SIV lentiviral vectors pseudotyped with VSV ⁇ G and expressing the EGFP transgene from these same promoters (CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), EF1aS (SEQ ID NO: 7), PGK (SEQ ID NO: 8), the core SFTPB fragment (mSPB) (SEQ ID NO: 1), full ⁇ length SPB (fSPB) (SEQ ID NO: 4), minimal SPC (mSPC) (SEQ ID NO: 9) and full ⁇ length SPC (fSPC) (SEQ ID NO: 10)) at a multiplicity of infection (moi) of 0.5 or 1.0.
  • CMV SEQ ID NO: 6
  • hCEF SEQ ID NO: 5
  • EF1aS SEQ ID NO: 7
  • PGK SEQ ID NO: 8
  • the core SFTPB fragment mSPB
  • fSPB full ⁇ length
  • the % EGFP ⁇ positive cells is low for mSPB, fSPB, mSPC and fSPC promoter sequences in generic HEK293T cells, indicating that the activity of these may be cell ⁇ specific.
  • the mSPB promoter drives very low level expression in HEK293T cells (see, Figure 2).
  • the expression driven by the mSPB promoter is likely to be specific to lung parenchymal cells, particularly ATII cells.
  • Example 2 The core SFPB (mSPB) promoter sequence drives cell ⁇ specific gene expression in vivo
  • the specificity of the mSPB promoter was further investigated using an in vivo mouse model. As shown in Figure 3A, on day 0 mice were dosed by nasal instillation with SIV vector. 7 days after dosing, lung tissue was harvested, fixed, frozen and cryosectioned. The sections were then analysed for expression of EGFP.
  • mSPB core SFTPB promoter fragment
  • mSPB core SFTPB promoter fragment
  • Example 3 Design and production of improved mSPB promoters
  • the inventors then sought to further increase transgene expression by and/or activity of (i.e. the number of lung parenchyma cells in which the promoter is active) the mSPB promoter, whilst retaining specificity for the lung parenchyma.
  • the inventors generated a panel of improved mSPB promoters, each comprising mSPB and an enhancer.
  • Lung ⁇ specific enhancers were selected using the ATII cell gene expression database from LungGENS and a tissue expression database.
  • tissue expression databases including https://research.cchmc.org/pbge/lunggens/default.html and https://tissues.jensenlab.org/Search were interrogated to identify genes with high mean levels of expression in ATII cells.
  • Enhancer regions were selected (and transcription factor binding sites identified) using University of California at Santa Cruz (UCSC) genome browser https://genome.ucsc.edu.
  • the identified enhancer sequences were cloned into a construct comprising mSPB and an operably linked EGFP2ALux transgene.
  • enhancers were added to the constructs in the forward (F) and reverse (R) direction. Schematics of exemplary constructs are shown in Figure 6, with the mSPB promoter sequence of SEQ ID NO: 1 and enhancer sequences as per SEQ ID NOs: 13 to 20, 23 to 33 and 36 to 40.
  • Similar reporter transgene expression constructs were also generated incorporating the commonly used enhancers hB ⁇ Actin (SEQ ID NOs: 21 and 22), SV40 (SEQ ID NOs: 34 and 35) and CMV (SEQ ID NOs: 11 and 12).
  • the expression constructs were incorporated into recombinant lentiviral vectors.
  • the various lentiviral vectors were prepared for analysis of EGFP or Lux reporter transgene expression from each enhancer/promoter combination.
  • Human HEK293T cells, murine LA ⁇ 4 cells and human SALI cells were transfected with the mSPB promoter constructs to determine whether the addition or the enhancer (or the orientation) increased gene expression.
  • mSPB did not significantly increase gene expression relative to the na ⁇ ve control. Differences in expression driven by mSPB compared with mSPB + the different enhancers was also not significant. In contrast, significant differences in gene expression are seen in murine LA ⁇ 4 cells (see, Figure 8) and human SALI cells (see, Figure 9).
  • murine LA ⁇ 4 cells the addition of a CMV enhancer (forwards or reverse), ELF3 enhancer (forwards), SV40 enhancer (forwards), and two of the VEGFA enhancers – VEG1 (forwards or reverse) and VEG2 (forwards) to the mSPB promoter significantly increased expression relative to the mSPB promoter alone.
  • Example 4 Further characterisation of improved mSPB promoters Using the results of Example 3, the following seven candidate enhancers were selected for further screening: SV40, CMV, SLC34A2 ⁇ 1, SLC34A2 ⁇ 2, VEGFA ⁇ 2, ELF3 ⁇ 1 and ELF3 ⁇ 2.
  • hCEF SEQ ID NO: 5
  • CMV SEQ ID NO: 6
  • hCEF SEQ ID NO: 5
  • CMV SEQ ID NO: 6
  • Lentivirus was dosed intranasally (1x10 7 TU per mouse in 100uL TSSM Buffer). On days 2, 7 and 28 post ⁇ dosing, all mice were anaesthetised and imaged for luciferase expression.
  • Figure 12 shows the expression of luciferase as a heat map. Expression driven by hCEF or CMV promoters is shown as a reference. The inventors selected promoter which drive expression in the lungs, but not the nose. Localised expression in the lungs is important, because the nose has cells that are similar to (airway epithelial cells ciliated/non ⁇ ciliated epithelial) to those in the lungs, so it is important to determine that the transgene will not be expressed in the nose.
  • FIG. 13A The expression of luciferase is quantified in Figure 13A.
  • the ratio of expression in the lungs and the nose is shown in Figure 13B.
  • the level of luciferase signal (Figure 13A) was greatest with the CMVenh group (SEQ ID NO: 11) (**** compared to mSPB (SEQ ID NO: 1) although the specificity for lung expression (Figure 13B), determined by the ratio of signal in the lung and nose (L ⁇ to ⁇ N), was not significantly different from CMV (SEQ ID NO: 6)or hCEF (SEQ ID NO: 5) at day 28.
  • transgene expression from these constructs delivered as a non ⁇ viral (plasmid) formulation were similar to those obtained following delivery with viral (recombinant lentiviral) vectors.
  • the enhancer sequences CMVenh (SEQ ID NO: 11), Slc2 (SEQ ID NO: 31) and VEGFA ⁇ 2 (SEQ ID NO: 38) combined with the mSPB (SEQ ID NO: 1) promoter sequence were assigned the nomenclature Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 respectively, as shown in SEQ ID NOs: 46, 47 and 48 respectively.
  • Mouse lungs were processed for cryosections and imaging. Cryosections (7 ⁇ M) were subject to immunohistochemistry using primary antibodies to detect colocalization of EGFP and Pro/Mature Surfactant Protein ⁇ B, the latter being a marker for ATII cells.
  • this experiment confirmed that the mSPB promoter/enhancer combinations tested drove transgene expression primarily in cells of the parenchyma, making them attractive candidates for targeted gene therapy of diseases resulting from or associated with deficiency of expression and/or expression of defective proteins in lung parenchymal cells, such as surfactant deficiencies.
  • Repetition of this experiment using a different ATII cell marker, Surfactant Protein ⁇ C yielded the same pattern of expression, with Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 driving EGFP expression which colocalises with the ATII cell ⁇ specific marker SP ⁇ C (data not shown). Therefore, Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 have been reproducibly shown to drive expression in the lung parenchyma.
  • Example 6 Improved mSPB promoters drive long ⁇ term expression in vivo
  • Lentiviral vectors comprising mSPB, Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 were administered to intranasally. The mice were then monitored for luciferase expression.
  • FIG. 16 shows the expression of luciferase as a heat map. Expression driven by hCEF or CMV promoters is shown as a reference. The expression of luciferase in the lungs over the time course of the experiment is quantified in Figure 16B. The area under the curve is shown in Figure 16C and the ratio of expression in the lungs and the nose is shown in Figure 16D.
  • the results indicate that the Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 promoter drives high levels of expression in the lungs over at least a six ⁇ month period. Further, mSPB, Alv ⁇ 02 and Alv ⁇ 03 drive high levels of expression in the lungs, but not the nose.
  • Example 7 – Improved mSPB promoters drive targeted expression in cells of the lung parenchyma Further analysis was conducted to confirm that Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 result in transgene expression in the target cells, particularly that expression is focussed in the cells of the lung parenchyma, rather than airway cells. This was investigated using immunohistochemistry.
  • the hCEF promoter mainly drove expression in the cells of the airway, whereas as shown in Figure 17B, the mSPB promoter drove expression primarily in the parenchymal cells.
  • the % of EGFP expression in the parenchymal cells for hCEF, mSPB, Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03 was quantified.
  • Alv ⁇ 01, Alv ⁇ 02 and Alv ⁇ 03, particularly Alv ⁇ 02 and Alv ⁇ 03 were able to drive EGFP expression in the parenchyma compared with the hCEF promoter.
  • Example 8 Expression of SP ⁇ B restores transepithelial electrical resistance (TEER) in SFTPB knock out lung cells
  • TEER transepithelial electrical resistance
  • a Surfactant Air Liquid Interface (SALI) model was used with the human H441 lung cell line and the H441 SP ⁇ B KO cell line where the SFTPB gene had been ablated (Munis et al. (2021) Mo. Ther. Methods Clin. Dev 20:20:237 ⁇ 246, herein incorporated by reference).
  • TEER transepithelial electrical resistance
  • Example 9 Pulmonary surfactant has no effect on HEK293/T cell transduction with rSIV.F/HN vectors
  • synthetic surfactant such as BLES or Curosurf
  • rSIV.F/HN vector encoding EGFP under CMV promoter control was mixed 1:1 with TSSM (vehicle control) or with BLES or Curosurf and incubated at room temperature for 30mins. The mixtures were then diluted in OptiMEM ⁇ I (supplemented with polybrene for final 8 ⁇ g/mL working concentration) and used to transduce HEK293T cells.
  • Example 10 Murine lung transduction with rSIV.F/HN vectors is at least as effective in the presence of a pulmonary surfactant Having surprisingly shown that pulmonary surfactants do not compromise the integrity of rSIV.F/HN lentiviral vectors in an in vitro setting, the effect of pulmonary surfactants on rSIV.FHN transduction of the murine lung was then investigated.
  • rSIV.F/HN vector expressing firefly luciferase from the hCEF promoter was mixed 1:1 with TSSM (vehicle control) or BLES or Curosurf (to generate 2E8 TU in 100uL per dose).
  • TSSM vehicle control
  • BLES Curosurf
  • TSSM diluent served as vehicle control.
  • Mice were subjected to in vivo bioluminescent imaging to measure Firefly luciferase expression in the murine lungs.

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Biotechnology (AREA)
  • Biomedical Technology (AREA)
  • General Engineering & Computer Science (AREA)
  • Organic Chemistry (AREA)
  • Molecular Biology (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • General Health & Medical Sciences (AREA)
  • Biochemistry (AREA)
  • Public Health (AREA)
  • Physics & Mathematics (AREA)
  • Biophysics (AREA)
  • Veterinary Medicine (AREA)
  • Plant Pathology (AREA)
  • Animal Behavior & Ethology (AREA)
  • Microbiology (AREA)
  • Epidemiology (AREA)
  • Pharmacology & Pharmacy (AREA)
  • Medicinal Chemistry (AREA)
  • Medicines That Contain Protein Lipid Enzymes And Other Medicines (AREA)
  • Medicines Containing Material From Animals Or Micro-Organisms (AREA)
  • Micro-Organisms Or Cultivation Processes Thereof (AREA)
  • Medicinal Preparation (AREA)
  • Pharmaceuticals Containing Other Organic And Inorganic Compounds (AREA)

Abstract

The present invention relates to nucleic acid cassettes for gene therapy, particularly to promoter and promoter/enhancer combinations for improved expression of transgenes in a lung parenchyma‐specific/preferred manner. The invention further relates nucleic acid cassettes comprising said promoters and promoter/enhancer combinations, viral and non‐viral vectors comprising such nucleic acid cassettes, and the use of such nucleic acid cassettes and vectors to increase expression of therapeutic proteins by lung parenchyma cells.

Description

SYNTHETIC PROMOTERS    FIELD OF THE INVENTION  The  present  invention  relates  to  nucleic  acid  cassettes  for  gene  therapy,  particularly  to  promoter  and promoter/enhancer  combinations  for  improved  expression of  transgenes  in  a  lung  parenchyma‐specific/preferred  manner.  The  invention  further  relates  nucleic  acid  cassettes  comprising  said  promoters  and  promoter/enhancer  combinations,  viral  and  non‐viral  vectors  comprising  such nucleic  acid  cassettes,  and  the use of  such nucleic  acid  cassettes  and  vectors  to  increase expression of therapeutic proteins by lung parenchyma cells.    BACKGROUND TO THE INVENTION  Surfactant protein B (SFTPB) deficiency  is a severe monogenic  interstitial  lung disorder that  leads to loss of life in infants as a result of alveolar collapse and respiratory distress syndrome. The  only curative treatment  is thought  to be  lung transplantation; however, the  lack of suitable donor  organs makes this a non‐viable option in most circumstances.  The use of nucleic acids as medicine, or gene therapy, is a promising new treatment modality,  both  for SFTPB deficiency and other genetic diseases,  including genetic diseases of the respiratory  tract. The reason many gene  therapies currently  in use or under development are not effective at  curing diseases is because it is difficult to make sufficient protein to reach the therapeutic threshold  needed to treat or cure the disease. As such, generating sufficient gene expression in a given target  cell is a major barrier to the success of many gene therapies.  One approach  to  reaching  the  large doses needed  for gene  therapy  to be  successful  is  to  administer massive amounts of the gene therapy to the patient, over 1 trillion viruses per kg of body  mass.  Producing  so much  virus  is  expensive,  contributing  to  the  $  1,000,000  USD  cost  of  gene  therapies, and giving so much virus to a person can trigger immune responses that threaten the health  of the patient and the efficacy of the therapy. To circumvent these problems, research to‐date has  focused on gain of function mutations resulting in more potent proteins. Such an approach has been  used  previously  in  the  gene  therapies  for  haemophilia  B  (the  Padua mutation  in  Factor  IX)  and  lipoprotein lipase deficiency (the S447X variant of lipoprotein lipase). However, such gain‐of‐function  mutations are not available  for most gene  therapies. Furthermore, and even with gain‐of‐function  mutations, high doses of the gene therapy vector were still necessary to make the treatment effective.  Therefore,  such gain‐of‐function mutations along do not adequately address  the exiting problems  associated with producing sufficient quantities of vector, or the unwanted and clinically dangerous  side effects associated with the large doses required.   The  present  inventors  have  previously  developed  a  lentiviral  vector,  which  has  been  pseudotyped with  hemagglutinin‐neuraminidase  (HN)  and  fusion  (F)  proteins  from  a  respiratory  paramyxovirus, comprising a promoter and a transgene. Typically, the backbone of the vector is from  a simian immunodeficiency virus (SIV), such as SIV1 or African green monkey SIV (SIV‐AGM). Preferably  the backbone of a viral vector of  the  invention  is  from SIV‐AGM. The HN and F proteins  function,  respectively, to attach to sialic acids and mediate cell fusion for vector entry to target cells. The present  inventors discovered that this specifically F/HN‐pseudotyped lentiviral vector can efficiently transduce  airway  epithelium,  resulting  in  transgene  expression  sustained  for  periods  beyond  the  proposed  lifespan of airway epithelial cells. Importantly, the present inventors also found that re‐administration  does not result in a loss of efficacy. These features make the vectors of the present invention attractive  candidates for treating diseases via their use in expressing therapeutic proteins: (i) within the cells of  the respiratory tract; (ii) secreted into the lumen of the respiratory tract; and (iii) secreted into the  circulatory  system.  However,  even  using  this  state‐of‐the‐art  platform  technology,  the  levels  of  transgene expressed are at the lower predicted threshold required for clinical efficacy.  There are other approaches which can also potentially increase the expression of therapeutic  proteins by gene therapy vectors.  For example, the use of exogenous signal peptides can be used to  increase expression and secretion of  therapeutic proteins by airway cells. By using  the exogenous  signal peptides  it  is possible  to produce more protein  for every copy of a gene  therapy vector or  transgene  that  is put  into a cell,  increasing  the dose of  therapeutic protein without  increasing  the  amount of gene therapy vector given to a patient. However, the use of exogenous signal peptides is  not appropriate for all therapeutic proteins or all conditions.  For example, not all therapeutic proteins  are  secreted, and  for  some  there may be clinical  reasons why manipulating  the  signal peptides  is  undesirable.  Generally,  gene  therapy  vectors  comprising promoters providing high  and  sustained  gene  expression in a variety of cell types are preferred, especially in a therapeutic context. For this reason,  the  inventors previously used a hCEF promoter to drive strong and persistent expression  in mouse  lung with non‐viral formulations. However for some conditions, expression in a specific tissue or cell  type is required to (i) achieve desired therapeutic target, (ii) avoid gene expression‐related toxicities,  and (iii) circumvent immune responses to the therapeutic agent stemming from gene expression in  undesired cell  types.  In particular,  it would be desirable  to express surfactants  in a  tissue‐ or cell‐ specific manner,  and with  similar  levels  of  gene  expression  to  that  of  ubiquitously  used  strong  promoters,  for the treatment of genetic diseases, particularly genetic respiratory diseases, such as  SFTPB deficiency.   There is therefore an unmet clinical need for new technologies to improve the tissue‐specific  and/or  cell‐specific  expression  of  transgenes  from  a  gene  therapy  vector.  It  is  an  object  of  the  invention to address one or more of these problems. In particular, it is an object of the invention to  provide new promoters, nucleic acid cassettes and gene therapy vectors which enable high expression  of transgenes in the lung parenchyma. Such promoters, cassettes and vectors may be of particular use  in  the  treatment  of  genetic  diseases,  particularly  genetic  respiratory  diseases,  such  as  surfactant  protein deficiency.    SUMMARY OF THE INVENTION  At present, there remains a pressing need for technology that enables high levels of cell‐ or  tissue‐specific expression of transgenes for gene therapy, including from the inventors’ own lentiviral  platform. In particular, there is a need for novel promoters that drive cell‐specific expression in the  lung parenchyma.   The present  inventors have developed a panel of novel promoter  sequences comprising a  functional fragment of the human SFTPB gene promoter. As demonstrated herein, the inventors have  surprisingly shown  that  the SFTPB promoter  fragment of  the  invention drives  increased  transgene  expression  compared with  the  full‐length SFTPB promoter.   Furthermore,  the  inventors have also  surprisingly demonstrated that the SFTPB promoter fragment of the invention can be combined with  particular enhancers, including some enhancers which are not associated with cell‐specific expression  in the lung parenchyma, to further improve transgene expression.   Accordingly, the present invention provides an SFTPB promoter fragment which comprises or  consists of at  least part of exon 1 of the SFTPB gene, but does not comprise  the SFTPB gene start  codon. Typically  said promoter  is  less  than 800 bases  in  length, preferably  less  than 700 bases  in  length. Said promoter may comprise or consist of: (a) SEQ ID NO: 1  or a sequence with at least 80%   identity to SEQ ID NO: 1; or (b) bases 81‐710 of SEQ ID NO: 2, or a sequence with at least 80% identity  to  bases  81‐710  of  SEQ  ID  NO:  2; wherein  optionally  (i)  said  SFTPB  promoter  fragment  further  comprises up to 20 bases at the 5’ end, which may optionally correspond to up to 20 bases 5’ to base  81 of SEQ ID NO: 2; and/or (ii) said SFTPB promoter fragment further comprises up to 4 bases at the  3’ end, which may optionally correspond to up to 4 bases 3’ to base 710 of SEQ ID NO: 2.  The SFTPB  promoter fragment may comprise or consist of SEQ ID NO: 1, or a sequence with at least 90% identity  to SEQ ID NO: 1.  The SFTPB promoter  fragment of  the  invention may  further  comprise a 5’ enhancer.  Said  enhancer may be in (i) the forward, or (ii) the reverse, orientation; preferably wherein the enhancer  is  in  the  forward orientation.   Said enhancer may be selected  from a SLC34A2 enhancer, a VEGFA  enhancer,  a  CMV  enhancer,  an  SV40  enhancer,  an  ELF3  enhancer,  an  actin  enhancer,  an  LMO7  enhancer, a SFTPC enhancer or a SFTPB enhancer. Preferably the enhancer may be selected from a  SLC34A2 enhancer, a VEGFA enhancer or a CMV enhancer.  The enhancer may be selected from (a) an  SLC34A2 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to  any one of SEQ ID NOs: 29‐33, preferably SEQ ID NO: 31; (b) a VEGFA enhancer which comprises or  consists  of  a  nucleotide  sequence with  at  least  90%  identity  to  any  one  of  SEQ  ID  NOs:  36‐40,  preferably SEQ ID NO: 38; (c) a CMV enhancer which comprises or consists of a nucleotide sequence  with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11; (d) a SV40 enhancer which  comprises or consists of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or 35; (e)  an ELF3 enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to  any one of SEQ  ID NOs: 13‐20;  (f) an actin enhancer which  comprises or  consists of a nucleotide  sequence with at least 90% identity to SEQ ID NO: 21 or 22; (g) an LMO7 enhancer which comprises  or consists of a nucleotide sequence with at least 90% identity to any one of SEQ ID NOs: 23—25; (h)  an SFTPC enhancer which comprises or consists of a nucleotide sequence with at least 90% identity to  SEQ ID NO:  27 or 28; or (i) an SFTPB enhancer which comprises or consists of a nucleotide sequence  with at least 90% identity to SEQ ID NO: 26.  The SFTPB promoter  fragment of  the  invention may  comprise or  consist of a nucleic acid  sequence of any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 49, or a sequence with at least 90%  identity to any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 48.  Preferably the SFTPB promoter  fragment of the invention comprises or consists of a nucleic acid sequence of any one of SEQ ID NOs:  46, 47 or 48, or  a sequence with at least 90% identity to any one of SEQ ID NOs: 46, 47 or 48.  The  invention  also  provides  a  nucleic  acid  cassette  comprising:  (a)  an  SFTPB  promoter  fragment of the  invention; and (b) a transgene.   Said transgene may encode a   therapeutic protein  selected from: (a) secreted therapeutic protein selected from: Surfactant Protein B (SP‐B), Surfactant  Protein C  (SP‐C), AAT, Factor VIII, Factor VII, Factor  IX, Factor X, Factor XI, van Willebrand Factor,  Granulocyte‐Macrophage Colony‐Stimulating Factor (GM‐CSF), decorin, an anti‐inflammatory protein  (e.g. IL‐10 or TGFβ) or monoclonal antibody, an anti‐inflammatory decoy, or a monoclonal antibody  against an  infectious agent; or  (b) ATP‐binding cassette  sub‐family A member 3  (ABCA3), TRIM72,  CSF2RA, or CSF2RB.  The SFTPB promoter fragment may increase expression of the transgene by lung parenchymal  cells, optionally compared with the full‐length SFTPB promoter. Expression of the transgene  may be  increased by at least 2‐fold, preferably at least 5‐fold compared with the full‐length SFTPB promoter.   The lung parenchymal cells may comprise one or more cell type selected from: alveolar type I epithelial  (ATI) cells, alveolar  type  II epithelial cells  (ATII), and/or club cells, preferably ATII and/or ATI cells.   Expression of the transgene by a nucleic acid cassette or promoter of the invention may be specific to  lung parenchymal cells; and/or the ratio of lung expression: nose expression of the transgene by the  SFTPB promoter fragment is at least 2:1.   The  invention further provides a gene therapy vector, comprising a nucleic acid cassette of  the invention.  Said gene therapy vector may be a non‐viral vector, wherein optionally: (a) the non‐ viral vector  is a plasmid; and/or (b) the non‐viral vector  is comprised  in a cationic  liposome, which  preferably comprises GL67A.  Said gene therapy vector may be a viral vector, optionally selected from:  a lentiviral vector; an AAV vector; and an adenoviral vector.  Said lentiviral vector may be pseudotyped  with haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus,  optionally from a Sendai virus. Said lentiviral vector may be selected from the group consisting of a  Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a Feline  immunodeficiency  virus  (FIV)  vector,  an  Equine  infectious  anaemia  virus  (EIAV)  vector,  and  a  Visna/maedi virus vector. Preferably said lentiviral vector is a SIV vector.  The invention further provides a method of expressing a therapeutic protein in a target cell,  comprising delivering a nucleic acid cassette of the invention or a gene therapy vector of the invention  into  the  target  cells.  Said  delivering may  comprise  integrating  said  nucleic  acid  cassette  or  gene  therapy vector into said target cell's genome.   The  invention also provides a gene therapy vector of the  invention  for use  in a method of  treating a disease.  Said disease may be a genetic disease.   The disease may be:  (a) a  respiratory  disease, particularly a genetic respiratory disease; or (b) a cardiovascular disease or blood disorder,  particularly a genetic  cardiovascular disease or blood disorder. The disease may be  selected  from  Surfactant  Protein  B  (SP‐B)  Deficiency;  Surfactant  Protein  C  (SP‐C)  deficiency;  ABCA3  deficiency;  Pulmonary  surfactant  metabolism  dysfunction  2  (SMDP2);  Pulmonary  surfactant  metabolism  dysfunction  3  (SMDP3);  Primary  Ciliary  Dyskinesia  (PCD);  Alpha  1‐antitrypsin  Deficiency  (A1AD);  Pulmonary  Alveolar  Proteinosis  (PAP);  Chronic  obstructive  pulmonary  disease  (COPD);  Acute  respiratory distress syndrome (ARDS); COVID‐19; a pulmonary fibrotic disease; a pulmonary allergic  condition;  a  pulmonary  bacterial  infection;  lung  cancer;  a  dysplastic  change  in  the  lungs;  and  haemophilia.  The invention further provides a cell comprising a nucleic acid cassette of the invention or a  gene therapy vector of the invention.  The invention also provides a composition comprising a nucleic acid cassette of the invention  or  a  gene  therapy  vector  of  the  invention  and  a  pharmaceutically  acceptable  carrier,  diluent  or  excipient.  In addition,  the  inventors have shown  for  the  first  time  that  lentiviral vectors, such as  the  lentiviral vector pseudotyped with haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a  respiratory paramyxovirus as exemplified herein, can be co‐administered with a synthetic surfactant,  and  achieve  at  least  as  efficient  transduction  into  target  cells,  and  potentially  even  enhanced  transduction, compared with transduction of the lentiviral vector in vehicle alone.  Whilst it is known  in the art that AAV vectors can be administered with surfactant,  it  is surprising that this  is possible  with lentiviral vectors due to fundamental structural differences between AAV and lentiviral vectors.   In particular,  lentiviral vectors are surrounded by a  lipid envelope, whereas AAV are not.    It would  therefore be expected  that  the  lentiviral vector envelope would be disrupted by  the hydrophobic  portions of a surfactant, having a negative effect on the lentiviral structure.  Accordingly, the invention also provides lentiviral vector pseudotyped with haemagglutinin‐ neuraminidase  (HN)  and  fusion  (F)  proteins  from  a  respiratory  paramyxovirus  and  comprising  a  transgene operably  linked  to  a promoter  for use  in  a method of  treating  a  disease, wherein  the  lentiviral vector is administered simultaneously or sequentially with a surfactant.  Optionally said disease is: (a) a genetic disease; (b) a respiratory disease, particularly a genetic  respiratory disease; (c) a cardiovascular disease or blood disorder, particularly a genetic cardiovascular  disease or blood disorder; and/or (d) selected from Surfactant Protein B (SP‐B) Deficiency; Surfactant  Protein  C  (SP‐C)  deficiency;  ABCA3  deficiency;  Pulmonary  surfactant  metabolism  dysfunction  2  (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD);  Alpha  1‐antitrypsin Deficiency  (A1AD);  Pulmonary  Alveolar  Proteinosis  (PAP);  Chronic  obstructive  pulmonary  disease  (COPD);  Acute  respiratory  distress  syndrome  (ARDS);  COVID‐19;  a  pulmonary  fibrotic  disease;  a  pulmonary  allergic  condition;  a  pulmonary  bacterial  infection;  lung  cancer;  a  dysplastic change in the lungs; and haemophilia.  In such a  lentiviral vector, the respiratory paramyxovirus may be a Sendai virus; and/or the  lentiviral vector may be selected from the group consisting of a Human immunodeficiency virus (HIV)  vector, a Simian immunodeficiency virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an  Equine infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector, preferably a SIV vector  The transgene may encode a therapeutic protein selected  from:  (a) a secreted therapeutic  protein selected from: Surfactant Protein B (SP‐B), Surfactant Protein C (SP‐C), AAT, Factor VIII, Factor  VII, Factor IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte‐Macrophage Colony‐Stimulating  Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody,  an anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP‐binding  cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.  In such a lentiviral vector, the promoter may comprise a SFTPB promoter fragment as defined  herein; and/or  the lentiviral vector may comprise a nucleic acid cassette as defined herein.  The  lentiviral  vector may  be  administered  before  the  surfactant.    The  surfactant may  be  administered before the lentiviral vector.  The lentiviral vector and surfactant may be administered  simultaneously,  optionally  wherein  the  lentiviral  vector  and  surfactant  are  mixed  prior  to  administration.    BRIEF DESCRIPTION OF THE DRAWINGS    Figure 1: Flow cytometry plots illustrating EGFP expression levels in human SALI cultures transduced   with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing the EGFP transgene at  a dose of 1x10Transducing Units (TU) from a range of promoters (CMV (SEQ ID NO: 6), hCEF ( (SEQ  ID NO: 5), EF1aS ( (SEQ ID NO: 7), PGK ( (SEQ ID NO: 8), mSPB ( (SEQ ID NO: 1), fSPB (SEQ ID NO: 4),  mSPC (SEQ ID NO: 9), fSPC (SEQ ID NO: 10)).     Figure 2: Graphs quantifying EGFP expression in human HEK293T cells transduced with recombinant  SIV  lentiviral vectors pseudotyped with VSV‐G and expressing  the EGFP  transgene  from a range of  promoters (CMV (SEQ ID NO: 6), hCEF ( (SEQ ID NO: 5), EF1aS ( (SEQ ID NO: 7), PGK ( (SEQ ID NO: 8),  mSPB ( (SEQ ID NO: 1), fSPB (SEQ ID NO: 4), mSPC (SEQ ID NO: 9), fSPC(SEQ ID NO: 10)) at a multiplicity  of infection (moi) of 0.5 or 1.0. The % EGFP‐positive cells and the mean fluorescence intensity (MFI)  at both moi are plotted. The negative control sample  (Mock; non‐transduced, treated with buffer‐ only) represents the background levels.     Figure 3: Experimental  schematics  for experiments  to  test EGFP expression  in vivo using  lentiviral  vectors expressing EGFP from different promoters. A Schematic showing timing of dosing, on Day 0  mice were dosed with  lentiviral vectors via nasal  instillation and 7 days post‐dosing the mice were  culled and  lung  tissue harvested  for cryosections. B Schematic  showing  treatment groups. Female  BALB/c mice (4‐6 weeks; n=3 per group) were dosed on Day 0 with recombinant SIV lentiviral vectors  pseudotyped  with  F/HN  and  expressing  the  EGFP  transgene  (1x106  TU  in  100µL  TSSM  Buffer)  expressing EGFP  from either  the mSPB  (SEQ  ID NO: 1) or  fSPB  (SEQ  ID NO: 4) promoter sequence  (mSPB and  fSPB groups,  respectively). Female BALB/c mice  (n=2) were  treated with 100 µL TSSM  buffer only (naïve group) as a negative control.    Figure 4: Representative  immunohistochemistry  images showing expression of EGFP  in  lung  tissue  samples taken from mice treated with  lentiviral vectors expressing EGFP from different promoters.  Images of native EFGP fluorescence (n=6 per mouse) were analysed in lung cryosections of mice dosed  as shown in Figure 3. EGFP fluorescence in lung sections was analysed for the mSPB (SEQ ID NO: 1)   and fSPB (SEQ ID NO: 4) groups and compared with mice in the naïve group imaged in parallel. The  observed  fluorescence appeared  ‘punctate’ and scattered through the  lung parenchyma  (Arrows =  foci corresponding to EGFP expression), which is indicative of expression in ATII cells. Very little EGFP  fluorescence was visible in the airway epithelia with these promoters, which contrasts with previously  observed EGFP expression from promiscuous promoters.    Figure 5: Panel of ATII specific genes (and enhancers) generated by interrogation of LungGENS and a  tissue expression database.     Figure 6: Schematics of exemplary mSPB/enhancer constructs. Different  lengths  (indicated  in base  pairs (bp)) of the newly identified enhancer sequences (boxes) were sub‐cloned in front of the mSPB   (SEQ ID NO: 1) promoter sequence (arrow boxes) in both forward (f) and reverse (r) orientations to  generate  candidate  synthetic  promoters  expressing  the  EGFP2ALux  reporter  transgene.  Similar  reporter  transgene  expression  constructs were  also  generated  incorporating  the  commonly  used  enhancers hB‐Actin, SV40 and CMV. Some f and r permutations were not constructed.     Figure 7: Graph showing expression of EGFP in HEK293T cells by different mSPB/enhancer constructs.  Recombinant  SIV  lentiviral  vectors pseudotyped with VSV‐G and expressing EGFP2ALux  transgene  from each of the enhancer/promoter combinations, were produced at small scale. These were used  to transduce human HEK293T cells in a 24‐well plate (seeded @1x105 cells/well) at a multiplicity of  infection (moi) of approximately 0.1. Luciferase levels (Relative Light Units (RLU)) were determined in  n=4  independent  biological  replicates.  Naïve  (non‐transduced)  cells  were  used  as  a  control  to  determine background RLU level. There were no statistically significant differences observed between  naïve and mSPB (SEQ ID NO: 1), or between mSPB (SEQ ID NO: 1) and all other enhancer constructs  (Kruskal‐Wallis with Dunn’s multiple comparison test).    Figure  8:  Graph  showing  expression  of  EGFP  in  murine  LA‐4  cells  by  different  mSPB/enhancer  constructs. P values are given where significant expression was observed. Recombinant SIV lentiviral  vectors  pseudotyped  with  VSV‐G  and  expressing  EGFP2ALux  transgene  from  each  of  the  enhancer/promoter  combinations, were  produced  at  small  scale.  These were  used  to  transduce  murine  LA‐4  cells  in  a  24‐well  plate  (1x105  cells/well)  at  a  multiplicity  of  infection  (moi)  of  approximately 0.1. Luciferase levels (Relative Light Units (RLU)) were determined in n=4 independent  biological replicates. Naïve (non‐transduced) cells were used as a control to determine background  RLU level. P values are given where statistically significant differences were observed (Kruskal‐Wallis  with Dunn’s multiple comparison test).    Figure  9:  Graph  showing  expression  of  EGFP  in  human  SALI  cells  by  different  mSPB/enhancer  constructs. P values are given where significant expression was observed. Recombinant SIV lentiviral  vectors  pseudotyped  with  VSV‐G  and  expressing  EGFP2ALux  transgene  from  each  of  the  enhancer/promoter  combinations, were  produced  at  small  scale.  These were  used  to  transduce  human  SALI  cultures  in  (1x105  Transducing  Units/well)  at  a  multiplicity  of  infection  (moi)  of  approximately 0.1. Luciferase levels (Relative Light Units (RLU)) were determined in n=4 independent  biological replicates. Naïve (non‐transduced) cells were used as a control to determine background  RLU level. P values are given where statistically significant differences were observed (Kruskal‐Wallis  with Dunn’s multiple comparison test).     Figure 10: Graph showing expression of EGFP in human SALI cells by lentiviral vectors with different  mSPB/enhancers in a secondary screen. P values are given where significant expression was observed.  Recombinant  SIV  lentiviral  vectors  pseudotyped  with  F/HN  and  expressing  EGFP2ALux  reporter  transgene,  were  used  to  transduce  human  SALI  cultures  (n=6  replicates)  in  a  repeat  secondary  screening experiment, which included the following promoter sequences: CMV (SEQ ID NO: 6), hCEF  (SEQ ID NO: 5), and mSPB ( (SEQ ID NO: 1), as well as with selected enhancer/promoter combinations  SV40 (f)[SEQ ID NO: 34], CMVenh (f) [SEQ ID NO: 11], Elf1 (f)[SEQ ID NO: 13], Elf2 (f)[SEQ ID NO: 15],  Slc1 (f)[SEQ ID NO: 29], Slc2 (f)[SEQ ID NO: 31], and Veg2 (f)[SEQ ID NO: 38]. Luciferase expression  levels (Relative Light Units (RLU)) were determined and compared with naïve (non‐transduced) control  cells. Vector Copy Number (VCN) was also determined in the transduced cell samples (n=3). Results  from six independent biological replicates are shown; P values are given where statistically significant  differences were observed (Kruskal‐Wallis with Dunn’s multiple comparison test).     Figure 11: Graph  showing expression of EGFP  in HEK293T cells by  lentiviral vectors with different  mSPB/enhancers in a secondary screen. P values are given where significant expression was observed.  Recombinant SIV  lentiviral vectors pseudotyped with the F/HN and expressing EGFP2ALux reporter  transgene, were used to transduce human HEK293T cell cultures (n=6 replicates) in a repeat secondary  screening experiment, which  included  the  following  selected enhancer/promoter  sequences: CMV  (SEQ  ID  NO:  6),  hCEF  (SEQ  ID  NO:  5),  and  mSPB  (  (SEQ  ID  NO:  1),  as  well  as  with  selected  enhancer/promoter combinations SV40 (f)[SEQ ID NO: 34], CMVenh (f) [SEQ ID NO: 11], Elf1 (f)[SEQ  ID NO: 13], Elf2 (f)[SEQ ID NO: 15], Slc1 (f)[SEQ ID NO: 29], Slc2 (f)[SEQ ID NO: 31], and Veg2 (f)[SEQ ID  NO: 38]. Luciferase expression levels (Relative Light Units (RLU)) were determined and compared with  naïve  (non‐transduced)  control  cells.  Vector  Copy  Number  (VCN)  was  also  determined  in  the  transduced  cell  samples  (n=3). Results  from  six  independent  biological  replicates  are  shown;  n.s:  Results were not statistically significantly different (Kruskal‐Wallis with Dunn’s multiple comparison  test).     Figure 12: Heat map of  luciferase expression  in mice  treated with  lentiviral vectors with different  mSPB/enhancers. Female BALB/c mice (6 weeks; n=5 per group) were dosed with recombinant SIV  lentiviral vectors pseudotyped with  F/HN and expressing EGFP2ALux  reporter  transgene  from  the  following promoter sequences: CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), and mSPB ( (SEQ ID NO: 1),  as well as with selected enhancer/promoter combinations SV40 (f)[SEQ ID NO: 34], CMVenh (f) [SEQ  ID NO: 11], Elf1 (f)[SEQ ID NO: 13], Elf2 (f)[SEQ ID NO: 15], Slc1 (f)[SEQ ID NO: 29], Slc2 (f)[SEQ ID NO:  31], and Veg2 (f)[SEQ ID NO: 38]. Lentivirus was dosed intranasally (1x107 TU per mouse in 100uL TSSM  Buffer). On  days  2,  7  and  28  post‐dosing,  all mice were  anaesthetised  and  imaged  for  luciferase  expression.  Observed  areas  of  luciferase  positive  signal  correspond  with  luciferase  transgene  expression in the nose and/or lungs. A naïve group (n=3) of (non‐transduced) animals were included  as  a  negative  control  and  showed  no  luciferase  signal  (background  levels).  Luciferase  signal was  observed in the nose and lung areas at all timepoints with the non‐specific CMV[SEQ ID NO: 6] and  hCEF[ SEQ ID NO: 5] lung promoter in line with expectations. Luciferase signal in the mSPB[SEQ IDN O:  1] group was overall lower and restricted to the lung, also as expected.     Figure  13:  Graphs  showing  in  vivo  luciferase  expression  by  lentiviral  vectors  with  different  mSPB/enhancers (A) and the ratio of luciferase expression in the lungs and the nose (B). As described  in Figure 12, female BALB/c mice (6 weeks; n=5 per group) were dosed with recombinant SIV lentiviral  vectors pseudotyped with F/HN and expressing EGFP2ALux  reporter  transgene  from  the  following  enhancer/promoter sequences: CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), and mSPB ( (SEQ ID NO: 1),  as well as the selected enhancer/promoter combinations SV40 (f)[SEQ ID NO: 34], CMVenh (f) [SEQ ID  NO: 11], Elf1 (f)[SEQ ID NO: 13], Elf2 (f)[SEQ ID NO: 15], Slc1 (f)[SEQ ID NO: 29], Slc2 (f)[SEQ ID NO: 31],  and Veg2 (f)[SEQ ID NO: 38]. Lentivirus was dosed intranasally (1x107 TU per mouse in 100uL TSSM  Buffer).  On  days  2,  7  and  28  post‐dosing,  mice  were  anaesthetised  and  imaged  for  luciferase  expression. A) The day 28  luciferase signal (Radiance (photons/s/cm2/sr) and B) the ratio of in vivo  day  28  luciferase  signal  in  the  mouse  lungs  and  nose,  are  plotted  for  each  enhancer/mSPB  combination tested.      Figure  14:  Graphs  showing  luciferase  expression  by  plasmids with  different mSPB/enhancers  in  human SALI cells (A), in vivo (B), and in HEK293T cells (C). The enhancer/mSPB promoter constructs  expressing Lux reporter transgene were evaluated in the context of a non‐viral formulation, to deliver  plasmids containing the CMV[SEQ ID NO: 6], hCEF[SEQ ID NO: 5] and mSPB[SEQ ID NO: 1] promoter  sequences as well as the selected enhancer/mSPB promoter constructs: slc2[SEQ ID NO: 31], veg2[SEQ  ID NO: 38], sv40[SEQ ID NO: 34], and CMVenh[SEQ ID NO: 11]. (A) Human SALI cultures (n=8) were  transfected with plasmid DNA  (2ug plasmid per  culture) complexed with  linear polyethyleneimine  (PEIPro; 100µL per culture) and  luciferase activity  in Relative Light Units (RLU) determined. A naïve  group (n=8; non‐transfected) was included as a negative control. Luciferase activity in the veg2[SEQ  ID NO: 38] and CMVenh[SEQ ID NO: 11] groups were significantly different from mSPB [SEQ ID NO: 1]  (** and  ****, respectively; Kruskal‐Wallis). NB: a CMV[SEQ ID NO: 6] group was not included in this  experiment. (B) Female BalbC mice (n=6 per group) were each dosed intranasally (100µL per mouse)  with 60µg of plasmid complexed with 77.4µg 25KDa branched polyethyleneimine (PEI). After 72hours  the  lungs were  harvested  and  processed  to measure  luciferase  activity  (RLU/mg  protein)  in  lung  homogenates. A naïve group (n=3; non‐transfected) was  included as a negative control. (C) Human  HEK293T cells were seeded (3x105 cells per well) and after 24 hours each well (n=4) was transfected  with plasmid DNA (2µg plasmid per well) complexed with linear polyethyleneimine (PEIPro). After 48  hours cells were lysed and assayed for luciferase activity measured in duplicate and expressed as RLU  per mg of protein.     Figure  15:  Representative  immunohistochemistry  images  showing  EGFP  expression  in  lung  parenchyma sections taken from mice treated with lentiviral vectors expressing EGFP from different  enhancer/mSPB constructs.  Representative images from (A) CMVenh (Alv‐01, SEQ ID NO: 46), (B) Slc2  (Alv‐02, SEQ ID NO: 47) and (C) Vegf2 (Alv‐03, SEQ ID NO: 48) groups are shown. On a background of  (DAPI stained blue) cell nuclei, white arrows  indicate examples of co‐localisation of  (yellow) signal  from EGFP transgene expression (green) and ATII cell‐specific marker SP‐B (red).  Magnification: scale  bar shown.    Figure 16:   To examine the  level and duration of expression from these promoters, female BALB/c  mice (n=10) were dosed with rSIV.F/HN encoding firefly  luciferase under the control of CMV, hCEF,  mSPB, Alv1, Alv2 or Alv3 promoter, or formulation buffer (TSSM) as a negative control (2.5e8 TU per  mouse via intranasal administration). (A)7, 14, 28 days, and 3 and 6 months after transduction, mice  were administered D‐luciferin and imaged for luciferase activity in the lung and nasal cavity. Signal in  regions  of  interest  (ROI)  capturing  the  chest  and  nasal  areas  were  measured  in  photons/second/cm2/sr. Average radiance in the lung (B) in addition to area under the curve (AUC)  analysis (C) were plotted. (D) Specificity for expression  in the  lung parenchyma was determined by  calculating the ratio of signal in the lung (indicative of alveolar and airway cell transduction) to signal  in the nasal cavity (indicative of airway cell transduction) 28 days after dosing. Significant differences  were determined by Kruskall‐Wallis (H(6)=55.53, p<0.0001) with Dunn’s post hoc multiple comparison  test  comparing  vector  groups using  the hCEF promoter. Data  are  shown  as  individual  values  and  mean±SEM. A calculated p value of <0.05 was deemed significant (p<0.01 ,**; p<0.0001 ,****).     Figure 17:  The hCEF promoter drives expression in multiple lung cell types, including cells lining the  airway epithelia, whereas the novel  lung promoters mSPB, Alv‐1, Alv‐2 and Alv‐3 were designed to  drive targeted expression in ATII cells in the parenchyma. To investigate the expression profile, BALB/c  mice (n=5‐10 per group) were dosed with SIV.F/HN expressing EGFP from the hCEF, mSPB or Alv‐1,  Alv‐2 or Alv‐3 promoters. Mice were culled on day 14 post‐dosing and whole  lungs processed  for  cryosectioning. Sections were DAPI stained and imaged (using Axioscanner) to visualise native EGFP‐ positive cells. Representative images of EGFP positive cells in the airway (Aw) and parenchyma (P) are  shown after administration of (A) rSIV.F/HN hCEF EGFP or (B) rSIV.F/HN mSP‐B EGFP. Images of lung  sections were  further  analysed  (using  Visiopharm  software) which  required manual  indication  of  airways and parenchyma. The percentage of EGFP‐positive cells  from  the  total  lung and  from  the  airways was used to (C) estimate the percentage of EGFP‐positive cells observed in the parenchyma  from  each  promoter.  Collectively,  the  data  presented  indicate  that  the  novel  lung  promoters  (especially Alv‐2  and Alv‐3)  show  expression mainly  in  the  parenchyma  compared with  the  hCEF  promoter.    Figure  18:  The  effect  of mSPB,  Alv‐1,  Alv‐2  and  Alv‐3  driven  SFP‐B  expression  on  transepithelial  electrical resistance (TEER) in a Surfactant Air Liquid Interface (SALI) model.  (A & B) TEER values from  SALI cultures generated from the H441 SP‐B KO cells were lower than the parental H441 cell line at 14  days post airlift constituting a phenotypic defect. (C) At 5 days post transduction, rSIV.F/HN expressing  EGFP  control  vector  did  not  increase  the  TEER  in  H441  SP‐B  KO  cells  SALI  cultures,  however  transduction with any/all of the vectors expressing SP‐B showed (D) correction of the TEER towards  normal, calculated as a percentage of the mock‐transduced parental H441 cells. ns: statistically non‐ significant.    Figure 19:  Graph showing the  % transduction of HEK293T cells with rSIV.F/HN vector encoding EGFP  under CMV promoter control was mixed with synthetic surfactants Beractat or Proactant alfa, or with  TSSM (vehicle control).  The presence of either synthetic surfactant had no significant effect on the %  transduction of the HEK293T cells with the rSIV.F/HN vector.      Figure  20:    The  effect  of  pulmonary  surfactants  on  rSIV.FHN  transduction  of  the  murine  lung.   rSIV.F/HN vector expressing firefly luciferase from the hCEF promoter (rSIV.F/HN hCEF Flux) was mixed  1:1 with TSSM (vehicle control) or BLES or Curosurf (to generate 2E8 TU in 100uL per dose). (A) Time  course of luciferase expression in the lungs of each mouse on days 7, 14, 28, 112 days post‐dosing. (B)  Graph  of  luciferase  activity  in  the  lungs  plotted  as  area  under  the  curve  (AUC)  analysis  (log10‐ transformed). (C)   Graph of the fold difference  in  luciferase expression relative to mice dosed with  control TSSM:vector.      DETAILED DESCRIPTION OF THE INVENTION    Definitions  Unless  defined  otherwise,  all  technical  and  scientific  terms  used  herein  have  the  same  meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.  Singleton, et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY, 20 ED., John Wiley and  Sons, New York (1994), and Hale & Marham, THE HARPER COLLINS DICTIONARY OF BIOLOGY, Harper  Perennial, NY (1991) provide the skilled person with a general dictionary of many of the terms used in  this disclosure. The meaning and scope of the terms should be clear; however,  in the event of any  latent  ambiguity,  definitions  provided  herein  take  precedent  over  any  dictionary  or  extrinsic  definition. It should be understood that this invention is not limited to the particular methodology,  protocols, and reagents, etc., described herein and as such can vary.  This disclosure is not limited by the exemplary methods and materials disclosed herein, and  any methods and materials similar or equivalent to those described herein can be used in the practice  or  testing of embodiments of  this disclosure.   The  terminology used herein  is  for  the purpose of  describing  particular  embodiments  only,  and  is  not  intended  to  limit  the  scope  of  the  present  invention, which is defined solely by the claims.  The description of embodiments of the disclosure is not intended to be exhaustive or to limit  the disclosure to the precise form disclosed. While specific embodiments of, and examples for, the  disclosure are described herein for illustrative purposes, various equivalent modifications are possible  within the scope of the disclosure, as those skilled in the relevant art will recognize. For example, while  method  steps or  functions are presented  in a given order, alternative embodiments may perform  functions in a different order, or functions may be performed substantially concurrently. The teachings  of the disclosure provided herein can be applied to other procedures or methods as appropriate. The  various embodiments described herein can be combined to provide further embodiments. Aspects of  the disclosure can be modified, if necessary, to employ the compositions, functions and concepts of  the  above  references  and  application  to  provide  yet  further  embodiments  of  the  disclosure.  Moreover, due  to biological  functional equivalency  considerations,  some  changes  can be made  in  protein structure without affecting  the biological or chemical action  in kind or amount. These and  other changes can be made to the disclosure in light of the detailed description. All such modifications  are intended to be included within the scope of the appended claims.  The headings provided herein are not limitations of the various aspects or embodiments of  this disclosure.   As used herein,  the  term "capable of' when used with a verb, encompasses or means  the  action  of  the  corresponding  verb.  For  example,  "capable  of  interacting"  also means  interacting,  "capable of  cleaving" also means  cleaves,  "capable of binding" also means binds and  "capable of  specifically targeting…" also means specifically targets.  Numeric ranges are inclusive of the numbers defining the range. Where a range of values is  provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless  the  context  clearly  dictates  otherwise,  between  the  upper  and  lower  limits  of  that  range  is  also  specifically disclosed. Each smaller range between any stated value or  intervening value  in a stated  range  and  any other  stated or  intervening  value  in  that  stated  range  is  encompassed within  this  disclosure. The upper and  lower  limits of  these  smaller  ranges may  independently be  included or  excluded in the range, and each range where either, neither or both limits are included in the smaller  ranges  is also encompassed within  this disclosure,  subject  to any  specifically excluded  limit  in  the  stated range. Where the stated range includes one or both of the limits, ranges excluding either or  both of those included limits are also included in this disclosure.  As used herein, the articles "a" and “an” may refer to one or to more than one (e.g. to at least  one) of the grammatical object of the article. Further, unless otherwise required by context, singular  terms shall include pluralities and plural terms shall include the singular. In this application, the use of  "or" means "and/or" unless stated otherwise. Furthermore, the use of the term "including", as well as  other forms, such as "includes" and "included", is not limiting.  “About” may generally mean an acceptable degree of error for the quantity measured given  the nature or precision of the measurements. Exemplary degrees of error are within 20 percent (%),  typically, within 10%, and more typically, within 5% of a given value or range of values. Preferably, the  term “about” shall be understood herein as plus or minus (±) 5%, preferably ± 4%, ± 3%, ± 2%, ± 1%, ±  0.5%, ± 0.1%, of the numerical value of the number with which it is being used.  The term "consisting of'' refers to compositions, methods, and respective components thereof  as  described  herein,  which  are  exclusive  of  any  element  not  recited  in  that  description  of  the  invention.  As used herein the term "consisting essentially of'' refers to those elements required for a  given  invention. The term permits the presence of elements that do not materially affect the basic  and  novel  or  functional  characteristic(s)  of  that  invention  (i.e.  inactive  or  non‐immunogenic  ingredients).  Embodiments described herein as “comprising” one or more features may also be considered  as disclosure of  the  corresponding embodiments “consisting of” and/or “consisting essentially of”  such features.   Concentrations,  amounts,  volumes,  percentages  and  other  numerical  values  may  be  presented herein in a range format. It is also to be understood that such range format is used merely  for convenience and brevity and should be interpreted flexibly to include not only the numerical values  explicitly recited as the  limits of the range but also to  include all the  individual numerical values or  sub‐ranges  encompassed within  that  range  as  if  each  numerical  value  and  sub‐range  is  explicitly  recited.  A "vector" or "construct" (sometimes referred to as gene delivery or gene transfer "vehicle")  refers to a macromolecule or complex of molecules comprising a polynucleotide to be delivered to a  host cell, either  in vitro or  in vivo. A vector can be a  linear or a circular molecule. A vector of  the  invention may be viral or non‐viral. All disclosure herein in relation vectors of the invention applies  equally to viral and non‐viral vectors unless otherwise stated. All disclosure in relation to viral vectors  of the invention applies equally and without reservation to lentiviral (e.g. SIV) vectors, particularly to  lentiviral (e.g. SIV) vectors that are pseudotyped with hemagglutinin‐neuraminidase (HN) and fusion  (F) proteins from a respiratory paramyxovirus (also referred to herein as SIV F/HN or SIV‐FHN).  As used herein, the term "plasmid", refers to a common type of non‐viral vector. A plasmid is  an  extra‐chromosomal  DNA molecule  separate  from  the  chromosomal  DNA which  is  capable  of  replicating  independently  of  the  chromosomal DNA.  Preferably  a  plasmid  is  circular  and may  be  double‐stranded.  The terms "nucleic acid cassette”, “nucleic acid construct", "expression cassette" and "nucleic  acid expression cassette" are used interchangeably to mean a nucleic acid molecule that is capable of  directing  transcription.  A  nucleic  acid  cassette  includes,  at  the  least,  a  promoter  or  a  structure  functionally equivalent to a promoter and a nucleic acid sequence to be transcribed. Thus, a nucleic  acid cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter  and a nucleic acid sequence encoding a protein of  interest. In the present  invention, a nucleic acid  cassette includes, at the least, a promoter or a structure functionally equivalent to a promoter, and a  nucleic acid encoding a therapeutic protein. A nucleic acid cassette may include additional elements,  such as an enhancer, and/or a transcription termination signal.  As used herein, the terms “transduced” and “modified” are used interchangeably to describe  cells which have been modified to express a transgene of interest. Typically the modification occurs  through transduction of the cells.  As used herein, the terms “titre” and “yield” are used interchangeably to mean the amount  of viral  (e.g.  lentiviral, particularly SIV) vector produced by a method of  the  invention. Titre  is  the  primary benchmark  characterising manufacturing efficiency, with higher  titres generally  indicating  that more vector is manufactured (e.g. using the same amount of reagents). Titre or yield may relate  to the number of vector genomes that have integrated into the genome of a target cell (integration  titre), which is a measure of “active” virus particles, i.e. the number of particles capable of transducing  a cell. Transducing units (TU/mL also referred to as TTU/mL) is a biological readout of the number of  host cells that get transduced under certain tissue culture/virus dilutions conditions, and is a measure  of the number of “active” virus particles. The total number of (active+inactive) virus particles may also  be determined using any appropriate means, such as by measuring either how much Gag is present in  the test solution or how many copies of viral RNA are in the test solution. Assumptions are then made  that a viral (e.g. lentivirus, particularly SIV) particle contains either 2000 Gag molecules or 2 viral RNA  molecules.  Once  total  particle  number  and  a  transducing  titre/TU  have  been  measured,  a  particle:infectivity ratio calculated.   Amino  acids  are  referred  to  herein  using  the  name  of  the  amino  acid,  the  three‐letter  abbreviation or the single letter abbreviation.   Unless otherwise  indicated, any nucleic acid  sequences are written  left  to  right  in 5'  to 3'  orientation;  amino  acid  sequences  are  written  left  to  right  in  amino  to  carboxy  orientation,  respectively.  As used herein,  the  terms "protein" and "polypeptide" are used  interchangeably herein  to  designate a series of amino acid residues, connected to each other by peptide bonds between the  alpha‐amino and carboxyl groups of adjacent residues. The terms "protein", and "polypeptide" refer  to  a  polymer  of  amino  acids,  including  modified  amino  acids  (e.g.,  phosphorylated,  glycated,  glycosylated,  etc.)  and  amino  acid  analogues,  regardless  of  its  size  or  function.  "Protein"  and  "polypeptide" are often used in reference to relatively large polypeptides, whereas the term "peptide"  is often used  in reference to small polypeptides, but usage of these terms  in the art overlaps. The  terms "protein" and "polypeptide" are used interchangeably herein when referring to a gene product  and  fragments  thereof. Thus, exemplary polypeptides or proteins  include gene products, naturally  occurring  proteins,  homologs,  orthologs,  paralogs,  fragments  and  other  equivalents,  variants,  fragments, and analogues of the foregoing.  As used herein, the terms “polynucleotides”, "nucleic acid" and "nucleic acid sequence" refers  to  any  molecule,  preferably  a  polymeric  molecule,  incorporating  units  of  ribonucleic  acid,  deoxyribonucleic  acid  or  an  analogue  thereof.  The  nucleic  acid  can  be  either  single‐stranded  or  double‐stranded. A single‐stranded nucleic acid can be one nucleic acid strand of a denatured double‐  stranded DNA Alternatively,  it can be a single‐stranded nucleic acid not derived  from any double‐ stranded DNA. In one aspect, the nucleic acid can be DNA. In another aspect, the nucleic acid can be  RNA Suitable nucleic acid molecules are DNA, including genomic DNA or cDNA. Other suitable nucleic  acid  molecules  are  RNA,  including  siRNA,  shRNA,  and  antisense  oligonucleotides.  The  terms  “transgene”  and  “gene”  are  also  used  interchangeably  and  both  terms  encompass  fragments  or  variants thereof encoding the target protein.  The  transgenes  of  the  present  invention  include  nucleic  acid  sequences  that  have  been  removed  from  their  naturally  occurring  environment,  recombinant  or  cloned  DNA  isolates,  and  chemically synthesized analogues or analogues biologically synthesized by heterologous systems.  Minor variations  in  the amino acid  sequences of  the  invention are contemplated as being  encompassed by the present invention, providing that the variations in the amino acid sequence(s)  maintain at least 60%, at least 70%, more preferably at least 80%, at least 85%, at least 90%, at least  95%, and most preferably at least 97% or at least 99% sequence identity to the amino acid sequence  of the invention or a fragment thereof as defined anywhere herein. The term homology is used herein  to mean identity. As such, the sequence of a variant or analogue sequence of an amino acid sequence  of the invention may differ on the basis of substitution (typically conservative substitution) deletion  or insertion. Proteins comprising such variations are referred to herein as variants.  Proteins of the invention may include variants in which amino acid residues from one species  are  substituted  for  the  corresponding  residue  in another  species, either at  the conserved or non‐ conserved positions. Variants of protein molecules disclosed herein may be produced and used in the  present  invention.  Following  the  lead  of  computational  chemistry  in  applying  multivariate  data  analysis  techniques  to  the  structure/property‐activity  relationships  [see  for  example, Wold,  et  al.  Multivariate data analysis in chemistry. Chemometrics‐Mathematics and Statistics in Chemistry (Ed.:  B.  Kowalski);  D.  Reidel  Publishing  Company,  Dordrecht,  Holland,  1984  (ISBN  90‐277‐1846‐6]  quantitative activity‐property relationships of proteins can be derived using well‐known mathematical  techniques,  such  as  statistical  regression,  pattern  recognition  and  classification  [see  for  example  Norman  et  al.  Applied  Regression  Analysis.  Wiley‐lnterscience;  3rd  edition  (April  1998)  ISBN:  0471170828; Kandel, Abraham et al. Computer‐Assisted Reasoning in Cluster Analysis. Prentice Hall  PTR,  (May 11, 1995),  ISBN: 0133418847; Krzanowski, Wojtek. Principles of Multivariate Analysis: A  User's  Perspective  (Oxford  Statistical  Science  Series,  No  22  (Paper)).  Oxford  University  Press;  (December 2000),  ISBN: 0198507089; Witten,  Ian H. et al Data Mining: Practical Machine Learning  Tools  and  Techniques  with  Java  Implementations.  Morgan  Kaufmann;  (October  11,  1999),  ISBN:1558605525; Denison David G. T. (Editor) et al Bayesian Methods for Nonlinear Classification and  Regression  (Wiley  Series  in  Probability  and  Statistics).  John  Wiley  &  Sons;  (July  2002),  ISBN:  0471490369; Ghose, Arup K. et al. Combinatorial Library Design and Evaluation Principles, Software,  Tools, and Applications  in Drug Discovery.  ISBN: 0‐8247‐0487‐8]. The properties of proteins can be  derived  from empirical and  theoretical models  (for example, analysis of  likely  contact  residues or  calculated  physicochemical  property)  of  proteins  sequence,  functional  and  three‐dimensional  structures and these properties can be considered individually and in combination.  Amino  acids  are  referred  to  herein  using  the  name  of  the  amino  acid,  the  three‐letter  abbreviation or the single letter abbreviation. The term “protein", as used herein, includes proteins,  polypeptides, and peptides. As used herein, the term “amino acid sequence” is synonymous with the  term “polypeptide” and/or the term “protein”. In some instances, the term “amino acid sequence” is  synonymous  with  the  term  “peptide”.  The  terms  "protein"  and  "polypeptide"  are  used  interchangeably herein. In the present disclosure and claims, the conventional one‐letter and three‐ letter codes  for amino acid residues may be used. The 3‐letter code  for amino acids as defined  in  conformity with  the  IUPACIUB  Joint  Commission  on  Biochemical Nomenclature  (JCBN).  It  is  also  understood that a polypeptide may be coded for by more than one nucleotide sequence due to the  degeneracy of the genetic code.  Amino acid residues at non‐conserved positions may be substituted with conservative or non‐ conservative residues. In particular, conservative amino acid replacements are contemplated.   A “conservative amino acid substitution” is one in which the amino acid residue is replaced  with an amino acid residue having a similar side chain. Families of amino acid residues having similar  side chains have been defined in the art, including basic side chains (e.g., lysine, arginine, or histidine),  acidic  side  chains  (e.g.,  aspartic  acid or  glutamic  acid), uncharged polar  side  chains  (e.g.,  glycine,  asparagine, glutamine, serine, threonine, tyrosine, or cysteine), nonpolar side chains (e.g., alanine,  valine,  leucine,  isoleucine, proline, phenylalanine, methionine, or  tryptophan), beta‐branched  side  chains  (e.g.,  threonine,  valine,  isoleucine)  and  aromatic  side  chains  (e.g.,  tyrosine, phenylalanine,  tryptophan, or histidine). Thus, if an amino acid in a polypeptide is replaced with another amino acid  from the same side chain family, the amino acid substitution  is considered to be conservative. The  inclusion of conservatively modified variants in a protein of the invention does not exclude other forms  of variant, for example polymorphic variants, interspecies homologs, and alleles.  “Non‐conservative amino acid substitutions”  include those  in which (i) a residue having an  electropositive side chain (e.g., Arg, His or Lys)  is substituted for, or by, an electronegative residue  (e.g., Glu or Asp), (ii) a hydrophilic residue (e.g., Ser or Thr)  is substituted for, or by, a hydrophobic  residue (e.g., Ala, Leu,  Ile, Phe or Val), (iii) a cysteine or proline  is substituted for, or by, any other  residue, or (iv) a residue having a bulky hydrophobic or aromatic side chain (e.g., Val, His, Ile or Trp) is  substituted for, or by, one having a smaller side chain (e.g., Ala or Ser) or no side chain (e.g., Gly).  “Insertions” or  “deletions” are  typically  in  the  range of about 1, 2, or 3 amino acids. The  variation  allowed may  be  experimentally  determined  by  systematically  introducing  insertions  or  deletions of amino acids  in a protein using recombinant DNA techniques and assaying the resulting  recombinant variants for activity. This does not require more than routine experiments for a skilled  person.  A “fragment” of a polypeptide comprises at least 50%, at least 60%, at least 70%, at least 80%,  at least 90%, at least 95%, at least 97% or more of the original polypeptide. For example, a fragment  may comprise at least 5, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least  16, at least 17, at least 18, at least 19, at least 20 or more amino acids of the protein from which it is  derived. A fragment may be continuous or discontinuous, preferably continuous.   The polynucleotides of the present  invention may be prepared by any means known  in the  art. For example, large amounts of the polynucleotides may be produced by replication in a suitable  host cell. The natural or synthetic DNA fragments coding for a desired fragment will be incorporated  into recombinant nucleic acid constructs, typically DNA constructs, capable of introduction into and  replication  in  a  prokaryotic  or  eukaryotic  cell.  Usually  the  DNA  constructs  will  be  suitable  for  autonomous replication in a unicellular host, such as yeast or bacteria, but may also be intended for  introduction to and  integration within the genome of a cultured  insect, mammalian, plant or other  eukaryotic cell lines.  The polynucleotides of the present  invention may also be produced by chemical synthesis,  e.g. by the phosphoramidite method or the tri‐ester method, and may be performed on commercial  automated  oligonucleotide  synthesizers.  A  double‐stranded  fragment may  be  obtained  from  the  single stranded product of chemical synthesis either by synthesizing the complementary strand and  annealing the strand together under appropriate conditions or by adding the complementary strand  using DNA polymerase with an appropriate primer sequence.  When applied to a nucleic acid sequence, the term “isolated”  in the context of the present  invention denotes that the polynucleotide sequence has been removed from its natural genetic milieu  and  is  thus  free  of  other  extraneous  or  unwanted  coding  sequences  (but may  include  naturally  occurring 5' and 3' untranslated regions such as promoters and terminators), and is in a form suitable  for use within genetically engineered protein production systems. Such isolated molecules are those  that are separated from their natural environment.   In view of the degeneracy of the genetic code, considerable sequence variation  is possible  among the polynucleotides of the present  invention. Degenerate codons encompassing all possible  codons for a given amino acid are set forth below:  Amino Acid  Codons  Degenerate Codon  Cys  TGC TGT  TGY  Ser  AGC AGT TCA TCC TCG TCT  WSN  Thr  ACA ACC ACG ACT  ACN  Pro  CCA CCC CCG CCT   CCN  Ala  GCA GCC GCG GCT  GCN  Gly  GGA GGC GGG GGT  GGN  Asn  AAC AAT  AAY  Asp  GAC GAT  GAY  Glu  GAA GAG  GAR  Gln  CAA CAG  CAR  His  CAC CAT  CAY  Arg  AGA AGG CGA CGC CGG CGT  MGN  Lys  AAA AAG  AAR  Met  ATG  ATG  Ile  ATA ATC ATT  ATH  Leu  CTA CTC CTG CTT TTA TTG  YTN  Val  GTA GTC GTG GTT  GTN  Phe  TTC TTT  TTY  Tyr  TAC TAT  TAY  Trp  TGG  TGG  Ter  TAA TAG TGA  TRR  Asn/ Asp    RAY  Glu/ Gln    SAR  Any    NNN    One  of  ordinary  skill  in  the  art will  appreciate  that  flexibility  exists when  determining  a  degenerate codon, representative of all possible codons encoding each amino acid. For example, some  polynucleotides  encompassed  by  the  degenerate  sequence  may  encode  variant  amino  acid  sequences, but one of ordinary skill in the art can easily identify such variant sequences by reference  to the amino acid sequences of the present invention.  A  “variant”  nucleic  acid  sequence  has  substantial  homology  or  substantial  similarity  to  a  reference nucleic acid sequence (or a fragment thereof). A nucleic acid sequence or fragment thereof  is “substantially homologous” (or “substantially identical”) to a reference sequence if, when optimally  aligned  (with  appropriate  nucleotide  insertions  or  deletions)  with  the  other  nucleic  acid  (or  its  complementary strand), there is nucleotide sequence identity in at least about 70%, 75%, 80%, 85, 90,  91,  92,  93,  94,  95,  96,  97,  98,  99  or  more%  of  the  nucleotide  bases.  Methods  for  homology  determination of nucleic acid sequences are known in the art.  Alternatively,  a  “variant”  nucleic  acid  sequence  is  substantially  homologous  with  (or  substantially  identical  to)  a  reference  sequence  (or  a  fragment  thereof)  if  the  “variant”  and  the  reference sequence they are capable of hybridizing under stringent (e.g. highly stringent) hybridization  conditions.  Nucleic  acid  sequence  hybridization  will  be  affected  by  such  conditions  as  salt  concentration  (e.g. NaCl),  temperature,  or  organic  solvents,  in  addition  to  the  base  composition,  length of the complementary strands, and the number of nucleotide base mismatches between the  hybridizing  nucleic  acids,  as  will  be  readily  appreciated  by  those  skilled  in  the  art.  Stringent  temperature conditions are preferably employed, and generally  include  temperatures  in excess of  30°C,  typically  in  excess  of  37°C  and  preferably  in  excess  of  45°C.  Stringent  salt  conditions will  ordinarily be less than 1000 mM, typically less than 500 mM, and preferably less than 200 mM. The  pH is typically between 7.0 and 8.3. The combination of parameters is much more important than any  single parameter.   Methods of determining nucleic acid percentage sequence identity are known in the art. By  way of example, when assessing nucleic acid sequence identity, a sequence having a defined number  of contiguous nucleotides may be aligned with a nucleic acid sequence (having the same number of  contiguous nucleotides)  from  the corresponding portion of a nucleic acid sequence of  the present  invention. Tools known in the art for determining nucleic acid percentage sequence identity include  Nucleotide BLAST (as described below).  One of ordinary skill in the art appreciates that different species exhibit “preferential codon  usage”. As used herein, the term “preferential codon usage” refers to codons that are most frequently  used in cells of a certain species, thus favouring one or a few representatives of the possible codons  encoding each amino acid. For example, the amino acid threonine (Thr) may be encoded by ACA, ACC,  ACG, or ACT, but in mammalian host cells ACC is the most commonly used codon; in other species,  different codons may be preferential. Preferential codons  for a particular host cell  species can be  introduced  into the polynucleotides of the present  invention by a variety of methods known  in the  art. Introduction of preferential codon sequences into recombinant DNA can, for example, enhance  production of the protein by making protein translation more efficient within a particular cell type or  species. Thus, according to the invention, in addition to the gag‐pol genes any nucleic acid sequence  may be codon‐optimised for expression in a host or target cell. In particular, the vector genome (or  corresponding plasmid),  the REV gene  (or  corresponding plasmid),  the  fusion protein  (F) gene  (or  correspond plasmid) and/or the hemagglutinin‐neuraminidase (HN) gene (or corresponding plasmid,  or any combination thereof may be codon‐optimised.  A “fragment” of a polynucleotide of  interest comprises a series of consecutive nucleotides  from  the  sequence  of  said  full‐length  polynucleotide.  By  way  of  example,  a  “fragment”  of  a  polynucleotide of interest may comprise (or consist of) at least 600 consecutive nucleotides from the  sequence of said polynucleotide (e.g. at  least 600, 650, 700, 750, 800 850, 900, or 950 consecutive  nucleic acid residues of said polynucleotide). Typically, a fragment as defined herein retains the same  function as the full‐length polynucleotide.  The  terms  "decrease",  "reduced",  "reduction",  or  "inhibit"  are  all  used  herein  to mean  a  decrease  by  a  statistically  significant  amount.  The  terms  "reduce,"  "reduction"  or  "decrease"  or  "inhibit" typically means a decrease by at least 10% as compared to a reference level (e.g. the absence  of a given treatment) and can include, for example, a decrease by at least about 10%, at least about  20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about  45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about  70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about  95%, at  least about 98%, at  least about 99%  , or more. As used herein, "reduction" or "inhibition"  encompasses  a  complete  inhibition  or  reduction  as  compared  to  a  reference  level.  "Complete  inhibition" is a 100% inhibition (i.e. abrogation) as compared to a reference level.   The terms "increased", "increase", "enhance", or "activate" are all used herein to mean an  increase by a statically significant amount. The terms "increased", "increase", "enhance", or "activate"  can mean an increase of at least 25%, at least 50% as compared to a reference level, for example an  increase of at least about 50%, or at least about 75%, or at least about 80%, or at least about 90%, at  least about 95%, or at least about 98%, or at least about 99%, or at least about 100%, or at least about  250% or more compared with a reference level, or at least about a 1.5‐fold, or at least about a 2‐fold,  or at least about a 2.5‐fold, or at least about a 3‐fold, or at least about a 4‐fold, or at least about a 5‐ fold or at least about a 10‐fold increase, or any increase between 1.5‐fold and 10‐fold or greater as  compared  to a reference  level.  In  the context of a yield or  titre, an "increase"  is an observable or  statistically significant increase in such level.  The terms "individual”, "subject”, and "patient”, are used interchangeably herein to refer to  a mammalian subject for whom diagnosis, prognosis, disease monitoring, treatment, therapy, and/or  therapy  optimisation  is  desired.  The mammal  can  be  (without  limitation)  a  human,  non‐human  primate, mouse, rat, dog, cat, horse, or cow. In a preferred embodiment, the individual, subject, or  patient is a human. An “individual” may be an adult, juvenile or infant. An “individual” may be male  or female.  A "subject  in need" of treatment for a particular condition can be an  individual having that  condition, diagnosed as having that condition, or at risk of developing that condition.  A subject can be one who has been previously diagnosed with or identified as suffering from  or having a condition in need of treatment or one or more complications or symptoms related to such  a condition, and optionally, have already undergone treatment for a condition as defined herein or  the one or more complications or symptoms related to said condition. Alternatively, a subject can also  be one who has not been previously diagnosed as having a condition as defined herein or one or more  or  symptoms or  complications  related  to  said  condition.  For  example,  a  subject  can be one who  exhibits one or more risk factors for a condition, or one or more or symptoms or complications related  to said condition or a subject who does not exhibit risk factors.  As used herein, the term “healthy  individual” refers to an  individual or group of  individuals  who are in a healthy state, e.g. individuals who have not shown any symptoms of the disease, have  not been diagnosed with the disease and/or are not likely to develop the disease e.g. cystic fibrosis  (CF) or any other disease described herein). Preferably said healthy individual(s) is not on medication  affecting CF and has not been diagnosed with any other disease. The one or more healthy individuals  may have a  similar  sex, age, and/or body mass  index  (BMI) as compared with  the  test  individual.  Application of standard statistical methods used in medicine permits determination of normal levels  of expression in healthy individuals, and significant deviations from such normal levels.  Herein the terms “control” and “reference population” are used interchangeably.   The  term  “pharmaceutically  acceptable”  as  used  herein means  approved  by  a  regulatory  agency  of  the  Federal  or  a  state  government,  or  listed  in  the  U.S.  Pharmacopeia,  European  Pharmacopeia or other generally recognized pharmacopeia  The publications discussed herein are provided solely for their disclosure prior to the filing  date  of  the  present  application.  Nothing  herein  is  to  be  construed  as  an  admission  that  such  publications constitute prior art to the claims appended hereto.   Disclosure related to the various methods of the invention are intended to be applied equally  to other methods, therapeutic uses or methods, the data storage medium or device, the computer  program product, and vice versa.    Surfactant Protein B (SFTPB) promoter fragments  The full‐length SFTPB (surfactant protein B; SFTPB) promoter has previously been defined as  a  sequence  that  could  be  amplified  from  human  genomic  DNA  with  the  PCR  primers  ATTTGAGCTCTTCTTTCTGCTGAACCATCG  (sense,  SEQ  ID  NO:  49)  and  TCTTAGATCTGTCAGACAGCTCTGGGTTCC  (antisense,  SEQ  ID  NO:  50),  wherein  the  underlines  sequences  align with  GenBank  NCBI  Reference  Sequence:  NG_016967.1  (version  1,  accessed  18  November  2022)  while  the  additional  5’  sequences  provide  SacI  (sense)  and  BglII  (antisense)  restriction enzyme sites. The forward primer binds to the sense strand from bases 4845 to 4864  in  NG_016967.1.  The  reverse  primer  binds  to  the  reverse  complement  of  bases  5797  to  5816  in  NG_016967.1. Thus, the primer pair define a 972 bp genomic fragment. The 972 bp genomic fragment  includes all of exon 1 of the SFTPB gene (bases 5543 to 5565 of NG_016967.1), wherein the A at base  5543 is reported as the starting nucleotide of the mRNA generated by the SFTPB promoter and the  initiating ATG at bases 5559 to 5561 of NG_016967.1 encodes the first methionine of pre‐pro‐SFTPB.   This  972  bp  genomic  fragment  further  includes  a  part  of  intron  1  from  bases  5626  to  5816  of  NG_016967.1.  Preferably, references herein to a full‐length SFTPB promoter refer specifically to this  972 bp genomic fragment, which is present SEQ ID NO: 2.  As  described  and  exemplified  herein,  the  present  inventors  have  identified  and  isolated  functional  fragments  of  the  SFTPB  (surfactant  protein  B)  gene  promoter.  The  SFTPB    promoter  fragments generated by the inventors are able to drive cell‐specific expression of transgenes in the  lung parenchyma. The fragments of the SFTPB gene promoter are shorter than the 972 bp genomic  fragment previously identified, i.e. are shorter than the full‐length SFTPB promoter of SEQ ID NO: 2  previously reported in the art. Therefore, the SFTPB  promoter fragments of the invention may be also  referred to as a core SFTPB promoters.   As discussed in more detail below and as exemplified herein,  the SFTPB  promoter fragments of the invention surprisingly increase transgene expression compared  with  the  full‐length SFTPB   promoter, and can do  so  in a  lung parenchymal cell preferred/specific  manner.  A promoter of the invention is an SFTPB promoter fragment as described herein.  Said SFTPB  promoter fragment is functional, also as described herein.  As used herein, the term “promoter of the  invention” is used interchangeably with the terms “SFTPB promoter fragment” and “functional SFTPB  promoter fragment”.   Thus, all disclosure herein to a (functional) SFTPB promoter fragment applies  equally and without reservation to all promoters of the invention.  An SFTPB promoter fragment of the invention comprises or consists of a core SFTPB promoter  fragment, as described herein, or a variant thereof.    An SFTPB promoter fragment of the invention may comprise or consist of a fragment of SEQ  ID NO: 2, which lack all or part of SFTPB intron 1.  SFTPB intron 1 begins at base 5626 of NG_016967.1  (corresponding  to  residue  782  of  SEQ  ID  NO:  2)  and  corresponds  to    bases  5626  to  5816  of  NG_016967.1  (corresponding  to  residues 782  to 972 of SEQ  ID NO: 2).   Thus, an SFTPB promoter  fragment of the invention may comprise part (but not all) of the SFTPB intron 1.  An SFTPB promoter  fragment  of  the  invention  may  not  comprise  one  or  more  bases  from  bases  5626  to  5816  of  NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2). By way of non‐limiting example,  an SFTPB promoter fragment of the invention may comprise fewer than 180, fewer than 170, fewer  than 160, fewer than 150, fewer than 140, fewer than 130, fewer than 120, fewer than 110, fewer  than 100, fewer than 100, fewer than 90, fewer than 80, fewer than 70, fewer than 60, fewer than 50,  fewer than 40, fewer than 30, fewer than 20, or fewer than 10 contiguous bases from 5626 to 5816 of  NG_016967.1 (corresponding to residues 782 to 972 of SEQ ID NO: 2). Preferably, an SFTPB promoter  fragment of the invention does not comprise bases  5626 to 5816 of NG_016967.1 (corresponding to  residues 782 to 972 of SEQ ID NO: 2). Typically, the SFTPB promoter fragments of the invention do not  comprise intron 1 of the SFTPB gene (or any portion thereof).  Typically, an SFTPB promoter fragment of the invention comprises at least part of exon 1 of  the SFTPB gene.  SFTPB exon 1 begins at base 5543 of NG_016967.1 (corresponding to residue 699 of  SEQ ID NO: 2) and corresponds to   bases 5543 to 5565 of NG_016967.1 (corresponding to residues  699 to 781 of SEQ ID NO: 2).     Thus, an SFTPB promoter fragment of the invention may comprise the contiguous nucleotide  sequence of bases 5543 to 5565 of NG_016967.1 (corresponding to residues 699 to 781 of SEQ ID NO:  2). An SFTPB promoter fragment of the invention may comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13,  14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24 or 25 bases, preferably  1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13,  14, 15, or 16 bases, 3’ to base 5542 of NG_016967.1 (corresponding to residue 698 of SEQ ID NO: 2),  and wherein the additional bases correspond to a fragment of exon 1 of the SFTPB gene. By way of  non‐limiting example, when an SFTPB promoter fragment of the invention comprises 1 base 3’ to base  5542 of NG_016967.1, the additional base corresponds to base 5543 of NG_016967.1 (corresponding  to residue 699 of SEQ ID NO: 2), when an SFTPB promoter fragment of the invention comprises 2 bases  3’  to  base  5542  of  NG_016967.1,  the  additional  bases  correspond  to  bases  5543  to  5544  of  NG_016967.1  (corresponding  to  residues 699  to 700 of  SEQ  ID NO: 2), when an  SFTPB promoter  fragment of the invention comprises 3 bases 3’ to base 5542 of NG_016967.1, the additional bases  correspond to bases 5543 to 5545 of NG_016967.1  (corresponding to residues 699 to 701 of SEQ ID  NO: 2), when an SFTPB promoter  fragment of  the  invention comprises 4 bases 3’  to base 5542 of  NG_016967.1,  the  additional  bases  correspond  to  bases  5543  to  5546  of  NG_016967.1   (corresponding to residues 699 to 702 of SEQ ID NO: 2), and so on. Preferably, an SFTPB promoter  fragment of the invention comprises a portion of SFTPB corresponding to SEQ ID NO: 3.  An SFTPB promoter fragment of the invention may not comprise the SFTPB gene start codon.    Typically, an SFTPB promoter fragment of the invention may comprise of at least part of exon 1 of the  SFTPB gene, but does not comprise the SFTPB gene start codon.  Thus, an SFTPB promoter fragment  of the invention typically does not comprise the initiating ATG at bases 5559 to 5561 of NG_016967.1  (corresponding to residues 715 to 717 of SEQ ID NO: 2).  Preferably an SFTPB promoter fragment of  the invention also does not comprise any of the SFTPB gene sequence 3’ of this start codon.      An SFTPB promoter fragment of the invention may comprise fewer than 900 bases in length,  fewer than 850 bases in length, fewer than 800 bases in length, fewer than 750 bases in length, fewer  than 700 bases in length, fewer than 650 bases in length, fewer than 640 bases in length, fewer than  630 bases in length, fewer than 620 bases in length, fewer than 610 bases in length or fewer than 600  bases in length. Preferably, an SFTPB promoter fragment of the invention comprises  fewer than 800  bases in length. A particularly preferred SFTPB promoter fragment of the invention comprises fewer  than 700 bases in length. An SFTPB promoter fragment of the invention may consist of fewer than 900  bases in length, fewer than 850 bases in length, fewer than 800 bases in length, fewer than 750 bases  in  length, fewer than 700 bases  in  length, fewer than 650 bases  in  length, fewer than 640 bases  in  length, fewer than 630 bases in length, fewer than 620 bases in length, fewer than 610 bases in length  or fewer than 600 bases in length. Preferably, an SFTPB promoter fragment of the invention consists  of fewer than 800 bases in length. A particularly preferred SFTPB promoter fragment of the invention  may consist of  fewer  than 700 bases  in  length.   The exemplified SFTPB promoter  fragment of  the  invention consists of 630 or 635 bases in length.     An SFTPB promoter fragment of the invention  may comprise from 600 bases to 800 bases in  length, from 600 bases to 750 bases  in  length, from 600 bases to 700 bases  in  length, or from 600  bases to 650 bases in length. An SFTPB promoter fragment of the invention may  consist of from 600  bases to 800 bases in length, from 600 bases to 750 bases in length, from 600 bases to 700 bases in  length, or from 600 bases to 650 bases in length.   An SFTPB promoter fragment of the invention  may comprise from 630 bases to 800 bases in  length, from 635 bases to 800 bases in length, from 630 bases to 750 bases in length, from 635 bases  to 750 bases in length, from 630 bases to 700 bases in length, from 635 bases to 700 bases in length,  from 630 bases to 650 bases in length, or from 635 bases to 650 bases in length. An SFTPB promoter  fragment of the invention  may consist of from 630 bases to 800 bases in length, from 635 bases to  800 bases in  length, from 630 bases to 750 bases in  length, from 635 bases to 750 bases in  length,  from 630 bases to 700 bases in length, from 635 bases to 700 bases in length, from 630 bases to 650  bases in length, or from 635 bases to 650 bases in length.  An SFTPB promoter fragment of the invention may be a fragment of a mammalian or avian,  preferably a mammalian SFTPB promoter, i.e. an SFTPB promoter fragment of the invention may be a  mammalian or avian SFTPB promoter fragment.   By way of non‐limiting example, a mammalian SFTPB  promoter fragment may be a human, non‐human primate, mouse, rat, dog, cat, horse, or cow SFTPB  promoter fragment.   Preferably,  the  SFTPB promoter  fragment  is  a  fragment of  a human  SFTPB promoter,  i.e.  preferably the SFTPB promoter fragment of the invention is a human SFTPB promoter fragment. By  way  of  non‐limiting  example,  an  SFTPB  promoter  fragment  of  the  invention may  comprise  any  functional fragment of SEQ ID NO: 2.  Typically such an SFTPB promoter fragment of the invention is  of a length as described herein.  An SFTPB promoter fragment of the invention may comprise or consist  of a sequence having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at  least  95%,  at  least  96%,  at  least  97%,  at  least  98%,  at  least  99%,  at  least  99.5%,  or  at  least  99.8%sequence identity to a fragment of SEQ ID NO: 2, wherein the promoter retains the function of  the  SFTPB gene promoter, as defined herein.   Typically  such an  SFTPB promoter  fragment of  the  invention  is of a  length as described herein.   Alternatively or  in addition, such an SFTPB promoter  fragment of the invention may comprise any additional feature (e.g. lack the SFTPB gene start codon  and/or lack all or part of the SFTPB gene intron 1), as described herein.    An SFTPB promoter fragment of the invention may comprise a sequence having at least 70%,  at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least 97%,  at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to bases 4925‐ 5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). For example, an SFTPB  promoter fragment of the invention may comprise a sequence having at least 90% identity to bases  4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter  fragment of the  invention may comprise a sequence having at  least at  least 95%  identity  to bases  4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter  fragment  of  the  invention  may  comprise  the  sequence  of  bases  4925‐5554  of  NG_016967.1  (corresponding to residues 81 to 710 of SEQ ID NO: 2).    An SFTPB promoter fragment of the invention may consist of a sequence having at least 70%,  at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least  98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity or more to bases 4925‐5554 of  NG_016967.1  (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of  the invention may consist of a sequence having at least at least 90% identity to bases 4925‐5554 of  NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of  the invention may consist of a sequence having at least at least 95% identity to bases 4925‐5554 of  NG_016967.1  (corresponding to residues 81 to 710 of SEQ ID NO: 2). An SFTPB promoter fragment of  the  invention may consist of the sequence of bases 4925‐5554 of NG_016967.1   (corresponding to  residues 81 to 710 of SEQ ID NO: 2).  In some embodiments, an SFTPB promoter fragment of the invention derived from SEQ ID NO:  2,  such  as  those  described  above, may  comprise  a  substitution  at  residue  683  of  SEQ  ID NO:  2  (corresponding  to  base  5527  of NG_016967.1).    Preferably,  an  SFTPB  promoter  fragment  of  the  invention derived from SEQ ID NO: 2, such as those described above, may comprise substitution of an  alanine  residue  by  a  cytosine  at  residue  683  of  SEQ  ID  NO:  2  (corresponding  to  base  5527  of  NG_016967.1), in other words, may comprise an A683C substitution.   An exemplified SFTPB promoter fragment of the invention is SEQ ID NO: 1.   Accordingly, the  present invention provides an SFTPB promoter fragment which comprises a sequence having at least  70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least 96%, at least  97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or more to SEQ ID  NO: 1. For example, an SFTPB promoter fragment of the invention may comprise a sequence having  at least 90% identity to SEQ ID NO: 1. An SFTPB promoter fragment of the invention may comprise a  sequence  having  at  least  at  least  95%  identity  to  SEQ  ID NO:  1.  Preferably,  an  SFTPB  promoter  fragment of the invention may comprise the sequence of SEQ ID NO: 1.  The present  invention provides an SFTPB promoter fragment which consists of a sequence  having at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, at least 95%, at least  96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or at least 99.8% sequence identity, or  more to SEQ ID NO: 1. For example, an SFTPB promoter fragment of the invention may consist of a  sequence having at least 90% identity to SEQ ID NO: 1. An SFTPB promoter fragment of the invention  may consist of a sequence having at least at least 95% identity to SEQ ID NO: 1. Preferably, an SFTPB  promoter fragment of the invention may consist of the sequence of SEQ ID NO: 1.  An SFTPB promoter fragment of the invention may further comprise  up to 10 bases (e.g., 1,  2, 3, 4, 5, 6, 7, 8, 9 or 10 bases), up to 15 bases, up to 20 bases, up to 25 bases or up to 25 bases at the  5’ end.   Such additional bases may preferably be present when  said SFTPB promoter  fragment  (i)  comprises a sequence having at least 70% identity to SEQ ID NO: 1; and/or (ii) comprises a sequence  having at least 70% identity to bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710  of SEQ ID NO: 2); as described herein.    Where additional bases are present at  the 5’ end of an  SFTPB promoter  fragment of  the  invention, said additional bases may be bases which correspond to the corresponding number of bases  5’ to base 4925 of NG_016967.1 (corresponding to residue 81 of SEQ ID NO: 2). By way of non‐limiting  example, when an SFTPB promoter fragment of the invention comprises 10 bases 5’ to base 4925 of  NG_016967.1, the additional base corresponds to bases 4915 to 4924 of NG_016967.1 (corresponding  to residues 71 to 80 of SEQ ID NO: 2), or when an SFTPB promoter fragment of the invention comprises  20 bases 5’ to base 4925 of NG_016967.1, the additional bases correspond to bases 4905 to 4924 of  NG_016967.1 (corresponding to residues 61 to 80 of SEQ ID NO: 2), and so on. Preferably, an SFTPB  promoter fragment of the invention may further comprise up to 20 bases at the 5’ end, and particularly  preferably,  the up  to 20 additional bases correspond  to bases 5’ of  residue 4925 of NG_016967.1  (corresponding to residue 81 of SEQ ID NO: 2).  Alternatively  or  in  addition,  an  SFTPB  promoter  fragment  of  the  invention may  further  comprise  up to 10 bases (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 bases) bases at the 3’ end.  Such additional  bases may preferably be present when said SFTPB promoter fragment (i) comprises a sequence having  at least 70% identity to SEQ ID NO: 1; and/or (ii) comprises a sequence having at least 70% identity to  bases 4925‐5554 of NG_016967.1 (corresponding to residues 81 to 710 of SEQ ID NO: 2); as described  herein.    Where additional bases are present at  the 3’ end of an  SFTPB promoter  fragment of  the  invention, said additional bases may be bases which correspond to the corresponding number of bases  3’ to base 5554 of NG_016967.1 (corresponding to residue 710 of SEQ ID NO: 2). By way of non‐limiting  example, when an SFTPB promoter fragment of the invention comprises 10 bases 3’ to base 5554 of  NG_016967.1, the additional base corresponds to bases 5555 to 5564 of NG_016967.1 (corresponding  to residues 711 to 720 of SEQ ID NO: 2), when an SFTPB promoter fragment of the invention comprises  5 bases 3’ to base 5554 of NG_016967.1, the additional bases correspond to bases 5555 to 5559 of  NG_016967.1 (corresponding to residues 711 to 715 of SEQ ID NO: 2), or when an SFTPB promoter  fragment of the invention comprises 4 bases 3’ to base 5554 of NG_016967.1, the additional bases  correspond to bases 5555 to 5558 of NG_016967.1 (corresponding to residues 711 to 714 of SEQ ID  NO: 2), and so on. Preferably, an SFTPB promoter fragment of the invention may further comprise up  to 4 bases at the 3’ end, and particularly preferably, the up to 4 additional bases correspond to bases  3’ of residue 5554 of NG_016967.1 (corresponding to residue 710 of SEQ ID NO: 2).    An SFTPB promoter fragment of the invention may comprise additional sequences at the 5’  and/or 3’ end to facilitate molecular biology applications of said SFTPB promoter fragment.  By way of  non‐limiting example, one or more restriction enzyme site may be added at the 5’ and/or 3’ end of  the sequence.  Such restriction enzyme sites may be used to facilitate cloning of the SFTPB promoter  fragment  into a non‐viral vector  (e.g. plasmid) according to the  invention, or  into a manufacturing  plasmid for use in the production of a viral vector according to the invention.  When restriction enzyme  sites are present at both the 5’ and 3’ end of the SFTPB promoter fragment, each restriction enzyme  site may  be  selected  independently,  and  thus  the  5’  and  3’  sites may  be  the  same  or  different.   Examples of restriction enzyme sites that may be used include NheI and BgIII.  For example, a SFTPB  promoter fragment of the invention may have a 5’ BgIII restriction enzyme site (5’AGATCT3’) and/or  a 3’ NheI restriction enzyme site (5’GCTAGC3’).    Enhancer  As exemplified herein, the inventors have also  modified the SFTPB promoter fragments of the  invention  to  further  increase  transgene  expression  levels  and/or  to  increase  promoter  activity,  particularly  in  the  lung  parenchyma.  In  particular,  the  inventors  modified  the  SFTPB  promoter  fragment to include an enhancer sequence. Thus, the invention further provides an SFTPB promoter  fragment which further comprises an enhancer.  All disclosure herein to SFTPB promoter fragments  of the invention applies equally and without reservation to SFTPB promoter fragment which further  comprise an enhancer.  An enhancer is a cis‐acting DNA sequence which can increase gene transcription.   An enhancer of the invention may be from about 150 to about 900 bp in length, such as from  about 200 to about 900 bp, from about 200 to about 800 bp, from about 300 to about 700 bp, from  about 400 to about 800 bp, or from about 300 to about 600 bp in length.    An enhancer of the invention may be linked to an SFTPB promoter fragments of the invention  by a linker. Said linker is typically a short DNA sequence, which may be from about 1 to about 50 bp  in length, such as from about 1 to about 20 bp, from about 1 to about 10 bp, from about 5 to about  20 bp in length, or from about 5 to about 10 bp in length.  A linker may be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10  bp, particularly 6 bp, in length.  A non‐limiting example of a linker sequence is given in SEQ ID NO: 51.    Said  linker may  comprise  or  consist  of  one  or more  restriction  enzyme  site,  non‐limiting  examples of which are described herein.  By way of example, a promoter/enhancer combination of  the invention may be joined by a linker which comprises or consists of a NheI restriction site and/or a  BgIII restriction site, preferably a BgIII restriction site.      An enhancer of the invention may comprise additional sequences at the 5’ and/or 3’ end.  By  way of non‐limiting example, one or more restriction enzyme site may be added at the 5’ and/or 3’  end of  the enhancer.   When restriction enzyme sites are present at both  the 5’ and 3’ end of  the  enhancer, each restriction enzyme site may be selected  independently, and thus the 5’ and 3’ sites  may be the same or different.  Said 5’ or 3’ restriction site may comprise all or part of a linker joining  the enhancer to the SFTPB promoter fragment of the invention.  The enhancer may be (i) 5’ to an SFTPB promoter fragment of the invention, or (ii) 3’ to an  SFTPB promoter  fragment of  the  invention. Preferably,  the enhancer  is   5’  to an SFTPB promoter  fragment of the invention.  The (5’ or 3’, preferably 5’) enhancer may be in (i) the forward, or (ii) the reverse orientation.  Preferably, the enhancer is in the forward orientation.  The inventors surprisingly found that the (cell‐specific) expression driven by a SFTPB promoter  fragment further comprising an enhancer sequence as defined herein is higher than gene expression  levels when using an SFTPB promoter  fragment of  the  invention alone. Thus, an SFTPB promoter  fragment  further comprising an enhancer as provided herein has  the potential  to provide an even  greater  increase  in  transgene  expression  compared  with  the  full‐length  SFTPB  promoter  (i.e.  transgene  expression  increases  full‐length  SFTPB  promoter  <  SFTPB  promoter  fragment  of  the  invention < SFTPB promoter of the invention further comprising an enhancer).  Higher expression may  be defined as greater than 105%, greater than 110%, greater than 120%, greater than 130%, greater  than 140%, greater than 150%, greater than 175%, greater than 200%, greater than 225%, greater  than 250%,  greater  than 300%,  greater  than 350%,  greater  than 400% or more of  the  transgene  expression compared with a suitable control, such as the SFTPB promoter fragment alone, or the full‐ length  SFTPB  promoter.  By way  of  non‐limiting  example,  expression  of  a  transgene  by  an  SFTPB  promoter of the invention further comprising an enhancer may be greater than 105%, greater than  110%, greater than 120%, greater than 130%, greater than 140%, greater than 150%, greater than  175%, greater than 200%, greater than 225%, or greater than 250% of the transgene expression when  using the SFTPB promoter fragment alone. Transgene expression may be quantified at the nucleic acid  and/or protein level, and can be quantified by any suitable standard technique known to the person  skilled in the art, for example, by real‐time reverse transcription polymerase chain reaction (RT‐qPCR),  Western blotting and enzyme‐linked immunosorbent assay or ELISA.  The  inventors  also  surprisingly  found  that  the  (cell‐specific) expression driven by  a  SFTPB  promoter  fragment  comprising  an  enhancer  sequence  as  defined  herein  is  comparable  to  gene  expression  levels when  using  ubiquitously  used  strong  promoters  (e.g.  CMV,  hCEF).  Comparable  expression may be defined as at least about 70%, at least about 80%, at least about 90%, at least about  95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% or 100% of gene  expression relative to expression of the same transgene using the same strong promoter (e.g. CMV or  hCEF promoter [such as SEQ  ID NOs: 6 and 5 respectively]).   Expression of the transgene using the  SFTPB promoter fragment may be higher than expression using a strong promoter (e.g. CMV or hCEF)  promoter, for example at least about 105%, at least about 110%, at least about 120%, at least about  150% or more relative to expression of the same transgene using the same strong promoter (e.g. CMV  or hCEF promoter).  Where the SFTPB promoter fragment comprises a SFTPB promoter fragment as defined above  operably  linked  to an enhancer,  said enhancer may preferably be of viral origin, a  lung‐preferred  enhancer, a  lung‐parenchyma‐preferred enhancer, a pneumocyte‐preferred enhancer, an ATII and  club  cell‐preferred enhancer, or an ATII  cell preferred enhancer, a  lung‐specific enhancer, a  lung‐ parenchyma‐specific  enhancer,  a  pneumocyte‐specific  enhancer,  an  ATII  and  club  cell‐specific  enhancer, or an ATII cell specific enhancer. The terms “preferred” and “specific” are defined herein.   As described herein,  the  term  “preferred” may alternatively or additionally be defined as  higher expression in the lung relative to expression in the nose (particularly the nasal cavity), and thus  give rise to an increased ratio of lung expression: nose expression, as described herein.  Expression in  the lung and nasal cavity can be determined using an in vivo luciferase reporter assay (e.g., wherein  vectors comprising nucleic acid cassettes of the invention are administered to mice, and the relative  bioluminescence of the lungs and nasal cavity is quantified).   The  enhancer may be  selected  from  a hB‐actin  enhancer,  a  SLC34A2  (Sodium‐dependent  phosphate transport protein 2B) enhancer, a VEGFA (Vascular Endothelial Growth Factor A) enhancer,  a  CMV  (Cytomegalovirus)  enhancer,  an  SV40  (simian  virus  40)  enhancer,  an  ELF3  (E74  like  ETS  transcription factor 3) enhancer, an SFTPC (surfactant protein C) enhancer, an SFTPB enhancer, or a  LMO7  (LIM domain 7) enhancer. Preferably,  the enhancer  is selected  from a SLC34A2 enhancer, a  VEGFA enhancer, a CMV enhancer, an SV40 enhancer or an ELF3 enhancer. For example,  in some  preferred embodiments, the enhancer is a SLC34A2 enhancer. In other preferred embodiments, the  enhancer is a VEGFA enhancer. In other preferred embodiments, the enhancer is a CMV enhancer. In  other preferred embodiments, the enhancer is a SV40 enhancer. In other preferred embodiments, the  enhancer  is  an  ELF3  enhancer.  Combinations  of  enhancers,  typically  those  identified  herein,  and  combinations of one or more preferred enhancer described herein, may be used according  to  the  present invention.  The CMV enhancer may be in the forwards orientation. Alternatively, the CMV enhancer is in  the reverse orientation. The ELF3 enhancer may be in the forwards orientation. The SV40 enhancer  may be in the forwards orientation. The SV40 enhancer may be in the reverse orientation. The VEGFA  enhancer may be in the forwards orientation. Alternatively, the VEGFA enhancer may be in the reverse  orientation.  The SLC34A2 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%,  at least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ  ID NOs:  29  to  33,  preferably  SEQ  ID NO:  31.  The  SLC34A2  enhancer may  comprise  a  nucleotide  sequence with at least 90% sequence identity to any one of SEQ ID NOs: 29 to 33, preferably SQE ID  NO: 31. The SLC34A2 enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 29  to 33, preferably SEQ ID NO: 31. The SLC34A2 enhancer may consist of a nucleotide sequence with at  least 70%, at  least 75%, at  least 80%, at  least 85%, at  least 90%  identity, or at  least 95% sequence  identity to any one of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31. The SLC34A2 enhancer may  consist of a nucleotide  sequence with  at  least 90%  identity  to any one of  SEQ  ID NOs: 29  to 33,  preferably SEQ ID NO: 31.  The SLC34A2 enhancer may consist of the nucleotide sequence of any one  of SEQ ID NOs: 29 to 33, preferably SEQ ID NO: 31.  The VEGFA enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID  NOs: 36‐40, preferably SEQ ID NO: 38. The VEGFA enhancer may comprise a nucleotide sequence with  at least 90% sequence identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38. The VEGFA  enhancer may comprise the nucleotide sequence of any one of SEQ ID NOs: 36‐40, preferably SEQ ID  NO: 38. The VEGFA enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID  NOs: 36‐40, preferably SEQ ID NO: 38. The VEGFA enhancer may consist of a nucleotide sequence with  at least 90% identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38.  The VEGFA enhancer  may consist of the nucleotide sequence of any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38.  The CMV enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 11 or  12, preferably SEQ ID NO: 11. The CMV enhancer may comprise a nucleotide sequence with at least  90% sequence  identity to SEQ  ID NO: 11 or 12, preferably SEQ  ID NO: 11. The CMV enhancer may  comprise the nucleotide sequence of SEQ ID NO: 5. The CMV enhancer may consist of a nucleotide  sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least  95% sequence  identity to SEQ  ID NO: 11 or 12, preferably SEQ  ID NO: 11. The CMV enhancer may  consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 11 or 12, preferably SEQ ID  NO: 11.  The CMV enhancer may consist of the nucleotide sequence of SEQ ID NO: 11 or 12, preferably  SEQ ID NO: 11.  This exemplary CMV enhancer sequence is CpG‐free.  CpG‐containing variants of this  CMV  enhancer  (or  other  CpG‐comprising  CMV  enhancers)  are  also  encompassed  by  the  present  invention.  The SV40 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 34 or  35. The SV40 enhancer may comprise a nucleotide sequence with at least 90% sequence identity to  SEQ ID NO: 34 or 35. the SV40 enhancer may comprise the nucleotide sequence of SEQ ID NO: 34 or  35. The SV40 enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least  80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 34 or 35. The  SV40 enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 34 or  35.  The SV40 enhancer may consist of the nucleotide sequence of SEQ ID NO: 34 or 35.  The ELF3 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID  NOs: 13 to 20. The ELF3 enhancer may comprise a nucleotide sequence with at least 90% sequence  identity to any one of SEQ ID NOs: 13 to 20. The ELF3 enhancer may comprise the nucleotide sequence  of any one of SEQ ID NOs: 13 to 20. The ELF3 enhancer may consist of a nucleotide sequence with at  least 70%, at  least 75%, at  least 80%, at  least 85%, at  least 90%  identity, or at  least 95% sequence  identity to any one of SEQ ID NOs: 13 to 20. The ELF3 enhancer may consist of a nucleotide sequence  with at least 90% identity to any one of SEQ ID NOs: 13 to 20.  The ELF3 enhancer may consist of the  nucleotide sequence of any one of SEQ ID NOs: 13 to 20.  The actin enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 21 or  22. The actin enhancer may comprise a nucleotide sequence with at least 90% sequence identity to  SEQ ID NO: 21 or 22. The actin enhancer may comprise the nucleotide sequence of SEQ ID NO: 21 or  22. The actin enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least  80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 21 or 22. The  actin enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 21 or  22.  The actin enhancer may consist of the nucleotide sequence of SEQ ID NO: 21 or 22.  The LMO7 enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to any one of SEQ ID  NOs: 23 to 25. The LMO7 enhancer may comprise a nucleotide sequence with at least 90% sequence  identity  to  any  one  of  SEQ  ID NOs:  23  to  25.  The  LMO7  enhancer may  comprise  the  nucleotide  sequence of  any one of  SEQ  ID NOs:  23  to  25.  The  LMO7  enhancer may  consist of  a nucleotide  sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least  95% sequence  identity  to any one of SEQ  ID NOs: 23  to 25. The LMO7 enhancer may consist of a  nucleotide  sequence with  at  least 90%  identity  to  any one of  SEQ  ID NOs:  23  to 25.    The  LMO7  enhancer may consist of the nucleotide sequence of any one of SEQ ID NOs: 23 to 25.  The SFTPC enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 27 or  28. The SFTPC enhancer may comprise a nucleotide sequence with at least 90% sequence identity to  SEQ ID NO: 27 or 28. The SFTPC enhancer may comprise the nucleotide sequence of SEQ ID NO: 27 or  28. The SFTPC enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least  80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 27 or 28. The  SFTPC enhancer may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 27 or  28.  The SFTPC enhancer may consist of the nucleotide sequence of SEQ ID NO: 27 or 28.  The SFTPB enhancer may comprise a nucleotide sequence with at least 70%, at least 75%, at  least 80%, at least 85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 26. The  SFTPB enhancer may comprise a nucleotide sequence with at least 90% sequence identity to SEQ ID  NO: 26. The SFTPB enhancer may comprise the nucleotide sequence of SEQ  ID NO: 26. The SFTPB  enhancer may consist of a nucleotide sequence with at least 70%, at least 75%, at least 80%, at least  85%, at least 90% identity, or at least 95% sequence identity to SEQ ID NO: 26. The SFTPB enhancer  may consist of a nucleotide sequence with at least 90% identity to SEQ ID NO: 26.  The SFTPB enhancer  may consist of the nucleotide sequence of SEQ ID NO: 26.  Any SFTPB promoter fragment of the invention may be combined with any enhancer of the  invention. For the avoidance of doubt, and by way of non‐limiting example, it is envisaged that any   preferred SFTPB promoter fragment of the invention may be combined with any preferred enhancer  (e.g.  an  SLC34A2  enhancer,  a  VEGFA  enhancer,  a  CMV  enhancer,  an  SV40  enhancer  or  an  ELF3  enhancer).   A preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise  or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least  90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity,  or more to SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT, identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  76).      Said  preferred  SFTPB  promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid  sequence with at least 90% identity to SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the  BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 76).   Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a  nucleic acid  sequence of SEQ  ID NO: 41, or  the  sequence of SEQ  ID NO: 41 wherein  the BgIII  site  (AGATCT,  identified  in  the  sequence  information  section herein)  is omitted  (SEQ  ID NO: 76). Said  preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic  acid sequence of SEQ ID NO: 41, or the sequence of SEQ ID NO: 41 wherein the BgIII site (AGATCT,  identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  76). A  particularly  preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist  of a nucleic acid sequence with at  least 70%, at  least 75%, at  least 80%, at  least 85%, at  least 90%  identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or  more to SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC, identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  82).      Said  preferred  SFTPB  promoter fragment further comprising an SLC34A2 enhancer may comprise or consist of a nucleic acid  sequence with at least 90% identity to SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the  Nhe1 site (GCTAGC, identified in the sequence information section herein) is omitted (SEQ ID NO: 82).   Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise a  nucleic acid sequence of SEQ  ID NO: 47, or the sequence of SEQ  ID NO: 47 wherein  the Nhe1 site  (GCTAGC,  identified  in  the  sequence  information  section herein)  is omitted  (SEQ  ID NO: 82). Said  preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may consist of a nucleic  acid sequence of SEQ ID NO: 47, or the sequence of SEQ ID NO: 47 wherein the Nhe1 site (GCTAGC,  identified  in  the  sequence  information  section herein)  is omitted  (SEQ  ID NO: 82).   A particularly  preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise or consist  of a nucleic acid sequence with at  least 70%, at  least 75%, at  least 80%, at  least 85%, at  least 90%  identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or  more to SEQ  ID NO: 47.     Said preferred SFTPB promoter fragment further comprising an SLC34A2  enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO:  47.  Said preferred SFTPB promoter fragment further comprising an SLC34A2 enhancer may comprise  a nucleic acid sequence of SEQ ID NO: 47. Said preferred SFTPB promoter fragment further comprising  an SLC34A2 enhancer may consist of a nucleic acid sequence of SEQ ID NO: 47.  A preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or  consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least  90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity,  or more to SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT, identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  77).      Said  preferred  SFTPB  promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid  sequence with at least 90% identity to SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the  BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 77).   Said  preferred  SFTPB  promoter  fragment  further  comprising  a  VEGFA  enhancer may  comprise  a  nucleic acid  sequence of SEQ  ID NO: 42, or  the  sequence of SEQ  ID NO: 42 wherein  the BgIII  site  (AGATCT,  identified  in  the  sequence  information  section herein)  is omitted  (SEQ  ID NO: 77). Said  preferred SFTPB promoter fragment further comprising a VEGFA enhancer may consist of a nucleic  acid sequence of SEQ ID NO: 42, or the sequence of SEQ ID NO: 42 wherein the BgIII site (AGATCT,  identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  77). A  particularly  preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or consist of  a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity,  or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to  SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the  sequence  information section herein)  is omitted (SEQ  ID NO: 83).     Said preferred SFTPB promoter  fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid sequence  with at least 90% identity to SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site  (GCTAGC,  identified  in  the sequence  information section herein)  is omitted  (SEQ  ID NO: 83).   Said  preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise a nucleic  acid sequence of SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC,  identified in the sequence information section herein) is omitted (SEQ ID NO: 83). Said preferred SFTPB  promoter fragment further comprising a VEGFA enhancer may consist of a nucleic acid sequence of  SEQ ID NO: 48, or the sequence of SEQ ID NO: 48 wherein the Nhe1 site (GCTAGC, identified in the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  83). A  particularly  preferred  SFTPB  promoter fragment further comprising a VEGFA enhancer may comprise or consist of a nucleic acid  sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least  95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, or more to SEQ ID NO:  48.   Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may comprise or  consist of a nucleic acid sequence with at least 90% identity to SEQ ID NO: 48.  Said preferred SFTPB  promoter fragment further comprising a VEGFA enhancer may comprise a nucleic acid sequence of  SEQ ID NO: 48. Said preferred SFTPB promoter fragment further comprising a VEGFA enhancer may  consist of a nucleic acid sequence of SEQ ID NO: 48.  A preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or  consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least  90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity,  or more to SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  78).      Said  preferred  SFTPB  promoter  fragment  further comprising a CMV enhancer may comprise or consist of a nucleic acid  sequence with at least 90% identity to SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the  BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 78).   Said preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise a nucleic  acid sequence of SEQ ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT,  identified in the sequence information section herein) is omitted (SEQ ID NO: 78). Said preferred SFTPB  promoter fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ  ID NO: 43, or the sequence of SEQ ID NO: 43 wherein the BgIII site (AGATCT, identified in the sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  78). A  particularly  preferred  SFTPB  promoter  fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with  at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least  96%, at  least 97%, at  least 98%, at  least 99% sequence  identity, or more to SEQ  ID NO: 46, or the  sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified in the sequence information  section  herein)  is  omitted  (SEQ  ID  NO:  81).      Said  preferred  SFTPB  promoter  fragment  further  comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with at  least 90%  identity to SEQ ID NO: 46, or the sequence of SEQ ID NO: 46 wherein the Nhe1 site (GCTAGC, identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID  NO:  81).    Said  preferred  SFTPB  promoter fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ  ID NO:  46,  or  the  sequence  of  SEQ  ID NO:  46 wherein  the Nhe1  site  (GCTAGC,  identified  in  the  sequence  information  section herein)  is omitted  (SEQ  ID NO: 81). Said preferred SFTPB promoter  fragment further comprising a CMV enhancer may consist of a nucleic acid sequence of SEQ ID NO:  46, or  the sequence of SEQ  ID NO: 46 wherein  the Nhe1 site  (GCTAGC,  identified  in  the sequence  information  section herein)  is omitted  (SEQ  ID NO: 81).   A particularly preferred SFTPB promoter  fragment further comprising a CMV enhancer may comprise or consist of a nucleic acid sequence with  at least 70%, at least 75%, at least 80%, at least 85%, at least 90% identity, or at least 95%, at least  96%, at  least 97%, at  least 98%, at  least 99% sequence  identity, or more  to SEQ  ID NO: 46.     Said  preferred SFTPB promoter fragment further comprising a CMV enhancer may comprise or consist of  a nucleic acid sequence with at least 90% identity to SEQ ID NO: 46.  Said preferred SFTPB promoter  fragment further comprising a CMV enhancer may comprise a nucleic acid sequence of SEQ ID NO: 46.  Said preferred SFTPB promoter fragment further comprising a CMV enhancer may consist of a nucleic  acid sequence of SEQ ID NO: 46.  A preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise or  consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least  90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity,  or more to SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT, identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  80).      Said  preferred  SFTPB  promoter fragment further comprising an SV40 enhancer may comprise or consist of a nucleic acid  sequence with at least 90% identity to SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the  BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 80).   Said preferred SFTPB promoter fragment further comprising an SV40 enhancer may comprise a nucleic  acid sequence of SEQ ID NO: 45, or the sequence of SEQ ID NO: 45 wherein the BgIII site (AGATCT,  identified in the sequence information section herein) is omitted (SEQ ID NO: 80). Said preferred SFTPB  promoter fragment further comprising an SV40 enhancer may consist of a nucleic acid sequence of  SEQ  ID NO: 45, or the sequence of SEQ  ID NO: 45 wherein the BgIII site (AGATCT,  identified  in the  sequence information section herein) is omitted (SEQ ID NO: 80).  A preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise or  consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least  90% identity, or at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity,  or more to SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT, identified  in  the  sequence  information  section  herein)  is  omitted  (SEQ  ID NO:  79).      Said  preferred  SFTPB  promoter fragment further comprising an ELF3 enhancer may comprise or consist of a nucleic acid  sequence with at least 90% identity to SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the  BgIII site (AGATCT, identified in the sequence information section herein) is omitted (SEQ ID NO: 79).   Said preferred SFTPB promoter fragment further comprising an ELF3 enhancer may comprise a nucleic  acid sequence of SEQ ID NO: 44, or the sequence of SEQ ID NO: 44 wherein the BgIII site (AGATCT,  identified in the sequence information section herein) is omitted (SEQ ID NO: 79). Said preferred SFTPB  promoter fragment further comprising an ELF3 enhancer may consist of a nucleic acid sequence of  SEQ  ID NO: 44, or the sequence of SEQ  ID NO: 44 wherein the BgIII site (AGATCT,  identified  in the  sequence information section herein) is omitted (SEQ ID NO: 79).  Thus, a preferred SFTPB promoter fragment further comprising an enhancer may comprise or  consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least 85%, at least  90% identity, at least 95%, at least 96%, at least 97%,a t least 98%, at least 99% identity or more to  any  one  of  SEQ  ID NOs:  41  to  48.  A  preferred  SFTPB  promoter  fragment  further  comprising  an  enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity to any one of  SEQ  ID NOs: 41  to 48. A preferred SFTPB promoter  fragment  further comprising an enhancer may  comprise a nucleic acid sequence of any one of SEQ ID NOs: 41 to 48. A preferred SFTPB promoter  fragment further comprising an enhancer may consist of a nucleic acid sequence of any one of SEQ ID  NOs: 41 to 48. A particularly preferred SFTPB promoter fragment further comprising an enhancer may  comprise or consist of a nucleic acid sequence with at least 70%, at least 75%, at least 80%, at least  85%, at least 90% identity, at least 95%, at least 96%, at least 97%,a t least 98%, at least 99% identity  or more to any one of SEQ ID NOs: 46 to 48. A particularly preferred SFTPB promoter fragment further  comprising an enhancer may comprise or consist of a nucleic acid sequence with at least 90% identity  to  any  one  of  SEQ  ID NOs:  46  to  48.  A  particularly  preferred  SFTPB  promoter  fragment  further  comprising an enhancer may comprise a nucleic acid sequence of any one of SEQ ID NOs: 46 to 48. A  particularly preferred SFTPB promoter  fragment  further  comprising an enhancer may  consist of a  nucleic acid sequence of any one of SEQ ID NOs: 46 to 48.    Lung parenchyma and specific/preferred expression therein  The respiratory system can be divided into airways and lung parenchyma. The airways consist  of the bronchus, which bifurcates off the trachea and divides into bronchioles and then further into  alveoli. The parenchyma is responsible for gas exchange and includes the alveoli, alveolar ducts, and  terminal and respiratory bronchioles. The most prominent structure  in the  lung parenchyma  is the  alveolus.  Two  types of  epithelial  cell  line  the  alveolus. Alveolar  type  I  (ATI)  cells  exhibit  a broad,  flattened morphology and cover around 95% of the surface area, whilst the cuboidal alveolar type II  cells (ATII cells) line the remainder of the alveolus. ATI cells provide a gas exchange interface with the  underlying endothelium, whereas ATII cells serve as both progenitors of ATI cells and also play a critical  role  in maintaining  the homeostasis of  the alveolus. The  latter  role  is  fulfilled by  the  secretion of  surfactant proteins from specialised organelles within ATII cells, so‐called ‘lamellar bodies’, into the  alveolar space.   Secretion of surfactant proteins maintain surface tension and prevents atelectasis at the end  of  expiration, whilst  contributing  to  the  varied  functions  of  the ATII  cells. ATII  cells  are  the  only  epithelial cell of the lung which synthesise and release all four surfactant proteins A, B, C and D, with  surfactant protein C being unique to the ATII cell. In addition to synthesising, storing and secreting  surfactant components, ATII cells have the following functions: (1) the transepithelial movement of  water and ions regulating the volume of the alveolar surface liquid (ASL) preventing alveoli flooding,  (2) the expression of immunomodulatory proteins necessary for host defence and the regulation of  innate immunity and (3) the regeneration of alveolar epithelium after injury.   In addition, surfactant proteins A, B and D are also synthesised by club cells (previously named  Clara Cells) founds in the terminal and respiratory bronchioles of humans. Club cells are non‐ciliated  epithelial cells found mainly in bronchioles as well as basal cells found in large airways. They have been  ascribed several protective roles, including airway repair after injury, secretion of anti‐inflammatory  and immunomodulatory proteins, and detoxification.  ATI dysfunction, ATII dysfunction and/or club cell dysfunction or dropout  is associated with  the pathogenesis of various parenchymal  lung diseases. Accordingly, the  lung parenchyma may be  targeted for treating genetic diseases such as surfactant deficiencies and interstitial lung disease.      The promoters of the invention drive transgene expression in the lung parenchyma, typically  one or more cell type of the  lung parenchyma.    In particular, the promoters of the  invention drive  transgene expression in (i) ATI cells, (ii) ATII cells, (iii) club cells, (iv) ATI cells and ATII cells, (v) ATI cells  and club cells, (vi) ATII cells and club cells, or (vii) ATI cells, ATII cells and club cells.  A  SFTPB  promoter  fragment  of  the  invention  is  functional.    As  used  herein,  the  term  “functional” may mean that an SFTPB promoter fragment of the invention retains the functionality of  the full‐length SFTPB promoter.  In other words, a promoter of the invention may be defined as being  capable of expressing a gene of interest (e.g., a transgene) in a cell type which, in a healthy subject,  would  express  SFTPB.  For  example,  the  an    SFTPB promoter  fragment may be used  to  express  a  transgene in ATII cells, club cells and/or ATI cells.   Additionally, or alternatively, an  SFTPB promoter fragment of the invention may be defined  as  functional  as  it  is  capable of preferentially or  specifically  expressing  a  gene of  interest  (e.g.,  a  transgene) in a tissue‐type or cell‐type which, in a healthy subject, would express SP‐B. Thus, tan SFTPB  promoter  fragment  of  the  invention may  be  a  lung‐parenchyma  preferred  promoter.  An  SFTPB  promoter fragment of the invention may be a lung‐parenchyma specific promoter. An SFTPB promoter  fragment  of  the  invention may  be  a  pneumocyte‐preferred  promoter, whereby  a  pneumocyte  is  defined as any of the specialized cells of the alveoli of the lungs.  An SFTPB promoter fragment of the  invention may be a pneumocyte‐specific promoter. An SFTPB promoter fragment of the invention may  be  an  ATII  cell‐preferred  promoter,  a  club  cell‐preferred  promoter  and/or  an  ATI  cell‐preferred  promoter. An SFTPB promoter fragment of the invention may be an ATII cell‐specific promoter, a club  cell‐specific expression and/or an ATI cell‐specific promoter.   An SFTPB promoter  fragment of  the  invention may preferably be an ATII cell‐preferred promoter. An SFTPB promoter  fragment of  the  invention may preferably be an ATII cell‐specific promoter.    Tissue or cell preferred expression may be defined as expression that is higher in said tissue  or cell than other tissue or cell types. For example,  lung‐parenchyma preferred expression may be  defined as expression that  is significantly higher  in the  lung‐parenchyma (or one or more cell type  therein, as described above)  than expression  in one or more of:  the brain,  the eye,  the endocrine  tissues, the proximal digestive tract, the gastrointestinal tract,  liver and gall bladder, pancreas, the  kidney  and/or  urinary  bladder, male  tissues  (i.e.,  the  testis,  epididymis,  prostate  and/or  seminal  vesicle) and female tissues  (i.e., the vagina, breast, cervix, endometrium, fallopian tube, ovary and  placenta), muscle tissues, connective and soft tissues, skin, bone marrow and lymphoid tissue. Lung‐ parenchyma preferred expression may be defined as expression that is at least about 5 times greater,  at least about 10 times greater, at least about 20 times greater, at least about 50 times greater or at  least  about  100  or more  times  greater  in  the  lung  parenchyma  than  one  or more  of  the  above  reference tissue types, especially the gastrointestinal tract and/or the brain.   Tissue or cell specific expression may be defined as expression that is at least about 10 times  greater, at least about 20 times greater, at least about 50 times greater or at least about 100 or more  times greater in said tissue or cell than any other tissue or cell types.   Preferential and/or specific expression may be assessed at the  level of RNA and/or protein  expression, preferably protein expression.   Expression  can be measured by  any  suitable  standard  technique known to the person skilled in the art. For example, RNA expression levels can be measured  by  quantitative  real‐time  PCR.  Protein  expression  can  be  measured  by  western  blotting  or  immunohistochemistry.   Advantageously,  restricting  the  expression  of  transgenes  to  cells  expressing  endogenous  SFTPB is expected to reduce the effects of off‐target gene expression, overexpression (e.g., toxicity/ER  stress/UPR, etc) and/or reduce immune responses. As such, the SFTPB promoters of the invention, as  a consequence of their preferential and/or specific expression in the lung parenchyma, or one or more  cell type thereof, have potential clinical benefits as a result of these advantageous properties.  By way  of non‐limiting example, compared with a hCEF promoter of SEQ ID NO: 5 and/or a CMV promoter of  SEQ ID NO: 6, an SFTPB promoter fragment of the invention may reduce the effects of off‐target gene  expression,  overexpression,  and/or  reduce  immune  responses.  Additionally,  or  alternatively,  compared to a hCEF promoter of SEQ  ID NO: 5and/or a CMV promoter of SEQ  ID NO: 6, an SFTPB  promoter fragment of the invention may preferentially or specifically express a transgene in the lung  parenchyma, pneumocytes,  such as ATII cells, ATI cells, and/or club cells.     As  described  and  exemplified  herein,  an  SFTPB  promoter  fragment  of  the  invention may  increase transgene expression compared with a full‐length SFTPB promoter, such as that described  herein.  Typically, an SFTPB promoter fragment of the invention increases transgene expression in the  lung parenchyma  (or one or more  lung parenchymal cell or cell type selected  from ATII cell(s), ATI  cell(s)  and/or  club  cell(s)).    An  SFTPB  promoter  fragment  of  the  invention  increases  transgene  expression in lung parenchyma (or one or more lung parenchymal cell or cell type selected from ATII  cell(s), ATI cell(s) and/or club cell(s)) by at least about 2‐fold,  at least about 2.5‐fold, at least about 3‐ fold, at  least about 4‐fold, at  least about 5‐fold, at  least about 7.5‐fold or more.   The  increase    in  transgene expression by an SFTPB promoter fragment of the invention may be quantified compared  with a suitable control, preferably compared with a full‐length SFTPB promoter, such as that described  herein.  Thus, an SFTPB promoter fragment of the invention increases transgene expression in the lung  parenchyma (or one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s)  and/or club cell(s)) compared with expression of the same transgene in the lung parenchyma (or one  or more  lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s))  compared with a  full‐length SFTPB promoter,  such as  that described herein.   An SFTPB promoter  fragment of the invention increases transgene expression in lung parenchyma (or one or more lung  parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about  2‐fold,  at least about 2.5‐fold, at least about 3‐fold, at least about 4‐fold, at least about 5‐fold, at least  about 7.5‐fold or more compared with expression of the same transgene in the lung parenchyma  (or  one or more lung parenchymal cell or cell type selected from ATII cell(s), ATI cell(s) and/or club cell(s))  compared with  a  full‐length  SFTPB  promoter,  such  as  that  described  herein.    In  some  preferred  embodiments, the SFTPB promoter fragment of the invention increases transgene expression in ATII  cells, and optionally one or more additional lung parenchymal cell type as described herein.  Again,  this increase in expression is preferably compared with expression of the same transgene in the same  cell type(s) by a full‐length SFTPB promoter, such as that described herein.  Alternatively  or  additionally,  as  described  and  exemplified  herein,  an  SFTPB  promoter  fragment of the invention may preferentially drive transgene expression in the lung (particularly the  lung parenchyma or one or more cell type thereof) compared with transgene expression in the nose  (or cells thereof).   Thus, an SFTPB promoter  fragment of the  invention may preferentially  increase  transgene expression  in  the  lung parenchyma  (or one or more  lung parenchymal  cell or  cell  type  selected from ATII cell(s), ATI cell(s) and/or club cell(s)) compared with transgene expression in the  nose  (or cells  thereof).   An SFTPB promoter  fragment of  the  invention may preferentially  increase  transgene expression  in  the  lung parenchyma  (or one or more  lung parenchymal  cell or  cell  type  selected from ATII cell(s), ATI cell(s) and/or club cell(s)) by at least about 2‐fold,  at least about 2.5‐ fold, at least about 3‐fold, at least about 4‐fold, at least about 5‐fold, at least about 7.5‐fold or more  compared with transgene expression in the nose (or cells thereof).  The ratio of lung expression: nose  expression by an SFTPB promoter fragment of the invention may be at least about 2:1, such as at least  about 2.5:1, at least about 3:1, at least about 4:1, at least about 5:1 or more.   Any  disclosure  herein  in  relation  to  promoters  (i.e.  SFTPB  promoter  fragments)  of  the  invention applies equally and without reservation to other aspects of the  invention comprising, or  relating to, said promoters.  Thus, any disclosure herein in relation to promoters (i.e. SFTPB promoter  fragments)  of  the  invention  applies  equally  and  without  reservation  to  promoter/enhancer  combinations, viral and non‐viral vectors of  the  invention.   By way of non‐limiting example, as  for  promoters (i.e. SFTPB promoter fragments) of the invention, promoter/enhancer combinations, viral  and non‐viral vectors of the  invention drive transgene expression  in the  lung parenchyma, typically  one or more cell  type of  the  lung parenchyma.    In particular,  the promoters  (i.e. SFTPB promoter  fragments) and promoter/enhancer combinations, viral and non‐viral vectors of the invention drive  transgene expression in (i) ATI cells, (ii) ATII cells, (iii) club cells, (iv) ATI cells and ATII cells, (v) ATI cells  and club cells, (vi) ATII cells and club cells, or (vii) ATI cells, ATII cells and club cells.    Nucleic Acid Cassettes  The  invention  also  provides  a  nucleic  acid  cassette.  In  particular,  the  present  invention  provides a nucleic acid cassette comprising (a) an SFTPB promoter fragment; and (b) a transgene.   As  defined  herein,  a  transgene  may  be  defined  as  anucleic  acid  sequence  encoding  a  therapeutic protein.Thus, the terms “nucleic acid sequence encoding a therapeutic protein” and the  term “transgene” may be used interchangeably. In a nucleic acid of the invention, an SFTPB promoter  fragment may increase expression of the transgene by the lung parenchyma (e.g. ATII cells), as defined  herein.   The  increase  in expression of a transgene by an SFTPB promoter fragment of the  invention  may  be  as  defined  herein,  including  disclosure  of  increasing  transgene  expression  using  SFTPB  promoter  fragments and/or SFTPB promoter  fragments  combined with an enhancer, as described  above. Accordingly,  any  disclosure  herein  in  relation  to  increasing  transgene  expression  using  an  SFTPB  promoter  fragment  and/or  SFTPB  promoter  fragment  combined with  an  enhancer  of  the  invention  applies  equally  and without  reservation  to  nucleic  acid  cassettes  of  the  invention.  For  example, the increase in expression of a transgene by an SFTPB promoter fragment of the invention  is an  increase of at  least about 50%, at  least about 60%, at  least about 70%, at  least about 80% or  more. Preferably, the  increase  in expression of a transgene by an SFTPB promoter fragment of the  invention is an increase of at least about 50%.  In a nucleic acid of the invention, an SFTPB promoter fragment of the invention may increase  expression of the transgene by the lung parenchyma (e.g. ATII cells) relative to a suitable control  For  example, the SFTPB promoter fragment of the invention may increase expression of the transgene by  the lung parenchyma (e.g. ATII cells) relative to expression using the full‐length SFTPB gene promoter  (e.g., the 972 bp genomic fragment defined above).   Alternatively or  in addition, expression of the transgene by the  lung parenchyma  (e.g. ATII  cells) using the SFTPB promoter fragment in a nucleic acid of the invention, may be  comparable to  the expression of the transgene using a ubiquitously used strong promoter (e.g. CMV or hCEF). For  example, the expression of the transgene using the SFTPB promoter fragment may be at least 70%, at  least  75%,  at  least  80%,  at  least  85%,  at  least  90%,  at  least  95% of  the  expression using  a CMV  promoter, as described herein. Alternatively, expression of the transgene using the SFTPB promoter  fragment in a nucleic acid of the invention, may be higher than expression using a CMV promoter, for  example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more  relative  to expression of  the  same  transgene using a CMV promoter. Alternatively or  in addition,  expression of the transgene using the SFTPB promoter fragment in a nucleic acid of the invention may  be at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95% of the expression  using  a hCEF promoter. Alternatively,  the  expression of  the  transgene using  the  SFTPB promoter  fragment in a nucleic acid of the invention may be higher than expression using a hCEF promoter, for  example at least about 105%, at least about 110%, at least about 120%, at least about 150% or more  relative to expression of the same transgene using a CMV promoter.  A nucleic acid cassette or vector of  the  invention enables  long‐term  transgene expression,  resulting in  long‐term expression of a transgene, which offers clinical benefits for the expression of  therapeutic proteins in patients. As described herein, the phrases “long‐term expression”, “sustained  expression”, “long‐lasting expression” and “persistent expression” are used  interchangeably. Long‐ term expression according to the present  invention means expression of a transgene, preferably at  therapeutic levels, for at least 45 days, at least 60 days, at least 90 days, at least 120 days, at least 180  days, at least 250 days, at least 360 days, at least 450 days, at least 730 days or more. Preferably long‐ term expression means expression for at least 90 days, at least 120 days, at least 180 days, at least  250 days, at least 360 days, at least 450 days, at least 720 days or more, more preferably at least 360  days, at least 450 days, at least 720 days or more. The long‐term expression is typically accompanied  by long‐term secretion or long‐term membrane insertion of the (therapeutic) protein encoded by the  transgene, depending on whether  the  (therapeutic) protein  is a secreted protein  (e.g. SFTPB) or a  membrane protein.   Long‐term secretion according to the present  invention means secretion of a (therapeutic)  protein encoded by a transgene of the invention, preferably at therapeutic levels, for at least 45 days,  at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360  days, at least 450 days, at least 730 days or more. Preferably long‐term secretion means secretion for  at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days, at least 450  days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at least 720 days  or more.   Long‐term membrane insertion according to the present invention means that a (therapeutic)  protein encoded by a transgene of the invention is inserted into and present in the cell membrane,  preferably at therapeutic  levels, for at  least 45 days, at  least 60 days, at  least 90 days, at  least 120  days, at  least 180 days, at  least 250 days, at  least 360 days, at  least 450 days, at  least 730 days or  more. Preferably long‐term expression means membrane insertion for at least 90 days, at least 120  days, at  least 180 days, at  least 250 days, at  least 360 days, at  least 450 days, at  least 720 days or  more, more preferably at least 360 days, at least 450 days, at least 720 days or more.   In particular, a nucleic acid cassette or vector of  the  invention may drive  (increased)  long‐ lasting  expression  of  a  transgene,  and/or  (preferably  and)  expression  and  secretion/membrane  insertion of a (therapeutic) protein encoded by said transgene  in one or more cell type of the  lung  parenchyma, as described herein, in vivo in a patient. Preferably, a nucleic acid cassette or vector of  the  invention  drives  expression  of  a  transgene,  and/or  (preferably  and)  expression  and  secretion/membrane insertion of a (therapeutic) protein encoded by a transgene of the invention in  one or more cell / cell type of the lung parenchyma, as described herein, for at least 45 days, more  preferably at least 90 days.   The  nucleic  acid  of  the  nucleic  acid  cassette may  be  as  defined  herein.  The  nucleic  acid  cassette comprise DNA or RNA. Preferably the nucleic acid cassette is DNA.   A nucleic acid cassette of the invention may optionally be codon optimised for expression in  a particular cell type, for example, eukaryotic cells (e.g. mammalian cells, yeast cells,  insect cells or  plants cells) or prokaryotic cells (e.g. E.coli). The term “codon optimised” refers to the replacement of  at least one codon within a base polynucleotide sequence with a codon that is preferentially used by  the host organism in which the polynucleotide is to be expressed. Typically, the most frequently used  codons in the host organism are used in the codon‐optimised polynucleotide sequence. Methods of  codon optimisation are well known in the art.     It will be understood by a skilled person that numerous different polynucleotides can encode  the same polypeptide as a result of the degeneracy of the genetic code.  It  is also understood that  skilled persons may, using routine techniques, make nucleotide substitutions that do not affect the  polypeptide  sequence  encoded  by  the  nucleic  acid molecules  to  reflect  the  codon  usage  of  any  particular host organism in which the polypeptides are to be expressed. Therefore, unless otherwise  specified, a nucleic acid  cassette  that  comprises or  consists of  the SFTPB promoter  fragment and  transgene (e.g., encoding a therapeutic protein) of the invention includes all polynucleotide sequences  that are degenerate versions of each other and that encode the same amino acid sequence.      A nucleic acid cassette of the  invention preferably comprises an SFTPB promoter fragment  comprising  an  SFTPB  promoter  fragment  operably  linked  to  an  enhancer  sequence,  as  described  herein.  A  nucleic  acid  cassette  of  the  invention  typically  comprises  a  SFTPB  promoter  fragment  comprising a  SFTPB promoter  fragment and an enhancer  sequence, wherein  the  SFTPB promoter  fragment is operably linked to a nucleic acid sequence comprising or consisting of a transgene (e.g.,  encoding a therapeutic protein). By operably linked, it is meant that the SFTPB promoter fragment is  configured to express the transgene (e.g., encoding the therapeutic protein). The transgene encoding  the  therapeutic protein may also be  linked  to a suitable  terminator sequence. Suitable  terminator  sequences are well known in the art.  The promoter  included  in  the nucleic  acid  cassettes  and  vectors of  the  invention may be  specifically  selected and/or modified  to  further  refine  regulation of expression of  the  therapeutic  gene. Again, suitable promoters and standard techniques for their modification are known in the art.  As a non‐limiting example, an SFTPB promoter fragment of the invention may be modified to reduce  the number of CpG dinucleotides, or to render the SFTPB promoter fragment CpG‐free.  A number of  (CpG‐free) promoters, and methods for the generation of CpG‐free promoters which are suitable for  use in the present invention are described in Pringle et al. (J. Mol. Med. Berl. 2012, 90(12): 1487‐96),  which is herein incorporated by reference in its entirety.   Preferably, the nucleic acid cassettes and  vectors of the invention comprise  an SFTPB promoter fragment having low or no CpG dinucleotide  content. Low CpG dinucleotide content may be defined as 10 CpG dinucleotides or less, preferably 5  CpG dinucleotides or less, such as 5, 4, 3, 2 or 1 CpG dinucleotides. An SFTPB promoter fragment may  have some or all CG dinucleotides replaced with any one of AG, TG or GT. The absence (or reduction)  of CpG dinucleotides further improves the performance of some nucleic acid cassettes and vectors of  the invention, particularly  lentiviral (e.g. SIV) vectors of the invention and in particular in situations  where  it  is  not  desired  to  induce  an  immune  response  against  an  expressed  antigen  or  an  inflammatory response against the delivered expression construct. The elimination or reduction of  CpG dinucleotides reduces the occurrence of flu‐like symptoms and inflammation which may result  from administration of constructs, particularly when administered to the airways.   The nucleic acid cassettes and vectors of the invention may be modified to allow shut down  of gene expression. Standard techniques for modifying the vector in this way are known in the art. As  a non‐limiting example, Tet‐responsive promoters are widely used.  The nucleic acid cassette of the invention (or a vector comprising said cassette) may have an  intron positioned between the promoter and the transgene. Non‐limiting examples of suitable introns  are found for example, in UK Application No. 2213936.4, which is herein incorporated by reference in  its entirety. Wherein the nucleic acid of the invention is present in a non‐viral vector (e.g. plasmid),  the presence of at least one intron between the SFTPB promoter fragment and the transgene may be  preferred, for example an intron as described in UK Application No. 2213936.4.    The nucleic acid cassettes and vectors of  the  invention may  include at  least one part of a  vector,  in particular,  regulatory elements. By way of non‐limiting example,  the promoter within a  nucleic acid cassette of the invention may be used to express more than one polypeptide, including  one or more therapeutic protein. Thus, the nucleic acid cassette may comprise a nucleic acid sequence  which, when  transcribed, gives rise to multiple polypeptides,  for  instance a  transcript may contain  multiple open reading frames  (ORFs) and also one or more  Internal Ribosome Entry Sites  (IRES) to  allow translation of ORFs after the first ORF. A transcript may be polycistronic, i.e. it may be translated  to give a polypeptide which is subsequently cleaved to give a plurality of polypeptides. Alternatively,  a nucleic acid cassette of the invention may comprise multiple promoters, including multiple SFTPB  promoter fragments of the invention, or multiple copies of any specific an SFTPB promoter fragment  of the invention, and hence give rise to a plurality of transcripts and hence a plurality of polypeptides,  including a plurality of  therapeutic proteins. Nucleic acid cassettes may,  for  instance, express one,  two,  three,  four or more polypeptides via a promoter or promoters,  including one or more SFTPB  promoter fragment of the invention.  A  nucleic  acid  cassette may  comprise  one  or more  translation  initiation  sequence  (TIS).  Translation  initiation  plays  an  important  role  in mRNA  translation,  canonically  a methionyl  tRNA  unique  for  initiation  (Met‐tRNAi)  identifies  the  AUG  start  codon  and  triggers  the  downstream  translation process. Non‐canonical start codons (e.g. CUG for valyl‐tRNA)/TIS may also be used.  The nucleic acid cassettes of the present  invention may comprise at  least one termination  signal. A “termination signal" or "terminator" is comprised of the DNA sequences involved in specific  termination of an RNA  transcript by an RNA polymerase. Thus, a  termination signal  that ends  the  production of an RNA transcript is contemplated according to the present invention. A terminator may  be necessary in vivo to achieve desirable message levels. In eukaryotic systems, a terminator region  may also comprise specific DNA sequences that permit site‐specific cleavage of the new transcript so  as to expose a polyadenylation site. This signals a specialized endogenous polymerase to add a stretch  of about 200 A residues (polyA) to the 3’ end of the transcript. RNA molecules modified with this polyA  tail appear to more stable and are translated more efficiently. Thus, when the nucleic acid cassette is  for expression  in eukaryotes, a terminator typically comprises a signal for the cleavage of the RNA,  and  it  is  preferred  that  the  terminator  signal  promotes  polyadenylation  of  the  message.  The  terminator  and/or  polyadenylation  site  elements  can  serve  to  enhance  message  levels  and  to  minimize read through from the cassette into other sequences.  Terminators  contemplated  for  use  in  the  invention  include  any  known  terminator  of  transcription described herein or known to one of ordinary skill in the art, including but not limited to,  for example, the termination sequences of genes, such as for example the bovine growth hormone  terminator  or  viral  termination  sequences,  such  as  for  example  the  SV40  terminator.  In  certain  embodiments, the termination signal may be a lack of transcribable or translatable sequence, such as  due to a sequence truncation.  The  invention also provides gene therapy vectors comprising a nucleic acid cassette of the  invention. Any and all disclosure herein in relation to nucleic acid cassettes of the invention applies  equally and without reservation to gene therapy vectors of the invention.  The nucleic acid cassettes and vectors of the invention are capable of expressing the transgene  in a given host cell. Any appropriate host cell may be used, such as mammalian, bacterial, insect, yeast,  and/or plant host cells. In addition, cell‐free expression systems may be used. Such expression systems  and host cells are standard in the art.  Typically the nucleic acid cassettes and vectors of the invention are capable of expressing the  transgene  in  the  lung. The nucleic acid  cassettes and vectors of  the  invention may be  capable of  expressing  the  transgene  in  the  lung  parenchyma.  The  nucleic  acid  cassettes  and  vectors  of  the  invention may be capable of expressing the transgene in one or more cell type selected from ATII cells,  ATI cells, club cells, and/or bronchioalveolar stem cells. The nucleic acid cassettes and vectors of the  invention are typically capable of expressing the transgene  in one or more ATII cells, ATI cells, club  cells, bronchioalveolar stem cells  in the terminal bronchioles. Preferably, the nucleic acid cassettes  and vectors of the invention are capable of expressing the transgene in ATII cells.  Thus, the nucleic  acid  cassettes  and  vectors  of  the  invention  are  capable  of  expressing  the  (therapeutic)  protein  encoded by the transgene in one or more cell/cell type of the lung parenchyma, particularly ATII cells,  ATI cells, club cells and/or bronchioalveolar stem cells in the terminal bronchioles, particularly in ATII  cells.  The nucleic acid cassettes of the invention may be made using any suitable process known in  the  art.  Thus,  the  nucleic  acid  cassettes  may  be  made  using  chemical  synthesis  techniques.  Alternatively,  the  nucleic  acid  cassettes  of  the  invention may  be made  using molecular  biology  techniques.    Non‐Viral Vectors    The  present  invention  also  provides  a  vector  comprising  a  nucleic  acid  cassette  of  the  invention. The vector(s) may be present in the form of a therapeutic composition or formulation. The  vector may be a non‐viral vector.  The non‐viral vector(s) may be a DNA vector, such as a DNA plasmid. The vector(s) may be an  RNA vector, such as a mRNA vector or a self‐amplifying RNA vector. The non‐viral vector may be an  exosomes or microvesicle (MV).  The non‐viral (e.g. DNA and/or RNA) vector(s) of the invention may  be capable of expression in eukaryotic and/or prokaryotic cells.   Typically, the non‐viral (e.g. DNA and/or RNA) vector(s) are capable of expression in a cell of  a subject, for example, a cell of a mammalian or avian subject to be immunised.   Typically the nucleic acid cassettes and vectors of the invention are capable of expressing a  transgene in airway cells, preferably lung parenchymal cells  (as described herein).    A non‐viral vector of  the present  invention may be a phage vector, such as an AAV/phage  hybrid vector as described  in Hajitou et al., Cell 2006; 125(2) pp. 385‐398; herein  incorporated by  reference.    Vector(s) of  the present  invention  (e.g. non‐viral DNA or RNA vectors) may be designed  in  silico, and then synthesised by conventional polynucleotide synthesis techniques.  Non‐viral plasmids cannot replicate in the subject to be treated, as they lack the viral genetic  material  which  hijacks  the  body's  normal  production  machinery.  However  they  are  capable  of  replicating in appropriate host cells, such as yeasts or bacteria including E. coli, and particularly airway  cells as defined herein.    The  term "plasmid" as used herein  refers  to a construction comprised of genetic material  designed  to direct  transformation of a  targeted  cell. The plasmid  contains a plasmid backbone. A  "plasmid backbone" as used herein contains multiple genetic elements positionally and sequentially  oriented with other necessary genetic elements such that the nucleic acid in the nucleic acid cassette  can be transcribed and when necessary translated in the transfected cells.   The plasmid backbone can contain one or more unique restriction sites within the backbone.  The plasmid may be capable of autonomous replication in a defined host or organism such that the  cloned sequence  is reproduced. The plasmid can confer some well‐defined phenotype on the host  organism which is either selectable or readily detected. The plasmid or plasmid backbone may have a  linear or circular configuration. The components of a plasmid can contain, but is not limited to, a DNA  molecule incorporating: (1) the plasmid backbone; (2) a sequence comprising or consisting of an SFTPB  promoter  fragment;  (3) a  transgene sequence encoding a  (therapeutic) protein; and optionally  (4)  additional regulatory elements for transcription, translation, RNA stability and replication.  The purpose of the plasmid in human gene therapy for the efficient delivery of nucleic acid  sequences to, and expression of therapeutic proteins in, a cell or tissue. In particular, the purpose of  the plasmid is to achieve high copy number, avoid potential causes of plasmid instability and provide  a means  for plasmid selection. As  for expression,  the nucleic acid cassette contains  the necessary  elements  for  expression  of  the  nucleic  acid within  the  cassette.  Expression  includes  the  efficient  transcription of an inserted gene, nucleic acid sequence, or nucleic acid cassette with the plasmid.     A DNA plasmid may be CpG‐free, or be optimised to reduce CpG dinucleotides as described  herein. A DNA plasmid of the invention may be codon‐optimised as described herein.  Methods of preparing plasmid DNA are well known in the art. Typically, they are capable of  autonomous replication in an appropriate host or producer cell.  The term "exosome" as used herein refers to an extracellular vesicle formed by exocytosis  from a cell of origin.  An exosome typically comprises a nucleic cassette of the invention.  Exosomes  may be used to transform a targeted cell. The nucleic acid within an exosome may contain one or  more genetic elements, which may be positionally and sequentially oriented with other necessary  genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed and when  necessary translated in the transformed cells.   The term "microvesicle" (MV) as used herein refers to an extracellular vesicle formed typically  between about 30 to about 1,000 nm in diameter.  An MV typically comprises a nucleic cassette of the  invention.  Exosomes may be used to transform a targeted cell. The nucleic acid within an MV may  contain one or more genetic elements, which may be positionally and sequentially oriented with other  necessary genetic elements such that the nucleic acid in the nucleic acid cassette can be transcribed  and when necessary translated in the transformed cells.   Host cells containing (e.g. transformed, transfected, or electroporated with) the plasmid may  be  prokaryotic  or  eukaryotic  in  nature,  either  stably  or  transiently  transformed,  transfected,  or  electroporated with the plasmid. Suitable host cells include bacterial, yeast, fungal, invertebrate, and  mammalian cells. Preferably the host cell is bacterial; more preferably E. coli.  Host cells can then be used in methods for the large scale production of the plasmid. The cells  are grown in a suitable culture medium under favourable conditions, and the desired plasmid isolated  from the cells, or from the medium in which the cells are grown, by any purification technique well  known to those skilled in the art; e.g. see Sambrook et al, supra.  Any appropriate delivery means can be used to deliver a non‐viral vector (e.g. plasmid) of the  invention  to a  target cell or patient. Suitable delivery means are known  in  the art and within  the  routine skill of one of ordinary skill in the art. Non‐limiting examples include the use of cationic lipids,  polymers (e.g. polyethyleneimine and poly‐L‐lysine) and electroporation.   Preferably  cationic  lipids may  be  used  to  deliver  non‐viral  (e.g.  plasmid)  vectors  of  the  invention  to  target  cells  or  to  a  patient. Non‐limiting  examples  of  cationic  lipids  suitable  for  use  according to the invention are GL67A and lipofectamine.  The cationic lipid mixture GL67A is a mixture of three components ‐ GL67 (Cholest‐5‐en‐3‐ol  (3β)‐,3‐[(3‐aminopropyl)[4‐[(3‐ aminopropyl)amino]butyl]carbamate],  (CAS Number: 179075‐30‐0)),  DOPE  (1,2‐dioleoyl‐sn‐glycero‐3‐phosphoethanolamine)  and  DMPE‐PEG5000  (1,2‐Dimyristoyl‐sn‐ Glycero‐3‐Phosphoethanolamine‐N‐[methoxy  (Polyethylene  glycol)5000]).  These  components  are  formulated at a 1:2:0.05 molar ratio to form GL67A. The composition of GL67A and methods for its  production are disclosed in WO2013/061091, as are methods for preparing mixtures of GL67A with  exemplary non‐viral vectors. The contents of WO2013/061091 are herein incorporated by reference  in their entirety.  Lipofectamine consists of a 3:1 mixture of DOSPA (2,3‐dioleoyloxy‐N‐  [2(sperminecarboxamido)ethyl]‐N,N‐dimethyl‐1‐propaniminium trifluoroacetate) and DOPE.    Viral vectors  The  present  invention  also  provides  a  vector  comprising  a  nucleic  acid  cassette  of  the  invention. The vector(s) may be present in the form of a therapeutic composition or formulation. The  vector may be a viral vector.  Any appropriate viral vector may be used to deliver a nucleic acid cassette of the invention.  By way of non‐limiting example, a viral vector of the invention may be a lentiviral vector, an adeno‐ associated virus (AAV) vector, an adenoviral vector, a poxvirus vector, a herpes simplex virus (HSV)  vector.  Derivatives of these viral vectors, such as lentivirus‐derived particles are also encompassed  within the invention.  Non‐limiting examples of adenoviral vectors include human serotypes such as AdHu5, simian  serotypes such as ChAd63, ChAdOX1 or ChAdOX2, and other forms. Non‐limiting examples of poxvirus  vectors  include  a  modified  vaccinia  Ankara  (MVA)).  ChAdOX1  and  ChAdOX2  are  disclosed  in  WO2012/172277 (herein incorporated by reference in its entirety). ChAdOX2 is a BAC‐derived and E4  modified AdC68‐based viral vector.     Viral vectors are usually non‐replicating or replication impaired vectors, which means that the  viral vector cannot  replicate  to any significant extent  in normal cells  (e.g. normal human cells), as  measured by conventional means – e.g. via measuring DNA synthesis and/or viral titre. Non‐replicating  or replication  impaired vectors may have become so naturally (i.e. they have been  isolated as such  from nature) or artificially (e.g. by breeding in vitro or by genetic manipulation). There will generally  be at least one cell‐type in which the replication‐impaired viral vector can be grown – for example,  modified vaccinia Ankara (MVA) can be grown in CEF cells.     Typically, the viral vector is incapable of causing a significant infection in an animal subject,  typically in a mammalian subject such as a human or other primate.    Viral vector(s) of  the present  invention may be designed  in silico, and  then synthesised by  conventional polynucleotide synthesis techniques.  Preferably the invention relates to retroviral vectors, particularly lentiviral vectors. The term  “lentivirus”  refers  to  a  family  of  retroviruses.  Retroviral/lentiviral  vectors  of  the  invention,  can  integrate  into  the  genome  of  transduced  cells  and  lead  to  long‐lasting  expression.  Examples  of  retroviruses  suitable  for  use  in  the present  invention  include  gammaretroviruses  such  as murine  leukaemia virus (MLV) and feline leukaemia virus (FLV). Examples of lentiviruses suitable for use in the  present invention include Simian immunodeficiency virus (SIV), Human immunodeficiency virus (HIV),  Feline immunodeficiency virus (FIV), Equine infectious anaemia virus (EIAV), and Visna/maedi virus. A  particularly preferred lentiviral vector is an SIV vector (including all strains and subtypes), such as a  SIV‐AGM (originally isolated from African green monkeys, Cercopithecus aethiops).   The retroviral/lentiviral (e.g. SIV) vectors of the present invention are typically pseudotyped  with hemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, or  with G glycoprotein from Vesicular Stomatitis Virus (G‐VSV). Preferably the lentiviral (e.g. SIV) vectors  of  the  present  invention  are  pseudotyped  with  HN  and  F  from  a  respiratory  paramyxovirus.  Particularly preferably the respiratory paramyxovirus is a Sendai virus (murine parainfluenza virus type  1).  The F protein may be a truncated F protein, typically one in which the cytoplasmic domain is  truncated.  Preferably the truncated F protein is Fct4, in which 38 amino acids have been truncated  from the C‐terminus of the F protein, with 4 amino acids of the F protein cytoplasmic domain being  retained. The F protein may be a truncated F protein, typically one in which the cytoplasmic domain  is truncated.  Preferably the truncated F protein is Fct4, in which 38 amino acids have been truncated  from the C‐terminus of the F protein, with 4 amino acids of the F protein cytoplasmic domain being  retained.  A retroviral/lentiviral (e.g. SIV) vector for use according to the  invention may be  integrase‐ competent (IC). Alternatively, the lentiviral (e.g. SIV) vector may be integrase‐deficient (ID).   Viral vectors of the invention, particularly retroviral/lentiviral (e.g. SIV) vectors as described  herein may  transduce one or more cells types as described herein to achieve  long term  transgene  expression.    The HN protein may be a truncated and/or chimeric HN protein, typically one  in which the  cytoplasmic domain is truncated or substituted.  Preferably, the HN protein is a chimeric HN protein  in  which  (i)  the  cytoplasmic  domain  of  the  HN  is  replaced  by  the  cytoplasmic  domain  of  the  transmembrane (TMP) protein; or (ii) the cytoplasmic domain of the TMP is added to the cytoplasmic  domain of the HN protein.  The HN protein may be as described in Kobayashi et al. (J. Virol. (2003)  77(4):2607‐2614), which is herein incorporated by reference in its entirety.  Particularly preferred truncated and/or chimeric forms of the F and HN proteins are described  in UK Patent Application No. 2212472.1, which is herein incorporated by reference in its entirety.    The viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the  present  invention enable high  levels of transgene expression. Together with the  increased  levels of  expression/secretion/membrane insertion of a therapeutic protein resulting from the use of an SFTPB  promoter  fragment of  the  invention,  these viral vectors  typically  result  in high  levels  (therapeutic  levels) of expression of the transgene, and  the (therapeutic) protein encoded by said transgene.   The  transgene  to  be  included  in  a  viral  vector  of  the  invention,  particularly  a  retroviral/lentiviral  (e.g. SIV) vector of  the  invention may be modified  to  facilitate expression. For  example, the transgene sequence may be in CpG‐depleted /low (or CpG‐fee) and/or codon‐optimised  form to facilitate gene expression. Standard techniques for modifying the transgene sequence in this  way are known in the art.  The viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the  invention  exhibit  enhanced  expression  of  the  transgene.  Accordingly,  the  viral  vectors  of  the  invention,  particularly  the  retroviral/lentiviral  (e.g.  SIV)  vectors  of  the  invention  are  capable  of  producing long‐lasting, repeatable, high‐level transgene expression, particularly in lung parenchyma  without inducing side effects (e.g., an undue immune response).   The viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the  present  invention  enable  long‐term  transgene  expression,  resulting  in  long‐term  expression  (and  secretion  or membrane  insertion)  of  a  (therapeutic)  protein  by  cells  of  the  lung  parenchyma  as  described herein. Long‐term expression according  to  the present  invention means expression of a  transgene gene and/or encoded (therapeutic) protein, preferably at therapeutic levels, for at least 45  days, at least 60 days, at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least  360  days,  at  least  450  days,  at  least  730  days  or more.  Preferably  long‐term  expression means  expression for at least 90 days, at least 120 days, at least 180 days, at least 250 days, at least 360 days,  at least 450 days, at least 720 days or more, more preferably at least 360 days, at least 450 days, at  least 720 days or more.   Preferably, the invention relates to the use of F/HN lentiviral vectors comprising a nucleic acid  cassette of the invention, particularly SIV F/HN vectors.   The  nucleic  acid  cassette  comprised  in  a  viral  vector  of  the  invention,  particularly  a  retroviral/lentiviral  (e.g. SIV) vector of  the  invention, may have no  intron positioned between  the  promoter  and  the  nucleic  acid  encoding  the  signal  peptide  and/or  the  nucleic  acid  encoding  the  therapeutic protein. Similarly,  there may be no  intron between  the promoter and  the nucleic acid  encoding the signal peptide and/or the nucleic acid encoding the therapeutic protein  in the vector  genome (pDNA1) plasmid (for example, pGM326 or pGM830 as illustrated in Figures 2A and B and the  corresponding sequences in UK Application No. 2102832.9, which is herein incorporated by reference  in its entirety).  The viral vectors of the invention may be made using any suitable process known in the art.  In particular, retroviral/lentiviral (e.g. SIV) vectors of the invention may be made using the methods  disclosed in UK Application No. 2102832.9, which is herein incorporated by reference in its entirety).  The viral vectors of the invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the  invention may  comprise  a  central  polypurine  tract  (cPPT)  and/or  the Woodchuck  hepatitis  virus  posttranscriptional regulatory elements (WPRE). An exemplary WPRE sequence is provided by SEQ ID  NO: 52.    Transgenes    A  nucleic  acid  cassette  of  the  invention  comprises  a  transgene.  Typically,  the  transgene  encodes    a  therapeutic  protein.  A  therapeutic  protein  is  one  which  has  potential  utility  in  the  treatment or prevention of a disease or condition, such as those describe herein. Thus, a nucleic acid  cassette of  the  invention may comprise a nucleic acid encoding a protein which has a  therapeutic  effect on a disease or condition to be treated.     A nucleic acid cassette of the invention may comprise a nucleic acid encoding a therapeutic  protein which is a functional or wild‐type form of a protein which is present in a patient to be treated  in a dysfunctional form (whether the dysfunction is inherent or acquired).  As used herein, the phrase  "inherent dysfunction" refers to a protein which is innately dysfunctional due to genetic factors and  the phrase "acquired dysfunction" refers to a protein which is dysfunctional due to environmental or  other factors after birth.     Thus, a nucleic acid cassette of the invention may comprise a nucleic acid encoding a protein  which is a functional or wild‐type form of a protein which is present in a patient, but which that has  become dysfunctional due to a genetic disease, such as a genetic respiratory disease.  The nucleic acid cassettes of the present invention are useful in the treatment of diseases via  their use  in expressing  therapeutic proteins  in  lung parenchyma as described herein  (e.g. ATI, ATII  cells). Preferably, the nucleic acid cassettes of the present  invention are useful  in the treatment of  diseases via their use in expressing therapeutic proteins in ATII cells, ATI cells and/or club cells.   The transgene may encode a therapeutic protein selected from:  (a) a secreted therapeutic  protein,  optionally  Surfactant  Protein  B  (SFTPB),  Surfactant  Protein  C  (SFTPC),  alpha‐1‐antitrypsin  (AAT),  Factor  VIII,  Factor  VII,  Factor  IX,  Factor  X,  Factor  XI,  von Willebrand  Factor,  Granulocyte‐ Macrophage Colony‐Stimulating Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or  TGGβ) or monoclonal antibody, an anti‐inflammatory decoy and a monoclonal antibody against an  infectious agent; or (b) ATP‐binding cassette sub‐family member A (ABCA3), TRIM72, CSF2RA, CSF2RB  and decorin.   In  some  embodiments,  the  therapeutic  protein  is  not  an  antibody,  particularly  not  a  monoclonal antibody. In such embodiments, the transgene may encode a therapeutic protein selected  from: (a) a secreted therapeutic protein, optionally SFTPB, SP‐C, AAT, Factor VIII, Factor VII, Factor IX,  Factor X, Factor XI, von Willebrand Factor, GM‐CSF, an anti‐inflammatory protein (e.g. IL‐10, or TGGβ)   and an anti‐inflammatory decoy; or (b) ABCA3, TRIM72, CSF2RA, CSF2RB and decorin.   The nucleic acid cassettes of the invention are particularly efficient at driving the expression,  secretion and/or membrane  insertion of proteins (e.g. therapeutic proteins as described herein) by  the  lung parenchyma. This  is particularly the case when such cassettes are comprised within F/HN  pseudotyped viral vectors of the invention (as described herein), which are efficient at targeting cells  in the lung parenchyma.  As such, for therapeutic applications the nucleic acid cassettes of the invention (and vectors  comprising said cassettes) are typically delivered to cells of the respiratory tract, particularly the cells  of  the  lung parenchyma.  In other words,  the nucleic acid  cassettes of  the  invention  (and  vectors  comprising said cassettes) are typically delivered to lung parenchyma as described herein. Accordingly,  the nucleic acid cassettes of  the  invention  (and vectors comprising  said cassettes) are particularly  suited  for  treatment  of  diseases  or  disorders  of  the  lung  parenchyma.  Typically,  the  nucleic  acid  cassettes of the invention (and vectors comprising said cassettes) may be used for the treatment of a  genetic respiratory disease.   A nucleic acid cassette of the invention (or vector comprising said cassette) may comprise a  nucleic acid encoding a polypeptide or protein that is therapeutic for the treatment of such diseases,  particularly a disease or disorder of the lung parenchyma.  The transgene and therapeutic protein of  the  invention are not  limited, one of ordinary skill  in the art will be able  to  identify transgenes an  therapeutic proteins which may be usefully delivered according to the invention, particularly in the  context of genetic diseases, particularly genetic respiratory diseases and diseases or disorders of the  lung parenchyma as those described herein.  Accordingly, a nucleic acid cassette of the invention (or vector comprising said cassette) may  comprise  a  nucleic  acid  sequence  encoding  a  therapeutic  protein  selected  from:  (a)  a  secreted  therapeutic protein, optionally SFTPB, SFTPC, AAT, Factor VIII,  Factor VII, Factor IX, Factor X, Factor  XI, von Willebrand Factor, GM‐CSF, an anti‐inflammatory protein (e.g. IL‐10, TGGβ, or TNF‐alpha) or  monoclonal antibody, an anti‐inflammatory decoy and a monoclonal antibody against an  infectious  agent;  or  (b) ABCA3,  TRIM72,  CSF2RA,  CSF2RB  or DCN. Other  preferred  examples  of  therapeutic  proteins that may be encoded by a nucleic acid sequence comprised in a nucleic acid cassette of the  invention  (or  vector  comprising  said  cassette)  include  genes  related  to  or  associated with  other  surfactant deficiencies.    The  therapeutic  protein  encoded  by  a  nucleic  acid  cassette  of  the  invention  (or  vector  comprising said cassette) may be SFTPB. Examples of an SFTPB therapeutic transgene are provided by  SEQ ID NOs: 53, 54 and 56. An exemplary codon‐optimised SFTPB transgene is provided by SEQ ID NO:  55.   The therapeutic protein encoded by said SFTPB transgene, may be exemplified by the polypeptide  of SEQ ID NO: 57. Variants thereof (as described therein) are also included, particularly variants with  at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 53 to 57.   The transgene may encode ABCA3. Examples of a ABACA3 transgene are provided by SEQ ID  NOs: 58 and 59. An exemplary codon‐optimised ABACA3 transgene is provided by SEQ ID NO: 60. The  polypeptide encoded by said ABACA3 transgene, may be exemplified by the polypeptide of SEQ ID NO:  61. Variants thereof (as described therein) are also included,  particularly variants with at least 90%  (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NOs: 58 to 61.  The  therapeutic  protein  encoded  by  a  nucleic  acid  cassette  of  the  invention  (or  vector  comprising said cassette) may be SFTPC. An example of an SFTPC therapeutic transgene is provided  by SEQ ID NO: 62. The therapeutic protein encoded by said SFTPC transgene, may be exemplified by  the polypeptide of SEQ ID NO: 63. Variants thereof (as described therein) are also included, particularly  variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID  NO: 62 or 63.   The  therapeutic  protein  encoded  by  a  nucleic  acid  cassette  of  the  invention  (or  vector  comprising said cassette) may be an AAT. An example of an AAT therapeutic transgene (SERPINA1) is  provided  by  SEQ  ID  NO:  70.  SEQ  ID  NO:  70  is  a  codon‐optimized  CpG  depleted  AAT  transgene  (SERPINA1) previously designed by the present inventors to enhance translation in human cells. Such  optimisation has been shown to enhance gene expression by up to 15‐fold. Variants of same sequence  (as defined herein) which possess the same technical effect of enhancing translation compared with  the unmodified (wild‐type) AAT gene sequence are also encompassed by the present invention. The  therapeutic protein encoded by said AAT transgene, may be exemplified by the polypeptide of SEQ ID  NO: 71. Variants thereof (as described therein) are also  included, particularly variants with at  least  90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any one of SEQ ID NO: 70 or 71.  The  therapeutic  protein  encoded  by  a  nucleic  acid  cassette  of  the  invention  (or  vector  comprising said cassette) may be an FVIII. Examples of a FVIII therapeutic transgene are provided by  SEQ ID NOs: 72 and 73. The polypeptide encoded by the FVIII transgene, may be exemplified by the  polypeptide  of  SEQ  ID NO:  74  and  75.  Variants  thereof  (as  described  therein)  are  also  included,  particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any  one of SEQ ID NOs: 72 to 75.  The  therapeutic  protein  encoded  by  a  nucleic  acid  cassette  of  the  invention  (or  vector  comprising said cassette) may be GM‐CSF. A GM‐CSF transgene may comprise or consist of SEQ ID NO:  64  (human).  The  polypeptide  encoded  by  the  GM‐CSF  transgene  may  be  exemplified  by  the  polypeptide of SEQ  ID NO: 65  (human). Variants  thereof  (as described  therein) are also  included,  particularly variants with at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to any  one of SEQ ID NOs: 64and 65.   The transgene may encode decorin. An example of a DCN transgene is provided by SEQ ID NO:  66.  The polypeptide encoded by said DCN transgene, may be exemplified by the polypeptide of SEQ  ID NO: 67. Variants thereof (as described therein) are also included,  particularly variants with at least  90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to SEQ ID NO: 66or 67.  The transgene may encode TRIM72. An example of a TRIM72 transgene is provided by SEQ ID  NO: 68. The polypeptide encoded by said TRIM72 transgene, may be exemplified by the polypeptide  of SEQ ID NO: 69. Variants thereof (as described therein) are also included,  particularly variants with  at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98, 99 or 100% to SEQ ID NO: 68or 69.  The  therapeutic  protein  encoded  by  a  nucleic  acid  cassette  of  the  invention  (or  vector  comprising said cassette) may be encoded by any one of SFTPB, SFTPC, Factor V, Factor VII, Factor IX,  Factor X and/or Factor XI, von Willebrand Factor, GM‐CSF, ABCA3, TRIM72 or DCN, or other known  related gene.   As the lung parenchyma is preferably targeted for delivery of the nucleic acid cassettes of the  invention  (and vectors comprising  said cassettes),  the  transgene may preferably be SFTPB, SFTPC,  ABCA3 or GM‐CSF. The therapeutic protein may be a monoclonal antibody (mAb) against an infectious  agent (bacterial, fungal or viral, e.g. the SARS‐Co‐V2 virus). The therapeutic protein may be anti‐TNF  alpha.  The  therapeutic protein may be one  implicated  in  an  inflammatory,  immune or metabolic  condition.   A nucleic acid cassette of the invention (or a vector comprising said cassette) may be delivered  to one or more cell/cell type of the lung parenchyma to allow production of proteins to be secreted  into circulatory system. In such embodiments, the therapeutic protein may be any one of Factor VII,  Factor VIII, Factor IX, Factor X, Factor XI and/or von Willebrand’s factor. Such a nucleic acid cassette  of  the  invention  (or a vector comprising  said cassette) may be used  in  the  treatment of diseases,  particularly cardiovascular diseases and blood disorders, preferably blood clotting deficiencies such as  haemophilia. Again, the therapeutic protein may be an mAb against an infectious agent or a protein  implicated in an inflammatory, immune or metabolic condition, such as, lysosomal storage disease.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may have no  intron  positioned  between  the  promoter  and  the  nucleic  acid  encoding  the  therapeutic  protein.  Similarly, when the nucleic acid cassette is comprised in a viral vector, there may be no intron between  the promoter and the transgene in the vector genome (pDNA1) plasmid used to make said viral vector,  as described herein. Alternatively, said nucleic acid cassette of the invention (or a vector comprising  said cassette) may have an intron positioned between the promoter and the transgene, particularly if  the cassette (or vector comprising said cassette) is non‐viral.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a SFTPB promoter fragment and an SFTPB transgene, including those described herein. Optionally said  nucleic  acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a SFTPB promoter fragment and an SFTPC transgene, including those described herein. Optionally said  nucleic  acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a  SFTPB promoter  fragment  and  an AAT  transgene  (SERPINA1),  including  those described herein.  Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have  no intron positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a SFTPB promoter fragment and an FVIII transgene, including those described herein. Optionally said  nucleic  acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a SFTPB promoter fragment and an FVII transgene, including those described herein. Optionally said  nucleic  acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a SFTPB promoter fragment and an FIX transgene, including those described herein. Optionally said  nucleic  acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a SFTPB promoter fragment and an FX transgene,  including those described herein. Optionally said  nucleic  acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a SFTPB promoter fragment and an FXI transgene, including those described herein. Optionally said  nucleic  acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a SFTPB promoter fragment and a von Willebrand Factor transgene, including those described herein.  Optionally said nucleic acid cassette of the invention (or a vector comprising said cassette) may have  no intron positioned between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) may comprise  a  SFTPB  promoter  fragment  and  a  Granulocyte‐Macrophage  Colony‐Stimulating  Factor  (GM‐CSF)  transgene, including those described herein. Optionally said nucleic acid cassette of the invention (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned  between  the  SFTPB  promoter  fragment and the transgene.   In some preferred embodiments, the lentiviral (e.g. SIV) vector comprises a SFTPB promoter  fragment  and  an  DCN  transgene,  including  those  described  herein.   Optionally  said  nucleic  acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned  between the SFTPB promoter fragment and the transgene.   In some preferred embodiments, the lentiviral (e.g. SIV) vector comprises a SFTPB promoter  fragment and a TRIM72  transgene,  including  those described herein.   Optionally  said nucleic acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned  between the SFTPB promoter fragment and the transgene.   In some preferred embodiments, the lentiviral (e.g. SIV) vector comprises a SFTPB promoter  fragment and a ABACA3  transgene,  including  those described herein.   Optionally  said nucleic acid  cassette  of  the  invention  (or  a  vector  comprising  said  cassette) may  have  no  intron  positioned  between the SFTPB promoter fragment and the transgene.   The nucleic acid cassette of the invention (or a vector comprising said cassette) comprises a  nucleic acid encoding a therapeutic protein (said nucleic acid is referred to interchangeably herein as  a  transgene).  The  nucleic  acid  sequence  encodes  a  gene  product,  e.g.,  a  protein,  particularly  a  therapeutic protein.   For example, the nucleic acid cassette of the invention (or a vector comprising said cassette)  comprises a nucleic acid sequence encoding an SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI,  von Willebrand Factor transgene, decorin, TRIM72 or ABAC3 and said nucleic acid sequence comprises  (or consists of) a nucleic acid sequence having at least 90% (such as at least 90, 92, 94, 95, 96, 97, 98,  99  or  100%)  sequence  identity  to  the  SFTPB,  SFTPC,  AAT, GM‐CSF,  FVIII,  FVII,  FVIX,  FX,  FXI,  von  Willebrand Factor transgene, decorin, TRIM72 or ABAC3 nucleic acid sequence respectively, examples  of which are described herein. The nucleic acid sequence encoding SFTPB, SFTPC, AAT, GM‐CSF, FVIII,  FVII,  FVIX,  FX,  FXI,  von  Willebrand  Factor  transgene,  decorin,  TRIM72  or  ABAC3  may  preferably  comprise (or consist of) a nucleic acid sequence having at least 95% (such as at least 95, 96, 97, 98, 99  or 100%) sequence identity to the SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand  Factor transgene, decorin, TRIM72 or ABAC3 nucleic acid sequence respectively, examples of which  are described herein.   The amino acid sequence of the (therapeutic) protein encoded by the transgene may be a  functional variant having at least 95% (such as at least 95, 96, 97, 98, 99 or 100%) sequence identity  to the functional protein. For example, an SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von  Willebrand  Factor  transgene,  decorin,  TRIM72  or  ABAC3  polypeptide  encoded  by  the  respective  SFTPB, SFTPC, AAT, GM‐CSF, FVIII, FVII, FVIX, FX, FXI, von Willebrand Factor transgene, decorin, TRIM72  or ABAC3 transgene may comprise (or consist of) an amino acid sequence having at least 95% (such  as at least 95, 96, 97, 98, 99 or 100%) sequence identity to the functional SFTPB, SFTPC, AAT, GM‐CSF,  FVIII,  FVII, FVIX, FX, FXI,  von Willebrand Factor  transgene, decorin, TRIM72 or ABAC3 polypeptide  sequence respectively.    An SFTPB promoter fragment and/or enhancer of the invention may be linked to a transgene  by a  linker. Said  linker  is  typically a short DNA sequence, as defined herein, and may comprise or  consist of one or more restriction enzyme site, examples of which are also described herein.  By way  of example, a SFTPB promoter fragment of the  invention may be  joined to a transgene by a  linker  which comprises or consists of a NheI restriction site and/or a BgIII restriction site, preferably a NheI  restriction site.      Signal peptides  The  transgene  encoding  for  a  (therapeutic)  protein may  further  comprise  a  nucleic  acid  sequence encoding for a signal peptide.  Said signal peptide may be the endogenous signal peptide of  the  (therapeutic) protein, or a  signal peptide exogenous  to  said  (therapeutic) protein.   When  the  transgene further comprises a nucleic acid encoding for an exogenous signal peptide, the transgene  preferably exclude a nucleic acid  sequence encoding  for  the endogenous  signal peptide.      In  such  instances,  the exogenous  signal peptide  is  typically  the sole signal peptide  linked with  (and hence  driving secretion and/or membrane insertion) of the therapeutic protein.    All  disclosure  herein  relates  to  both  transgenes  and  therapeutic  proteins  including  and  excluding  endogenous  signal  peptides  unless  explicitly  stated.    By way  of  non‐limiting  example,  sequence  identity of variants, and/or  lengths of fragments may be based on the sequence with or  without a signal peptide.     Any signal peptide and therapeutic combination may be used, provided that this combination  is  effective  in  increasing  the  expression,  secretion  and/or membrane  insertion  of  a  (therapeutic)  protein as defined herein.   Selection of a signal peptide may depend on  the specific  (therapeutic)  protein and/or  the specific  lung parenchymal cell  type by which  the  (therapeutic) protein  is  to be  expressed/secreted/inserted into the cell membrane.    Exogenous signal peptides which increase expression, secretion and/or membrane insertion  of a (therapeutic) protein have been described by the  inventors  in International Patent Application  No: WO2022/219333, which is herein incorporated by reference in its entirety.    Co‐administration of lentiviral vectors and surfactants  As  exemplified  herein,  lentiviral  vectors,  such  as  the  lentiviral  vector  pseudotyped  with  haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus, can be  administered in combination with a synthetic surfactant, and achieve at least as efficient transduction  into  target  cells, and potentially even enhanced  transduction,  compared with  transduction of  the  lentiviral  vector  in  vehicle alone.   This  is  surprising as  lentiviral  vectors are  surrounded by a  lipid  envelope, which would be predicted  to be disrupted by  the hydrophobic portions of a surfactant,  having a negative effect on the lentiviral structure.  Whilst the combination of a lentiviral vector and  a surfactant is exemplified herein with a specific lentiviral vector, SIV.F/HN with a luciferase transgene,  for  the  avoidance  of  doubt  this  example  provides  proof  of  concept  for  other  lentiviral  vectors,  particularly other SIV.F/HN vectors to be co‐administered with a surfactant in this way.  This is because  it is the nature of the lentiviral structure, i.e. the presence of an envelope that is determinative.  Having  shown that co‐administration is feasible with one lentiviral vector, a skilled person would understand  that co‐administration could be carried out with any lentiviral vector and a surfactant.  Similarly, the  nature  of  any  pseudotyping  proteins  will  not  affect  the  ability  of  a  lentiviral  vector  to  be  co‐ administered with a surfactant.   The  invention  therefore  provides  lentiviral  vector  pseudotyped  with  haemagglutinin‐ neuraminidase  (HN)  and  fusion  (F)  proteins  from  a  respiratory  paramyxovirus  and  comprising  a  transgene operably  linked  to  a promoter  for use  in  a method of  treating  a  disease, wherein  the  lentiviral vector is administered in combination with a surfactant.  Thus,  the  invention  therefore provides  lentiviral vector pseudotyped with haemagglutinin‐ neuraminidase  (HN)  and  fusion  (F)  proteins  from  a  respiratory  paramyxovirus  and  comprising  a  transgene operably  linked  to  a promoter  for use  in  a method of  treating  a  disease, wherein  the  lentiviral vector is administered simultaneously or sequentially with a surfactant.  The lentiviral (e.g. SIV) vector and surfactant are administered in combination.  Administered  "in  combination,"  encompasses  both  simultaneous  (also  referred  to  as  concurrent)  administration/delivery and sequential (also referred to as separate) administration/delivery.  For  , "simultaneous" or "concurrent delivery”, the delivery of the  lentiviral (e.g. SIV) vector  may still be occurring when the delivery of the surfactant begins, or the delivery of surfactant may still  be occurring when  the delivery of the  lentiviral  (e.g. SIV) vector begins, so  that there  is overlap  in  terms of administration. Simultaneous delivery may encompass delivery of  the  lentiviral  (e.g. SIV)  vector  and  surfactant within weeks  to months or  even  years of  each other,  typically  so  that  the  lentiviral (e.g. SIV) vector delivery overlaps with the delivery of the surfactant.   Alternatively, the delivery of the lentiviral (e.g. SIV) vector may end before the delivery of the  surfactant begins, or the delivery of the surfactant may end before delivery of the lentiviral (e.g. SIV)  vector begins. Sequential administration may  involve  the  lentiviral  (e.g. SIV) vector and surfactant  being administered within 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 6 hours, 12 hours or 24 hours  or longer of each other.   The  lentiviral  vector may  be  administered  before  the  surfactant.    The  surfactant may  be  administered before the lentiviral vector.  The lentiviral vector and surfactant may be administered  simultaneously,  optionally  wherein  the  lentiviral  vector  and  surfactant  are  mixed  prior  to  administration.  Typically the treatment is more effective because of combined administration. For example,  treatment with the lentiviral (e.g. SIV) vector may be more effective, e.g., an equivalent effect is seen  with  less of the  lentiviral (e.g. SIV) vector, or the  lentiviral (e.g. SIV) vector reduces symptoms to a  greater extent, than would be seen if the lentiviral (e.g. SIV) vector were administered in the absence  of the surfactant.  By way of further example, treatment with the surfactant may be more effective,  e.g., an equivalent effect is seen with less of the surfactant, or the surfactant reduces symptoms to a  greater extent, than would be seen if the surfactant were administered in the absence of the lentiviral  (e.g. SIV) vector.    Typically, delivery is such that the reduction in a symptom, or other parameter related to a  disease to be treated  is at  least equivalent to what would be observed with the  lentiviral (e.g. SIV)  vector  delivered  in  the  absence  of  the  surfactant,  or  the  analogous  situation  is  seen  with  the  surfactant.   Alternatively, a combination therapy of the invention may increase transgene expression by  at least 1.2 fold, at least 1.3 fold, at least 1.4 fold, at least 1.5 fold, at least 2 fold, at least 2.5 fold or  more  compared with  treatment with  the  lentiviral  (e.g. SIV) vector alone  (i.e.  compared with  the  increase in transgene expression achieved when treating with the lentiviral (e.g. SIV) alone).    Without being bound by theory, it is believed that the surfactant may aid the distribution of  the  lentiviral  (e.g.  SIV)  vector  within  the  lungs,  enabling  it  to  penetrate  more  deeply  into  the  respiratory tree and thus facilitating transduction of the lung parenchyma.  It will be appreciated  that appropriate dosage of  the  lentiviral  (e.g. SIV) vector and/or  the  surfactant, will depend on the specific agent, and can also vary from patient to patient.  Any surfactant may be used in a combination therapy according to the present invention.  It  will be appreciated that it is within the routine practice of one of ordinary skill in the art to select such  a  surfactant.    By way  of  non‐limiting  example,  two  animal‐derived  surfactants  in  clinical  use  are  Beractant (BLES) and Poractant alfa (Curosurf).  Other synthetic and/or modified surfactants may also  be used.  A disease to be treated with such a lentiviral vector and a surfactant may be a genetic disease.   Alternatively or  in addition,  the disease  to be  treated may be a  respiratory disease, particularly a  genetic  respiratory  disease;  or  a  cardiovascular  disease  or  blood  disorder,  particularly  a  genetic  cardiovascular disease or blood disorder.   Non‐limiting examples of diseases which may be treated  according  to  this aspect of the  invention  include Surfactant Protein B  (SP‐B) Deficiency; Surfactant  Protein  C  (SP‐C)  deficiency;  ABCA3  deficiency;  Pulmonary  surfactant  metabolism  dysfunction  2  (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary Dyskinesia (PCD);  Alpha  1‐antitrypsin Deficiency  (A1AD);  Pulmonary  Alveolar  Proteinosis  (PAP);  Chronic  obstructive  pulmonary  disease  (COPD);  Acute  respiratory  distress  syndrome  (ARDS);  COVID‐19;  a  pulmonary  fibrotic  disease;  a  pulmonary  allergic  condition;  a  pulmonary  bacterial  infection;  lung  cancer;  a  dysplastic change in the lungs; and haemophilia.  In some embodiments, include Surfactant Protein B  (SP‐B) Deficiency may be a preferred indication according to the invention.  A lentiviral vector for use in combination with a surfactant may comprise HN and F proteins  from a Sendai virus, as described herein.  Alternatively or in addition, a lentiviral vector for use in combination with a surfactant may be  selected  from  the  group  consisting  of  a  Human  immunodeficiency  virus  (HIV)  vector,  a  Simian  immunodeficiency  virus  (SIV)  vector,  a  Feline  immunodeficiency  virus  (FIV)  vector,  an  Equine  infectious anaemia virus (EIAV) vector, and a Visna/maedi virus vector.  Preferably said lentiviral vector  may be a SIV vector.  The transgene may encode any suitable therapeutic protein as described herein. By way of  non‐limiting  example,  said  transgene  may  be  selected  from:  (a)  a  secreted  therapeutic  protein  selected  from: Surfactant Protein B  (SP‐B), Surfactant Protein C  (SP‐C), AAT, Factor VIII, Factor VII,  Factor  IX, Factor X, Factor XI, van Willebrand Factor, Granulocyte‐Macrophage Colony‐Stimulating  Factor (GM‐CSF), decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody,  an anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or (b) ATP‐binding  cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.  In some preferred embodiments, a lentiviral vector for use in combination with a surfactant  may  comprise  a  SFTPB  promoter  fragment  as  defined  herein.    Alternatively  or  in  addition,  said  lentiviral vector may comprise a nucleic acid cassette as defined herein.    Therapeutic Indications  The nucleic acid cassettes and vectors of the present invention enable cell‐preferred or cell‐ specific expression of a transgene encoding a (therapeutic) protein. This may be further increased by  the nucleic acid cassette or vector facilitating efficient transgene expression. The nucleic acid cassettes  and vectors of  the  invention, and particularly  the F/HN‐pseudotyped  retroviral/lentiviral  (e.g. SIV)  vectors of  the  invention are  capable of:  (i)  transduction of one or more  cell/cell  type of  the  lung  parenchyma without  disruption  of  epithelial  integrity;  (ii)  persistent  gene  expression;  (iii)  lack  of  chronic  toxicity;  and/or  (iv)  efficient  repeat  administration.  Long  term/persistent  stable  gene  expression, preferably at a therapeutically‐effective  level, may be achieved using repeat doses of a  nucleic acid cassette or vector of the present invention. Alternatively, a single dose may be used to  achieve the desired long‐term expression.  Thus, advantageously, the nucleic acid cassettes and vectors of  the present  invention, and  particularly  the  retroviral/lentiviral  (e.g. SIV) vectors of  the present  invention can be used  in gene  therapy. Accordingly, the present invention provides a nucleic acid cassette or gene therapy vector as  defined herein for use in a method of treating or preventing a disease. The disease to be treated may  be chronic or acute.   The nucleic acid cassettes and vectors (viral and non‐viral) of the invention, particularly the  retroviral/lentiviral (e.g. SIV) vectors of the present invention may be used to deliver any transgene  useful  in gene therapy. Typically, the nucleic acid cassettes and vectors  (viral and non‐viral) of the  invention, particularly the retroviral/lentiviral (e.g. SIV) vectors of the present invention are for use in  gene therapy for the treatment of a disease or disorder of the lung parenchyma.   By way of example, efficient airway cell uptake properties of the nucleic acid cassettes and  vectors of  the present  invention,  and particularly  the  retroviral/lentiviral  (e.g.  SIV)  vectors of  the  invention make  them highly  suitable  for  treating  respiratory or  lung diseases, particularly genetic  respiratory diseases, particularly preferably those of/involving the lung parenchyma.   The  nucleic  acid  cassettes  and  vectors  of  the  present  invention,  and  particularly  the  retroviral/lentiviral (e.g. SIV) vectors of the invention can also be used in methods of gene therapy to  promote  secretion  of  therapeutic  proteins.  By  way  of  further  example,  the  invention  provides  secretion of therapeutic proteins into alveoli or lumen of the bronchioles. Administration of a nucleic  acid cassettes and vectors of the present  invention, and particularly a retroviral/lentiviral (e.g. SIV)  vector of the invention and its uptake by airway cells may be used to enable the use of the lungs as a  “factory” to produce a therapeutic protein that is then secreted and enters the general circulation at  therapeutic levels, where it can travel to cells/tissues of interest to elicit a therapeutic effect. Thus,  other diseases which are not respiratory tract diseases, such as cardiovascular diseases, particularly  genetic cardiovascular diseases or blood disorders, particularly blood clotting deficiencies, can also be  treated  by  the  nucleic  acid  cassettes  and  vectors  of  the  present  invention,  and  particularly  the  retroviral/lentiviral (e.g. SIV) vectors of the present invention.   Nucleic  acid  cassettes  and  vectors  of  the  present  invention,  and  particularly  the  retroviral/lentiviral  (e.g. SIV) vectors of the  invention can effectively treat a disease by providing a  transgene for the correction of the disease. For example, resulting in the expression and secretion of  SFTPB from cells of the  lung parenchyma, to compensate for the pathologically  low  levels of SFTPB  expression in patients with SFTPB deficiency. By way of further example, nucleic acid cassettes and  vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention  may  be  used  to  treat  alpha‐1‐antitrypsin  (AAT)  deficiency,  typically  by  gene  therapy with  a  AAT  transgene (SERPINA1) as described herein.  AAT is a secreted anti‐protease that is produced mainly in  the liver and then trafficked to the lung, with smaller amounts also being produced in the lung itself.  The main function of AAT is to bind and neutralise/inhibit neutrophil elastase. Gene therapy with AAT  according to the present invention is relevant to AAT deficient patient, as well as in other lung diseases  such as CF or chronic obstructive pulmonary disease (COPD), and offers the opportunity to overcome  some  of  the  problems  encountered  by  conventional  enzyme  replacement  therapy  (in which AAT  isolated from human blood and administered intravenously every week), providing stable, long‐lasting  expression  in  the  target  tissue  (lung/nasal  epithelium),  ease  of  administration  and  unlimited  availability.  Transduction  with  a  nucleic  acid  cassettes  and  vectors  of  the  present  invention,  and  particularly  a  retroviral/lentiviral  (e.g.  SIV)  vector  of  the  invention may  lead  to  secretion  of  the  recombinant protein into the lumen of the lung as well as into the circulation. One benefit of this is  that  the  therapeutic  protein  reaches  the  interstitium.  AAT  gene  therapy may  therefore  also  be  beneficial  in other disease  indications, non‐limiting examples of which  include  type 1  and  type 2  diabetes,  acute myocardial  infarction,  ischemic  heart  disease,  rheumatoid  arthritis,  inflammatory  bowel disease, transplant rejection, graft versus host (GvH) disease, multiple sclerosis, liver disease,  cirrhosis, vasculitides and infections, such as bacterial and/or viral infections.  AAT has numerous other anti‐inflammatory and tissue‐protective effects, for example in pre‐ clinical models of diabetes, graft versus host disease and inflammatory bowel disease. The production  of  AAT  in  the  lung  and/or  nose  following  transduction  according  to  the  present  invention may,  therefore, be more widely applicable, including to these indications.  Other examples of diseases  that may be  treated with gene  therapy of a  secreted protein  according to the present invention include cardiovascular diseases and blood disorders, particularly  blood clotting deficiencies  such as haemophilia  (A, B or C), von Willebrand disease and Factor VII  deficiency.   In  some  preferred  embodiments,  the  disease  to  be  treated  is  selected  from  a  surfactant  protein  deficiency,  such  as  Surfactant  Protein  B  (SFTPB) Deficiency,  Surfactant  Protein  C  (SFTPC)  deficiency, ABCA3 deficiency, Pulmonary surfactant metabolism dysfunction 2  (SMDP2) Pulmonary  surfactant  metabolism  dysfunction  3  (SMDP3),  or  other  surfactant  deficiencies;  Primary  Ciliary  Dyskinesia  (PCD);  Alpha  1‐antitrypsin  Deficiency  (A1AD);  Pulmonary  Alveolar  Proteinosis  (PAP,  hereditary  and/or  acquired);  Chronic  obstructive  pulmonary  disease  (COPD);  Acute  respiratory  distress syndrome (ARDS); COVID‐19; a pulmonary fibrotic disease   (including  idiopathic pulmonary  fibrosis);  a  pulmonary  allergic  condition;  a  pulmonary  bacterial  infection;  asthma;  lung  cancer;  a  dysplastic change in the lungs; and haemophilia.  Other examples of diseases or disorders to be treated include Primary Ciliary Dyskinesia (PCD),  acute  lung  injury,    and/or  inflammatory,  infectious,  immune  or  metabolic  conditions,  such  as  lysosomal storage diseases or a pulmonary bacterial infection, or any other lung disease or disorder.   The nucleic acid cassettes and vectors (viral and non‐viral) of the invention, particularly the  retroviral/lentiviral (e.g. SIV) vectors of the present invention, typically provide high expression levels  of a transgene of  interest, and the (therapeutic) protein encoded thereby, when administered to a  patient.  The  terms  high  expression  and  therapeutic  expression  are  used  interchangeably  herein.  Expression may  be measured  by  any  appropriate method  (qualitative  or  quantitative,  preferably  quantitative), and concentrations given in any appropriate unit of measurement, for example ng/ml  or µM.   Expression/secretion/membrane  insertion  of  a  transgene,  or  the  (therapeutic)  protein  of  interest  encoded  thereby  may  be  given  in  absolute  terms.  Alternatively,  expression/secretion/membrane insertion of a therapeutic protein may be given in relative terms, for  example relative to the expression/secretion/membrane insertion of the therapeutic protein encoded  by  a  corresponding  nucleic  acid  cassette  or  vector  of  the  invention without  the  SFTB  promoter  fragment  of  the  invention,  relative  to  the  expression/secretion/membrane  insertion  of  the  same  transgene  using  the  full‐length  SFTB  promoter  as  described  herein,  or  relative  to  the  expression/secretion/membrane insertion of the corresponding endogenous (defective) gene.    Expression may be measured in terms of mRNA or protein expression. The expression of the  therapeutic protein of the invention may be quantified relative to the endogenous protein or gene in  terms of protein concentration, mRNA copies per cell or any other appropriate unit. Secretion and/or  membrane  insertion  of  a  therapeutic  protein may  be  quantified  relative  to  secretion/membrane  insertion of the corresponding endogenous protein, or relative to the  level of secretion/membrane  insertion of the therapeutic protein introduced via an expression cassette with the same transgene  but without the SFTB promoter fragment of the invention or comprising the full‐length SFTB promoter  as described herein.   Expression  levels  of  a  nucleic  acid  encoding  a  (therapeutic)  protein  and/or  the  expression/secretion/membrane insertion of the encoded (therapeutic) protein of the invention may  be measured  ex  vivo  (e.g.  in  the  conditioned media  used  to  culture  the  cells  or within  the  cells  themselves)  or  in  vivo  (e.g.  in  the  lung  tissue,  epithelial  lining  fluid  and/or  serum/plasma)  as  appropriate. A high and/or therapeutic expression level may therefore refer to the concentration in  the lung, epithelial lining fluid and/or serum/plasma.   Repeated doses of nucleic acid cassettes and vectors (viral and non‐viral) of the  invention,  particularly the retroviral/lentiviral (e.g. SIV) vectors of the present  invention may be administered  twice‐daily, daily, twice‐weekly, weekly, monthly, every two months, every three months, every four  months, every six months, yearly, every two years, or more. Dosing may be continued for as long as  required, for example, for at least six months, at least one year, two years, three years, four years, five  years, ten years, fifteen years, twenty years, or more, up to for the lifetime of the patient to be treated.  The invention also provides nucleic acid cassettes and vectors of the present invention, and  particularly  retroviral/lentiviral  (e.g. SIV) vectors of  the  invention as described herein  for use  in a  method of gene therapy, wherein said method comprises the steps of: (a) transducing cells (e.g. lung  parenchyma)  ex  vivo  to  produce  modified  cells  expressing  a  transgene  of  interest;  and  (b)  administering the resulting modified cells.   The invention provides a method of treating a disease, the method comprising administering  a nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g.  SIV) vector of the invention to a subject. Any disease described herein may be treated according to  the invention. In particular, the invention provides a method of treating a lung disease using a nucleic  acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV) vector  of the invention. The disease to be treated may be a chronic disease.   The  invention also provides a nucleic acid cassette or vector of the present  invention, and  particularly a retroviral/lentiviral  (e.g. SIV) vector of  the  invention as described herein  for use  in a  method of treating a disease. Any disease described herein may be treated according to the invention.  In particular, the  invention provides a nucleic acid cassette or vector of the present  invention, and  particularly a retroviral/lentiviral (e.g. SIV) vector of the invention for use in a method of treating a  lung disease. The disease to be treated may be a chronic disease.   The  invention  also  provides  the  use  of  a  nucleic  acid  cassette  or  vector  of  the  present  invention, and particularly a retroviral/lentiviral (e.g. SIV) vector of the invention as described herein  in the manufacture of a medicament for use in a method of treating a disease. Any disease described  herein may be treated according to the invention. In particular, the invention provides the use of a  nucleic acid cassette or vector of the present invention, and particularly a retroviral/lentiviral (e.g. SIV)  vector of the invention for the manufacture of a medicament for use in a method of treating a lung  disease. The disease to be treated may be a chronic disease.   The invention also provides a cell comprising a nucleic acid cassette or vector of the present  invention, and particularly a retroviral/lentiviral (e.g. SIV) vector. Said cell may be a lung parenchyma  cell as described herein.  The  invention  further provides  a method of expressing  a  transgene,  typically  a  transgene  encoding a  (therapeutic) protein  in a  target cell, comprising delivering a nucleic acid cassette or a  vector of the invention into the target cells.  Said method may be carried out in vitro, ex vivo, or in  vivo, preferably  in vitro or ex vivo.   The target cell may be any appropriate cell type, such as those  described herein.  By way of non‐limiting example, the target cells may be prokaryotic or eukaryotic,  preferably  eukaryotic.    Particularly  preferred  are  mammalian  cells,  such  as  human,  non‐human  primate, mouse, rat, dog, cat, horse, or cow cells. The cells may be primary cells or cell lines.  Non‐ limiting examples of cells  include ATII cells, ATI cells, club cells and/or HEK293T cells.   The step of  delivering the nucleic acid cassette or vector may comprise integrating said nucleic acid cassette or  gene therapy vector into said target cell's genome. Any appropriate technique may be used to deliver  the nucleic acid cassette or vector, examples of which are known  in the art and within the routine  practice of one of ordinary skill in the art.  The method may further comprise a step of culturing cells  expressing the transgene, and/or isolating or purifying the expressed (therapeutic) protein from said  cells.  Any and all disclosure herein  in relation to nucleic acid cassettes or vectors of the present  invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention applies equally and  without reservation to the therapeutic uses and methods described herein.  Further, for the avoidance of doubt, any and all disclosure herein  in relation to therapeutic  uses and methods using nucleic acid cassettes or vectors of the present invention, applies equally and  without  reservation  to  the combination  therapies of  lentiviral  (e.g. SIV) vectors and surfactants as  described herein.  Long term/persistent stable gene expression, preferably at a therapeutically‐effective  level,  may be achieved using repeat doses of nucleic acid cassettes or vectors of the present invention, and  particularly  retroviral/lentiviral  (e.g.  SIV)  vectors  of  the  invention  of  the  present  invention.  Alternatively, a single dose may be used to achieve the desired long‐term expression.    Formulation and administration  The invention also provides a composition comprising a nucleic acid cassette or vector of the  present  invention,  and  particularly  a  retroviral/lentiviral  (e.g.  SIV)  vector  of  the  invention,  and  optionally a pharmaceutically acceptable carrier, excipient, buffer or diluent.  The  nucleic  acid  cassettes  or  vectors  of  the  present  invention,  and  particularly  retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered in any dosage appropriate  for achieving the desired therapeutic effect. Appropriate dosages may be determined by a clinician or  other medical practitioner using standard techniques and within the normal course of their work. Non‐ limiting examples of suitable dosages of viral vectors of the invention include 1x108 transduction units  (TU), 1x109 TU, 1x1010 TU, 1x1011 TU or more.  Non‐limiting examples of suitable dosages of non‐viral  vectors/delivery means of the invention include a maximum of 30 mL per dose, a maximum of 25 mL  per dose, a maximum of 20 mL per dose, a maximum of 15 mL per dose, a maximum of 10 mL per  dose, or less, preferably a maximum of 20 mL per dose.   Non‐limiting examples of pharmaceutically acceptable carriers  that may be comprised  in a  composition  of  the  invention  include  water,  saline,  and  phosphate‐buffered  saline.  In  some  embodiments,  however,  the  composition  is  in  lyophilized  form,  in  which  case  it may  include  a  stabilizer,  such  as  bovine  serum  albumin  (BSA).  In  some  embodiments,  it  may  be  desirable  to  formulate the composition with a preservative, such as thiomersal or sodium azide, to facilitate long‐ term storage.   The  nucleic  acid  cassettes  or  vectors  of  the  present  invention,  and  particularly  retroviral/lentiviral (e.g. SIV) vectors of the invention may be administered by any appropriate route.  It may be desired  to direct  the compositions of  the present  invention  (as described above)  to  the  respiratory system of a subject. Efficient transmission of a therapeutic/prophylactic composition or  medicament to the site of a disease or disorder  in the respiratory tract may be achieved by oral or  intra‐nasal administration, for example, as aerosols (e.g. nasal sprays), or by catheters. Typically the  nucleic acid cassettes or vectors of the present  invention, and particularly retroviral/lentiviral  (e.g.  SIV) vectors of the  invention are stable  in clinically relevant nebulisers,  inhalers (including metered  dose inhalers), catheters and aerosols, etc.   Other  routes of  administration,  including but not  limited  to  i.v.  administration,  intranasal  administration and  intraplural  injection are also encompassed by  the present  invention.    Suitable  administration routes are known in the art.   In some embodiments the nose is a preferred production site for a therapeutic protein using  nucleic acid cassettes or vectors of the present  invention, and particularly retroviral/lentiviral  (e.g.  SIV) vectors of the invention for at least one of the following reasons: (i) extracellular barriers such as  inflammatory cells and sputum are less pronounced in the nose; (ii) ease of vector administration; (iii)  smaller quantities of vector required; and  (iv) ethical considerations. Thus, nasal administration of  nucleic acid cassettes or vectors of the present  invention, and particularly retroviral/lentiviral  (e.g.  SIV) vectors of  the  invention may  result  in efficient  (high‐level) and  long‐lasting expression of  the  therapeutic protein of interest. Accordingly, nasal administration of nucleic acid cassettes or vectors  of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention may  be preferred.   Formulations for  intra‐nasal administration may be  in the form of nasal droplets or a nasal  spray. An intra‐nasal formulation may comprise droplets having approximate diameters in the range  of 100‐5000 µm, such as 500‐4000 µm, 1000‐3000 µm or 100‐1000 µm. Alternatively,  in  terms of  volume, the droplets may be in the range of about 0.001‐100 µl, such as 0.1‐50 µl or 1.0‐25 µl, or such  as 0.001‐1 µl.   The aerosol formulation may take the form of a powder, suspension or solution. The size of  aerosol particles is relevant to the delivery capability of an aerosol. Smaller particles may travel further  down the respiratory airway towards the alveoli than would larger particles. In one embodiment, the  aerosol particles have  a diameter distribution  to  facilitate delivery  along  the entire  length of  the  bronchi, bronchioles, and alveoli. Alternatively, the particle size distribution may be selected to target  a particular section of the respiratory airway, for example the alveoli. In the case of aerosol delivery  of  the medicament,  the  particles may  have  diameters  in  the  approximate  range  of  0.1‐50  µm,  preferably 1‐25 µm, more preferably 1‐5 µm.  Aerosol particles may be for delivery using a nebulizer (e.g. via the mouth) or nasal spray. An  aerosol formulation may optionally contain a propellant and/or surfactant.   The  formulation  of  pharmaceutical  aerosols  is  routine  to  those  skilled  in  the  art,  see  for  example, Sciarra, J. in Remington's Pharmaceutical Sciences (supra). The agents may be formulated as  solution  aerosols,  dispersion  or  suspension  aerosols  of  dry  powders,  emulsions  or  semisolid  preparations. The aerosol may be delivered using any propellant system known to those skilled in the  art. The aerosols may be applied to the upper respiratory tract, for example by nasal inhalation, or to  the lower respiratory tract or to both. The part of the lung that the medicament is delivered to may  be determined by  the disorder. Compositions  comprising nucleic  acid  cassettes or  vectors of  the  present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of the invention, in particular  where intranasal delivery is to be used, may comprise a humectant. This may help reduce or prevent  drying of  the mucus membrane and  to prevent  irritation of  the membranes. Suitable humectants  include,  for  instance,  sorbitol, mineral oil, vegetable oil and glycerol;  soothing agents; membrane  conditioners; sweeteners; and combinations thereof. The compositions may comprise a surfactant.  Suitable surfactants include non‐ionic, anionic and cationic surfactants. Examples of surfactants that  may be used include, for example, polyoxyethylene derivatives of fatty acid partial esters of sorbitol  anhydrides,  such  as  for  example,  Tween  80,  Polyoxyl  40  Stearate,  Polyoxy  ethylene  50  Stearate,  fusieates, bile salts and Octoxynol.  In  some  cases  after  an  initial  administration  a  subsequent  administration  of  nucleic  acid  cassettes or vectors of the present invention, and particularly retroviral/lentiviral (e.g. SIV) vectors of  the invention may be performed. The administration may, for instance, be at least a week, two weeks,  a month, two months, six months, a year or more after the initial administration. In some instances,  nucleic acid cassettes or vectors of the present  invention, and particularly retroviral/lentiviral  (e.g.  SIV) vectors of  the  invention may be administered at  least once a week, once a  fortnight, once a  month, every two months, every six months, annually or at longer intervals. Preferably, administration  is every six months, more preferably annually. The nucleic acid cassettes or vectors of the present  invention, and particularly retroviral/lentiviral (e.g. SIV) vectors may, for instance, be administered at  intervals dictated by when the effects of the previous administration are decreasing.  Further, for the avoidance of doubt, any and all disclosure herein in relation formulations of  nucleic acid cassettes or vectors of the present invention, applies equally and without reservation to  the formulations of  lentiviral (e.g. SIV) vectors for combination therapies of lentiviral (e.g. SIV) vectors  and surfactants as described herein.    SEQUENCE HOMOLOGY  Any of a variety of sequence alignment methods can be used to determine percent identity,  including, without  limitation,  global methods,  local methods  and  hybrid methods,  such  as,  e.g.,  segment approach methods. Protocols to determine percent identity are routine procedures within  the scope of one skilled in the art. Global methods align sequences from the beginning to the end of  the molecule and determine the best alignment by adding up scores of individual residue pairs and by  imposing gap penalties. Non‐limiting methods include, e.g., CLUSTAL W, see, e.g., Julie D. Thompson  et al., CLUSTAL W:  Improving  the Sensitivity of Progressive Multiple Sequence Alignment Through  Sequence Weighting, Position‐ Specific Gap Penalties and Weight Matrix Choice, 22(22) Nucleic Acids  Research  4673‐4680  (1994);  and  iterative  refinement,  see,  e.g.,  Osamu  Gotoh,  Significant  Improvement  in  Accuracy  of Multiple  Protein.  Sequence  Alignments  by  Iterative  Refinement  as  Assessed by Reference to Structural Alignments, 264(4) J. MoI. Biol. 823‐838 (1996). Local methods  align sequences by  identifying one or more conserved motifs shared by all of the  input sequences.  Non‐limiting methods include, e.g., Match‐box, see, e.g., Eric Depiereux and Ernest Feytmans, Match‐ Box: A Fundamentally New Algorithm for the Simultaneous Alignment of Several Protein Sequences,  8(5)  CABIOS  501  ‐509  (1992);  Gibbs  sampling,  see,  e.g.,  C.  E.  Lawrence  et  al.,  Detecting  Subtle  Sequence  Signals: A Gibbs  Sampling  Strategy  for Multiple Alignment, 262(5131  )  Science 208‐214  (1993); Align‐M, see, e.g., Ivo Van WaIIe et al., Align‐M ‐ A New Algorithm for Multiple Alignment of  Highly Divergent Sequences, 20(9) Bioinformatics:1428‐1435 (2004).  Thus, percent sequence  identity  is determined by conventional methods. See, for example,  Altschul et al., Bull. Math. Bio. 48: 603‐16, 1986 and Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA  89:10915‐19, 1992. Briefly, two amino acid sequences are aligned to optimize the alignment scores  using a gap opening penalty of 10, a gap extension penalty of 1, and the "blosum 62" scoring matrix  of Henikoff and Henikoff (ibid.) as shown below (amino acids are indicated by the standard one‐letter  codes).  The "percent sequence identity" between two or more nucleic acid or amino acid sequences  is a function of the number of identical positions shared by the sequences. Thus, % identity may be  calculated  as  the  number  of  identical  nucleotides  /  amino  acids  divided  by  the  total  number  of  nucleotides / amino acids, multiplied by 100. Calculations of % sequence identity may also take into  account  the number of gaps, and  the  length of each gap  that needs  to be  introduced  to optimize  alignment  of  two  or more  sequences.  Sequence  comparisons  and  the  determination  of  percent  identity between two or more sequences can be carried out using specific mathematical algorithms,  such as BLAST, which will be familiar to a skilled person.    ALIGNMENT SCORES FOR DETERMINING SEQUENCE IDENTITY     A R N D C Q E G H I L K M F P S T W Y V  A 4  R ‐1 5  N ‐2 0 6  D ‐2 ‐2 1 6  C 0 ‐3 ‐3 ‐3 9  Q ‐1 1 0 0 ‐3 5  E ‐1 0 0 2 ‐4 2 5  G 0 ‐2 0 ‐1 ‐3 ‐2 ‐2 6  H ‐2 0 1 ‐1 ‐3 0 0 ‐2 8  I ‐1 ‐3 ‐3 ‐3 ‐1 ‐3 ‐3 ‐4 ‐3 4  L ‐1 ‐2 ‐3 ‐4 ‐1 ‐2 ‐3 ‐4 ‐3 2 4  K ‐1 2 0 ‐1 ‐3 1 1 ‐2 ‐1 ‐3 ‐2 5  M ‐1 ‐1 ‐2 ‐3 ‐1 0 ‐2 ‐3 ‐2 1 2 ‐1 5  F ‐2 ‐3 ‐3 ‐3 ‐2 ‐3 ‐3 ‐3 ‐1 0 0 ‐3 0 6  P ‐1 ‐2 ‐2 ‐1 ‐3 ‐1 ‐1 ‐2 ‐2 ‐3 ‐3 ‐1 ‐2 ‐4 7  S 1 ‐1 1 0 ‐1 0 0 0 ‐1 ‐2 ‐2 0 ‐1 ‐2 ‐1 4  T 0 ‐1 0 ‐1 ‐1 ‐1 ‐1 ‐2 ‐2 ‐1 ‐1 ‐1 ‐1 ‐2 ‐1 1 5  W ‐3 ‐3 ‐4 ‐4 ‐2 ‐2 ‐3 ‐2 ‐2 ‐3 ‐2 ‐3 ‐1 1 ‐4 ‐3 ‐2 11  Y ‐2 ‐2 ‐2 ‐3 ‐2 ‐1 ‐2 ‐3 2 ‐1 ‐1 ‐2 ‐1 3 ‐3 ‐2 ‐2 2 7  V 0 ‐3 ‐3 ‐3 ‐1 ‐2 ‐2 ‐3 ‐3 3 1 ‐2 1 ‐1 ‐2 ‐2 0 ‐3 ‐1 4     The percent identity is then calculated as:       Total number of identical matches    __________________________________________ x 100   [length of the longer sequence plus the    number of gaps introduced into the longer   sequence in order to align the two sequences]    Substantially homologous polypeptides are characterized as having one or more amino acid  substitutions,  deletions  or  additions.  These  changes  are  preferably  of  a  minor  nature,  that  is  conservative  amino  acid  substitutions  (as  described  herein)  and  other  substitutions  that  do  not  significantly affect the folding or activity of the polypeptide; small deletions, typically of one to about  30  amino  acids;  and  small  amino‐  or  carboxyl‐terminal  extensions,  such  as  an  amino‐terminal  methionine residue, a small linker peptide of up to about 20‐25 residues, or an affinity tag.   In  addition  to  the  20  standard  amino  acids,  non‐standard  amino  acids  (such  as  4‐ hydroxyproline, 6‐N‐methyl  lysine, 2‐aminoisobutyric acid,  isovaline and α  ‐methyl  serine) may be  substituted for amino acid residues of the polypeptides of the present invention. A limited number of  non‐conservative amino acids, amino acids that are not encoded by the genetic code, and unnatural  amino acids may be substituted for polypeptide amino acid residues. The polypeptides of the present  invention can also comprise non‐naturally occurring amino acid residues.   Non‐naturally occurring amino acids  include, without  limitation, trans‐3‐methylproline, 2,4‐ methano‐proline,  cis‐4‐hydroxyproline,  trans‐4‐hydroxy‐proline,  N‐methylglycine,  allo‐threonine,  methyl‐threonine,  hydroxy‐ethylcysteine,  hydroxyethylhomo‐cysteine,  nitro‐glutamine,  homoglutamine, pipecolic acid,  tert‐leucine, norvaline, 2‐azaphenylalanine, 3‐azaphenyl‐alanine, 4‐ azaphenyl‐alanine, and 4‐fluorophenylalanine. Several methods are known in the art for incorporating  non‐naturally occurring amino acid  residues  into proteins. For example, an  in vitro  system can be  employed wherein nonsense mutations are suppressed using chemically aminoacylated suppressor  tRNAs.  Methods  for  synthesizing  amino  acids  and  aminoacylating  tRNA  are  known  in  the  art.  Transcription and translation of plasmids containing nonsense mutations is carried out in a cell free  system comprising an E. coli S30 extract and commercially available enzymes and other  reagents.  Proteins  are  purified  by  chromatography.  See,  for  example,  Robertson  et  al.,  J.  Am.  Chem.  Soc.  113:2722, 1991; Ellman et al., Methods Enzymol. 202:301, 1991; Chung et al., Science 259:806‐9,  1993; and Chung et al., Proc. Natl. Acad. Sci. USA 90:10145‐9, 1993). In a second method, translation  is carried out in Xenopus oocytes by microinjection of mutated mRNA and chemically aminoacylated  suppressor tRNAs (Turcatti et al., J. Biol. Chem. 271:19991‐8, 1996). Within a third method, E. coli cells  are cultured in the absence of a natural amino acid that is to be replaced (e.g., phenylalanine) and in  the  presence  of  the  desired  non‐naturally  occurring  amino  acid(s)  (e.g.,  2‐azaphenylalanine,  3‐ azaphenylalanine, 4‐azaphenylalanine, or 4‐fluorophenylalanine). The non‐naturally occurring amino  acid is incorporated into the polypeptide in place of its natural counterpart. See, Koide et al., Biochem.  33:7470‐6, 1994. Naturally occurring amino acid residues can be converted to non‐naturally occurring  species by in vitro chemical modification. Chemical modification can be combined with site‐directed  mutagenesis to further expand the range of substitutions (Wynn and Richards, Protein Sci. 2:395‐403,  1993).  A limited number of non‐conservative amino acids, amino acids that are not encoded by the  genetic code, non‐naturally occurring amino acids, and unnatural amino acids may be substituted for  amino acid residues of polypeptides of the present invention.  Essential amino acids in the polypeptides of the present invention can be identified according  to procedures known in the art, such as site‐directed mutagenesis or alanine‐scanning mutagenesis  (Cunningham  and Wells,  Science  244:  1081‐5,  1989).  Sites  of  biological  interaction  can  also  be  determined by physical analysis of structure, as determined by such techniques as nuclear magnetic  resonance, crystallography, electron diffraction or photoaffinity labeling, in conjunction with mutation  of putative contact site amino acids. See, for example, de Vos et al., Science 255:306‐12, 1992; Smith  et al., J. Mol. Biol. 224:899‐904, 1992; Wlodaver et al., FEBS Lett. 309:59‐64, 1992. The identities of  essential amino acids can also be inferred from analysis of homologies with related components (e.g.  the translocation or protease components) of the polypeptides of the present invention.  Multiple  amino  acid  substitutions  can  be  made  and  tested  using  known  methods  of  mutagenesis and screening, such as those disclosed by Reidhaar‐Olson and Sauer (Science 241:53‐7,  1988) or Bowie and Sauer (Proc. Natl. Acad. Sci. USA 86:2152‐6, 1989). Briefly, these authors disclose  methods  for  simultaneously  randomizing  two  or  more  positions  in  a  polypeptide,  selecting  for  functional  polypeptide,  and  then  sequencing  the  mutagenized  polypeptides  to  determine  the  spectrum of allowable substitutions at each position. Other methods that can be used include phage  display  (e.g., Lowman et al., Biochem. 30:10832‐7, 1991; Ladner et al., U.S. Patent No. 5,223,409;  Huse, WIPO  Publication WO  92/06204)  and  region‐directed mutagenesis  (Derbyshire  et  al., Gene  46:145, 1986; Ner et al., DNA 7:127, 1988).  Multiple  amino  acid  substitutions  can  be  made  and  tested  using  known  methods  of  mutagenesis and screening, such as those disclosed by Reidhaar‐Olson and Sauer (Science 241:53‐7,  1988) or Bowie and Sauer (Proc. Natl. Acad. Sci. USA 86:2152‐6, 1989). Briefly, these authors disclose  methods  for  simultaneously  randomizing  two  or  more  positions  in  a  polypeptide,  selecting  for  functional  polypeptide,  and  then  sequencing  the  mutagenized  polypeptides  to  determine  the  spectrum of allowable substitutions at each position. Other methods that can be used include phage  display  (e.g., Lowman et al., Biochem. 30:10832‐7, 1991; Ladner et al., U.S. Patent No. 5,223,409;  Huse, WIPO  Publication WO  92/06204)  and  region‐directed mutagenesis  (Derbyshire  et  al., Gene  46:145, 1986; Ner et al., DNA 7:127, 1988).    SEQUENCE INFORMATION    Key to Sequences    SEQ ID NO: 1  Core SFTPB promoter (mSPB)  SEQ ID NO: 2  972bp SFTPB genomic promoter sequence  SEQ ID NO: 3  5’ fragment portion of SFTPB exon 1   SEQ ID NO: 4  full length SFTPB promoter sequence (fSPB)  SEQ ID NO: 5  Exemplary hCEF promoter  SEQ ID NO: 6  Exemplary CMV promoter  SEQ ID NO: 7  Exemplary EF1aS promoter  SEQ ID NO: 8  Exemplary PGK promoter  SEQ ID NO: 9  Exemplary core SFTPC promoter (mSPC)  SEQ ID NO: 10  full length SFTPC promoter sequence (fSPC)  SEQ ID NO: 11  CpG‐free CMV enhancer forward  SEQ ID NO: 12  CpG‐free CMV enhancer reverse  SEQ ID NO: 13  ELF3‐1 enhancer forward  SEQ ID NO: 14  ELF3‐1 enhancer reverse  SEQ ID NO: 15  ELF3‐2 enhancer forward  SEQ ID NO: 16  ELF3‐2 enhancer reverse  SEQ ID NO: 17  ELF3‐3 enhancer forward  SEQ ID NO: 18  ELF3‐3 enhancer reverse  SEQ ID NO: 19  ELF3‐4 enhancer forward  SEQ ID NO: 20  ELF3‐4 enhancer reverse  SEQ ID NO: 21  actin enhancer forward  SEQ ID NO: 22  actin enhancer reverse  SEQ ID NO: 23  LMO7‐1 enhancer forward  SEQ ID NO: 24  LMO7‐1 enhancer reverse  SEQ ID NO: 25  LMO7‐2 enhancer reverse  SEQ ID NO: 26  SFTPB enhancer forward  SEQ ID NO: 27  SFTPC enhancer forward  SEQ ID NO: 28  SFTPC enhancer reverse  SEQ ID NO: 29  SLC34A2‐1 enhancer forward  SEQ ID NO: 30  SLC34A2‐1 enhancer reverse  SEQ ID NO: 31  SLC34A2‐2 enhancer forward  SEQ ID NO: 32  SLC34A2‐2 enhancer reverse  SEQ ID NO: 33  SLC34A2‐3 enhancer reverse  SEQ ID NO: 34  SV40 enhancer forward  SEQ ID NO: 35  SV40 enhancer reverse  SEQ ID NO: 36  VEGFA‐1 enhancer forward  SEQ ID NO: 37  VEGFA‐1 enhancer reverse  SEQ ID NO: 38  VEGFA‐2 enhancer forward  SEQ ID NO: 39  VEGFA‐3 enhancer forward  SEQ ID NO: 40  VEGFA‐3 enhancer reverse  SEQ ID NO: 41  SLC34A2 enhancer + core SFTPB promoter  SEQ ID NO: 42  VEGFA enhancer + core SFTPB promoter  SEQ ID NO: 43  CpG‐free CMV enhancer + core SFTPB promoter  SEQ ID NO: 44  ELF3 enhancer + core SFTPB promoter  SEQ ID NO: 45  SV40 enhancer + core SFTPB promoter  SEQ ID NO: 46  Alv‐01 (CMV forward enhancer + mSFPB promoter)  SEQ ID NO: 47  Alv‐02 (SLC34A2‐2 forward enhancer + mSFPB promoter)  SEQ ID NO: 48  Alv‐03 (VEGFA‐2 forward enhancer + mSFPB promoter)  SEQ ID NO: 49  SFTPB promoter forward primer  SEQ ID NO: 50  SFTPB promoter reverse primer  SEQ ID NO: 51  Exemplary linker between SFTPB promoter fragment and enhancer   SEQ ID NO: 52  Exemplary WPRE component (mWPRE)  SEQ ID NO: 53  Exemplary SFTPB transgene  SEQ ID NO: 54  human surfactant protein B (hSP‐B) transgene  SEQ ID NO: 55  codon‐optimised hSP‐B transgene  SEQ  ID NO: 56 Homo  sapiens  surfactant protein B  (SFTPB), RefSeqGene on  chromosome 2.  (NCBI  Reference Sequence: NG_016967.1) (18425bp)  SEQ ID NO: 57  Exemplary SFTPB polypeptide  SEQ ID NO: 58  Exemplified ABACA3 (ABCA3) transgene  SEQ ID NO: 59  Exemplified human ABACA3 (hABCA3) transgene  SEQ ID NO: 60  codon‐optimised hABCA3 transgene  SEQ ID NO: 61  Exemplified Human ABCA3 polypeptide  SEQ ID NO: 62  Exemplary SFTPC transgene  SEQ ID NO: 63  Exemplary SFTPC polypeptide  SEQ ID NO: 64  Exemplary hGM‐CSF transgene  SEQ ID NO: 65  Exemplary hGM‐CSF polypeptide  SEQ ID NO: 66  Exemplary Human DCN (Decorin) transgene  SEQ ID NO: 67  Exemplary Human Decorin polypeptide  SEQ ID NO: 68  Exemplary Human TRIM72 transgene  SEQ ID NO: 69  Exemplary Human TRIM72 polypeptide  SEQ ID NO: 70  Exemplary AAT transgene (SERPINA1)  SEQ ID NO: 71  Exemplary A1A1 polypeptide  SEQ ID NO: 72   Exemplary FVIII transgene (N6)  SEQ ID NO: 73   Exemplary FVIII transgene (V3)  SEQ ID NO: 74   Exemplary FVIII polypeptide (N6)  SEQ ID NO: 75   Exemplary FVIII polypeptide (V3)  SEQ ID NO: 76  corresponds to SLC34A2 enhancer + core SFTPB promoter of SEQ ID NO: 41 without  restriction site    SEQ ID NO: 77  corresponds to VEGFA enhancer + core SFTPB promoter of SEQ ID NO: 4241 without  restriction site    SEQ ID NO: 78  corresponds to CpG‐free CMV enhancer + core SFTPB promoter of SEQ ID NO: 43  without restriction site    SEQ ID NO: 79  corresponds to ELF3 enhancer + core SFTPB promoter of SEQ ID NO: 44 without  restriction site    SEQ ID NO: 80  corresponds to SV40 enhancer + core SFTPB promoter of SEQ ID NO: 45 without  restriction site    SEQ ID NO: 81  corresponds to Alv‐01 (CMV forward enhancer + mSFPB promoter) of SEQ ID NO: 46  without restriction site    SEQ ID NO: 82  corresponds to Alv‐02 (SLC34A2‐2 forward enhancer + mSFPB promoter) of SEQ ID  NO: 47 without restriction site  SEQ ID NO: 83  corresponds to Alv‐03 (VEGFA‐2 forward enhancer + mSFPB promoter) of SEQ ID NO:  48 without restriction site    SEQ ID NOs: 84‐86 correspond to the exemplary SFTPC transgene of SEQ ID NO: 62 without either/both  restriction site    SEQ ID NO: 1 core SFTPB Promoter* (635bp)  TATAGGGCTGTCTGGGAGCCACTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGGGTATTTTGTTTTCTGTTT TGTGTAAATGCTCTTCTGACTAATGCAAACCATGTGTCCATAGAACCAGAAGATTTTTCCAGGGGAAAAGGTA AGGAGGTGGTGAGAGTGTCCTGGGTCTGCCCTTCCAGGGCTTGCCCTGGGTTAAGAGCCAGGCAGGAAGCTC TCAAGAGCATTGCTCAAGAGTAGAGGGGGCCTGGGAGGCCCAGGGAGGGGATGGGAGGGGAACACCCAGG CTGCCCCCAACCAGATGCCCTCCACCCTCCTCAACCTCCCTCCCACGGCCTGGAGAGGTGGGACCAGGTATGG AGGCTTGAGAGCCCCTGGTTGGAGGAAGCCACAAGTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGG AGGCAGGAACAGGCCATCAGCCAGGACAGGTGGTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGG GATCAAGCACCTGGAGGGCTCTTCAGAGCAAAGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCA CGCCCCGCCCAGCTATAAGGGGCCATGCMCCAAGCAGGGTACCCAGGCTGCAGAGGGTGCC    *wherein M is C or A, preferably M is C    SEQ ID NO: 2 972bp SFTPB genomic promoter sequence – bases 4845‐5816 of NG_016967.1 (972bp)  TTCTTTCTGCTGAACCATCGCAGCTATGCCCCAGCCCCTACCCTGGAGGGGTCCCCAGGGGCCATGGG CAGCACCTCCTGTATAGGGCTGTCTGGGAGCCACTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGG GTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTAATGCAAACCATGTGTCCATAGAACCAGA AGATTTTTCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCTGGGTCTGCCCTTCCAGGGCTTGCC CTGGGTTAAGAGCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGTAGAGGGGGCCTGGGAGGCCC AGGGAGGGGATGGGAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCTCCACCCTCCTCAACCTCC CTCCCACGGCCTGGAGAGGTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTGGAGGAAGCCACAA GTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGCCAGGACAGGTG GTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCTTCAGAGCA AAGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGGCCAT GCACCAAGCAGGGTACCCAGGCTGCAGAGGTGCCATGGCTGAGTCACACCTGCTGCAGTGGCTGCTGC TGCTGCTGCCCACGCTCTGTGGCCCAGGCACTGGTGAGTCTCCCCCAGCCTCCCCTCTCCTAGGCAGC TCCACCACTCACTGAGCACTGCTTTGTGCTAGGCATTAACCCAAGTCTGTCCTCATTTTAAAGACAAG GCAGCTGGGGTTCAGAGAGGGTTCAGAGCTTATCCAAGGTCACACAGCTGGCGGGTCCAGGAGCAGGT GGAACCCAGAGCTGTCTGAC The bold and underlined 5’ and 3’ sequences were used  in the art to design primers to this 972bp  SFTPB promoter.    Exon 1 of the SFTPB gene (as described by NG_016967.1) is dash‐underlined (corresponding to bases  5543‐5565 of NG_016967.1). The first A of this exon is the starting nucleotide of the mRNA generated  by the SFTPB promoter.    The bold and  italicised ATG  (corresponding to bases 5559‐5561 of NG_016967.1) encodes the first  methionine of pre‐pro‐SFTPB.    3’ of the double‐underlined exon 1 is a partial portion of intron 1 (corresponding to bases 5626‐5816  of NG_016967.1).     The first base (base 81 of SEQ  ID NO: 2) and  last base (base 710 of SEQ ID NO: 2) of the core SFTP  promoter fragment of the invention are double‐underlined.    The wavy‐underlined sequence is the predicted TATA box of the SFTPB promoter.    SEQ ID NO: 3  5’ fragment portion of SFTPB exon 1   AGGCTGCAGAGG    SEQ ID NO: 4  full length SFTPB promoter sequence (1054bp)  GGATCCTCCCTCCTCGGCCTCCCAAAGTGCCAGGATTACAGGAGTGAGCCACCACACCCAGCCCCATCTCTTTT CATCATGGTACTAATTCCTGCCCGTCCACCCACAAAAGCACTGTAGTCGTTCCCGAGTATAGAGGCCTGTGAG CCTCCACTAGGGAGAGGGCTCCTGCAGAGATCAGATAAATTGATCACAATGGCTGGGGTGGTGGCAATGTGC TAATGCTCTCTTTCTTCCACTCAAGATATCCTCTGTCTCCCTCAGCCTGTGAGCTTTTTCTCCAGTGTGCTCTGCC AGTGGGGGCCTTGCCTGAGAGCCCCTGCAGCTGCAGAGGACAGTTTCTTTCTGCTGAACCATCGCAGCTATGC CCCAGCCCCTACCCTGGAGGGGTCCCCAGGGGCCATGGGCAGCACCTCCTGTATAGGGCTGTCTGGGAGCCA CTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGGGTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTA ATGCAAACCATGTGTCCATAGAACCAGAAGATTTTTCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCT GGGTCTGCCCTTCCAGGGCTTGCCCTGGGTTAAGAGCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGT AGAGGGGGCCTGGGAGGCCCAGGGAGGGGATGGGAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCT CCACCCTCCTCAACCTCCCTCCCACGGCCTGGAGAGGTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTG GAGGAAGCCACAAGTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGC CAGGACAGGTGGTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCT TCAGAGCAAAGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGG CCATGCCCCAAGCAGGGTACCCAGGCTGCAGAGGGTGCC    SEQ ID NO: 5 Exemplified hCEF promoter (562bp)  GTTACATAACTTATGGTAAATGGCCTGCCTGGCTGACTGCCCAATGACCCCTGCCCAATGATGTCAATAATGAT GTATGTTCCCATGTAATGCCAATAGGGACTTTCCATTGATGTCAATGGGTGGAGTATTTATGGTAACTGCCCAC TTGGCAGTACATCAAGTGTATCATATGCCAAGTATGCCCCCTATTGATGTCAATGATGGTAAATGGCCTGCCTG GCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTATGTATTAGTCATTGCTATTAC CATGGGAATTCACTAGTGGAGAAGAGCATGCTTGAGGGCTGAGTGCCCCTCAGTGGGCAGAGAGCACATGG CCCACAGTCCCTGAGAAGTTGGGGGGAGGGGTGGGCAATTGAACTGGTGCCTAGAGAAGGTGGGGCTTGGG TAAACTGGGAAAGTGATGTGGTGTACTGGCTCCACCTTTTTCCCCAGGGTGGGGGAGAACCATATATAAGTGC AGTAGTCTCTGTGAACATTCAAGCTTCTGCCTTCTCCCTCCTGTGAGTTT    SEQ ID NO: 6 Exemplified CMV promoter (855bp)  CAATATTGGCCATTAGCCATATTATTCATTGGTTATATAGCATAAATCAATATTGGCTATTGGCCATTGCATACG TTGTATCTATATCATAATATGTACATTTATATTGGCTCATGTCCAATATGACCGCCATGTTGGCATTGATTATTG ACTAGTTATTAATAGTAATCAATTACGGGGTCATTAGTTCATAGCCCATATATGGAGTTCCGCGTTACATAACT TACGGTAAATGGCCCGCCTGGCTGACCGCCCAACGACCCCCGCCCATTGACGTCAATAATGACGTATGTTCCC ATAGTAACGCCAATAGGGACTTTCCATTGACGTCAATGGGTGGAGTATTTACGGTAAACTGCCCACTTGGCAG TACATCAAGTGTATCATATGCCAAGTCCGCCCCCTATTGACGTCAATGACGGTAAATGGCCCGCCTGGCATTAT GCCCAGTACATGACCTTACGGGACTTTCCTACTTGGCAGTACATCTACGTATTAGTCATCGCTATTACCATGGT GATGCGGTTTTGGCAGTACACCAATGGGCGTGGATAGCGGTTTGACTCACGGGGATTTCCAAGTCTCCACCCC ATTGACGTCAATGGGAGTTTGTTTTGGCACCAAAATCAACGGGACTTTCCAAAATGTCGTAATAACCCCGCCCC GTTGACGCAAATGGGCGGTAGGCGTGTACGGTGGGAGGTCTATATAAGCAGAGCTCGTTTAGTGAACCGTCA GATCACTAGAAGCTTTATTGCGGTAGTTTATCACAGTTAAATTGCTAACGCAGTCAGTGCTTCTGACACAACAG TCTCGAACTTAAGCTGCAGAAGTTGGTCGTGAGGCACTGGGCAG    SEQ ID NO: 7  Exemplary EF1aS promoter (236bp)  CGTGAGGCTCCGGTGCCCGTCAGTGGGCAGAGCGCACATCGCCCACAGTCCCCGAGAAGTTGGGGGGAGGG GTCGGCAATTGAACCGGTGCCTAGAGAAGGTGGCGCGGGGTAAACTGGGAAAGTGATGTCGTGTACTGGCT CCGCCTTTTTCCCGAGGGTGGGGGAGAACCGTATATAAGTGCAGTAGTCGCCGTGAACGTTCTTTTTCGCAAC GGGTTTGCCGCCAGAACACAG    SEQ ID NO: 8  Exemplary PGK promoter (529bp)  GATCTTCGAATTCCCACGGGGTTGGGGTTGCGCCTTTTCCAAGGCAGCCCTGGGTTTGCGCAGGGACGCGGC TGCTCTGGGCGTGGTTCCGGGAAACGCAGCGGCGCCGACCCTGGGTCTCGCACATTCTTCACGTCCGTTCGCA GCGTCACCCGGATCTTCGCCGCTACCCTTGTGGGCCCCCCGGCGACGCTTCCTGCTCCGCCCCTAAGTCGGGA AGGTTCCTTGCGGTTCGCGGCGTGCCGGACGTGACAAACGGAAGCCGCACGTCTCACTAGTACCCTCGCAGA CGGACAGCGCCAGGGAGCAATGGCAGCGCGCCGACCGCGATGGGCTGTGGCCAATAGCGGCTGCTCAGCAG GGCGCGCCGAGAGCAGCGGCCGGGAAGGGGCGGTGCGGGAGGCGGGGTGTGGGGCGGTAGTGTGGGCCC TGTTCCTGCCCGCGCGGTGTTCCGCATTCTGCAAGCCTCCGGAGCGCACGTCGGCAGTCGGCTCCCTCGTTGA CCGAATCACCGACCTCTCTCCCCAGG    SEQ ID NO: 9  Exemplary core SFTPC promoter (345bp)  CAGGGCAGCAGGGGCAGGTGCCAGCAAGGAAGGCAGGCACGCCAGGAAGACACCCATGGTGAGAAGTGCA GATGGCCCGAGGGCAAGTTTGCTCAACTCACCCAGGTTTGCTCTTGCTGGGGCCAAGAGGACTCATGTGCCA GGGCCAAGGGCCCTTGGGGGCTCTCACAGGGGGCTTATCTGGGCTTCGGTTCTGGAGGGCCAGGAACAAAC AGGCTTCAAAGCCAAGGGCTTGGCTGGCACACAGGGGGCTTGGTCCTTCACCTCTGTCCCCTCTCCCTACGGA CACATATAAGACCCTGGTCACACCTGGGAGAGGAGGAGAGGAGAGCATAGCACCTGCAG    SEQ ID NO: 10  full length SFTPC promoter sequence (3813bp)  AGCTTGAAGACTGCTGCTCTCTACCACGTTAGCTCCCCTGTGGCGGAGGATGACTGTCCTAAGAGCTCATAGG CACGGGGCAAGGGGAGCCTGGCTGTCAGCTGCCTGGGCTTCTAGTCTTGTGCTTTTTGCACACCAGTCCAGGG AACGAAGACCACTGGCTTTAAGATGCTTCCCCAGCTGTCCCCAGACTCTGCCAGCAGGGGATTCTCTGGTCTG AGCTTAAGTTGTGTTCTCCCAGCCAGGGATGCCCCTGCCCTTTGATGTCTCCTTGCTGCCACACATTTAGCCGC CCTCCCCATGCCAGCTTGGGGGGAGGGAAGCAGTGAGGGTAGGGAGGTGGCTGGGGCAGCTGGGCAACTG TCCCCACCCGTCCCTGGCACGGCTCTGCCCAGTACACAAAGAGCAAAGTGAATCTTGTCCCCACCCCTGCAGCT GAGGGGCTGGAGGAGGAAACGGGGAGGCCCACACAGAAGGGGTGGCCACCGTGGGGCTGTCCATCACTCA GGGCTCTCAGAGGGAGTCAACCCAGAAACAGACAAAGAGGGTGAGTCTGGGCTGTGTTCTTAGCTAGTGAG AGGTCCCCTAGAGGATGAAGTAGATGATGCTAATGAGGATGACTGGATGTCACACCCATGATGCTATTAGGT CCTCATAATAGCATAGTGAGGTGGACAGCTAGTACCTGACCCATCTCACAGATGAACATACTAATGCCTAACA AAGCAGAACAACTCACGCTGGGTCCCAGAGCTGGCCAGTGGAAGCACTGAGACCTCCACATACTGAAGGCAT GGACTATTGACCGCTGTTGGTATTGGTCTCATCATTGACTATCATTAAGTGTTGGCTGTTTGCATGCTTCCTGCC CAGTGGCAGGTTCAAAGAAGCCCGCAGGAAGCGTGCTCCTTTCTTTCCCAGGGCCCGCAATTGGGCTGGAAG ATAGAGCAACAAAAAGCGCCCATGTAACTCATGGGAACATTCATGTGTGCTGAATGGCAGGTGAAGGTGCCA CAGAGAGGCTGAGGATTTCAGAGGGCACCATGAACTGGAGTGAGGTCGCAGAGCAGGTGCCATTGGCTCTT GGCCTGTTTGGGTGGGTGGCATTCAGAGAGGTGGAAGGTCAGATGCACTGTTCACGCCTGTAATCTCAGCAC TTTGGAAGGCCAAGGTAGGAGGATCACACGAGGCCAGGAGATCAACGCTGCAGTGAGCTATGAAGCTGTGA TTGCACCACTGCACTGCAGCTTGGGTAACAGAGTGAGACCCTGTCTCTAAATAATTAAATAAATAAAATAAAA ATAAAACCGGAGAAGTGGAGAGGGATTGGAGGTGGGCTTTCACAGAGGGAGAAACGGCTTAAGTACAGGC CAAAAAGTGAGAAGGCTGCAGACAGGGCTGGTAGGGGGAGGGGGAAATTTGGCACACCCAGCTAAAGGTC CTTCTGTGTCTCCTTCTCCAAGGAACCCAAGACCTTCACTTGGTTGGTGTGAGCACTCCAGGAGGCAGGCACCC TCCCTCAGCCCTCAAGCAAGCAAAAATGGGTTTAAAAAAAGAAGGAGAAGAAGCAGCAGCAGCAGCCGCCA CAGAGCTTGTGACAGCTACAGCCTAAGGGCAACAGGCAGGGGAGACCAAGGACCAGAAAGAGCAGAGGCTT TTTCAAAGAAAGAGATCCCTCTCCCAGCACCCAGCGATGGCGGCAAACCCCACCCACAGTGCCTGCTAAGAAC AAGTCCCACGTGAGAACAACATGGCCCCCCGAGATGCCCACAGGGACCCCGAGATGCCTGCAGTGCTTGGCT CTCCCGCTGGCCAGCTGCCCACCTGGCTCAGGCCCAGTACTCGTGAGTCAGCCGATCAATCCCAATGTTGCCA GGATGATGGGGGCGGGAGTAAGGGCCCTGGGGGAGGGCAGGGGTGGGCACTGCAGGCGAGCTGTCTCCCA CATCTGGCACCTGCACACAGCTGAGGCCGAGCTGAGAGGATGCTTCTGTGGGCTCCCCCTCCTCCCGGCACCC CTCCCCTCCTTTCACTGTCCTAGGACACTCTCTGGCTGCTGGAGTCTTAGGCAAATATTTAAAGGGGCAGCAAG GGGGTGAGGAGGGTGGTGGGAGCAAACACTTCCCTCCCTTTCTTCCTCCCTGCGCTTCTCAGGGGCTCTCAGT TCAGATGCCATGCTGTTATGCAACCTTGGGGCTGAAGGCCCTCCAGATGGAGAGGGGGACAGGGGAACCTG CCAGCTCATGACCGAAGGGCAGGGCCCAGGTGGGAGGGGGCTGGGGCAGGGGACAGGAACTGGGGTGGC ATGTTAAAGGACAGGAGGCTGGTTGGGCATGGTGGCTCACACCTGTAATCCTAGCACTTTAGGAGGCCGAGA TCACTTGAGCCCAGGAGTTCAAGACCAGCCTGGGCAACATGGTGAAAACTCATCTCTATAAAACAAGCAAAAA TTAGCTGGGCACAATGGCATGCACTGGTAGTCCCAGCTACTTGGGATGCTGAGGTGTGAGGATCACCGGAGT CCAGGAGGTCAACGCTGCAGTGAGCAGTGATCTCGCTACTGCATGCCAGCCTGGGTGATAAAGTGAGACCCT GTCTCACAACAAAACAAAACAAAACAAAACAAAAGGATAGGAGGTTAAGGGAGCGAGCCCAGGCCTGGACT CTGCCACAGTCACCTGAGTTTAGAGGTGAAGGGACTTTAGAGACCACCTGGCCCAAGGAGTGAAAGGGAACC TGAGAGAAAGGGTGAGCCAGCCCAAAGTCATTGGCAGATTTGACCTCGTAAATACATAGAGATGGCTTTGGG AAGGCACTAGGAAAGACAGAGAAAAGAGAAGGAGACAGTCCTCAAAGCTGATCGTATTTGGGTGAGTATTA TTCTCAGGGCAAATTTAGGATCAGAGGATGCAGAAAGGGGAGTCTAGAGGGGTAGAGTGTAGACCACAGGG TGAGTGAGCTGATTCGAGGATGGGGAGACTGGGAGCCCACCAGTGACCAGAGCCAGCCCTGTTCAGGGCTG TCCGGGCAGAAGAAAGCAGTGTCAGACCTGGAATCTGCCATCAGCACAGCCTGCAATTGACAGACAAGCCCA GAGCAAAGAAGGAAGCACTGCACATGAGTAAGAGCTTGCCACCAGTGGGGACAGAGTTTCCAGAATTAGGA AAATAATCACTGGGGGCAAGTTTGAGGTTGGTACCAGATATGTGGGAGGAGGCAAGGTAAGGGAAAGAGTA CTTGAAGTTGGAACTGGTCCTTGCAGGGAAATGCACATTTATGAAACCCCGAAAACTGATGTCAAAGCACCTC CTGCCTTGGGCAGAGTCCTCTCAGAGTCTACAGGTGCTGCCTCCAGAACCCTCTTCCTGGAGCGCATCCCTATG TATCTAGAAATTCTGCTGGGAAATATGATGGTCAGACCCTTGGCCACCTGAAAGGTTCAGGGTGGTAGAAGA AAAAGGAAAGCCACAGGGCAGCAGGGGCAGGTGCCAGCAAGGAAGGCAGGCACGCCAGGAAGACACCCAT GGTGAGAAGTGCAGATGGCCCGAGGGCAAGTTTGCTCAACTCACCCAGGTTTGCTCTTGCTGGGGCCAAGAG GACTCATGTGCCAGGGCCAAGGGCCCTTGGGGGCTCTCACAGGGGGCTTATCTGGGCTTCGGTTCTGGAGGG CCAGGAACAAACAGGCTTCAAAGCCAAGGGCTTGGCTGGCACACAGGGGGCTTGGTCCTTCACCTCTGTCCCC TCTCCCTACGGACACATATAAGACCCTGGTCACACCTGGGAGAGGAGGAGAGGAGAGCATAGCACCTGCAG    SEQ ID NO: 11 CpG‐Free CMV Enhancer Forward (302bp)  GTTACATAACTTATGGTAAATGGCCTGCCTGGCTGACTGCCCAATGACCCCTGCCCAATGATGTCAATAATGAT GTATGTTCCCATGTAATGCCAATAGGGACTTTCCATTGATGTCAATGGGTGGAGTATTTATGGTAACTGCCCAC TTGGCAGTACATCAAGTGTATCATATGCCAAGTATGCCCCCTATTGATGTCAATGATGGTAAATGGCCTGCCTG GCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTATGTATTAGTCATTGCTATTAC CATGG    SEQ ID NO: 12 CpG‐Free CMV Enhancer Reverse (302bp)  CCATGGTAATAGCAATGACTAATACATAGATGTACTGCCAAGTAGGAAAGTCCCATAAGGTCATGTACTGGGC ATAATGCCAGGCAGGCCATTTACCATCATTGACATCAATAGGGGGCATACTTGGCATATGATACACTTGATGT ACTGCCAAGTGGGCAGTTACCATAAATACTCCACCCATTGACATCAATGGAAAGTCCCTATTGGCATTACATG GGAACATACATCATTATTGACATCATTGGGCAGGGGTCATTGGGCAGTCAGCCAGGCAGGCCATTTACCATAA GTTATGTAAC    SEQ ID NO: 13  ELF3‐1 enhancer forward (233bp)  CAGGGGGCCTGGGGCTAGGGGACAGCGGGGCCTTTTTCTTACCTAAGAGGCTACAAAGGGGAGGGTGAGGT GCCTGAAAAGGTCGGCTCTGCTGCCACCGTGTGGCTAATACTAAGAACAGCAGCAAGCCTAGGCTGTTGGTCT GGTGGTTGGGGAATTGGGAGAGCAGAAAGTAATATAGAGTCACCCAACTCACATTTTCTCAATCAGCTCCTTC TCATCCAATGCTCCCG    SEQ ID NO: 14  ELF3‐1 enhancer reverse (233bp)  CGGGAGCATTGGATGAGAAGGAGCTGATTGAGAAAATGTGAGTTGGGTGACTCTATATTACTTTCTGCTCTCC CAATTCCCCAACCACCAGACCAACAGCCTAGGCTTGCTGCTGTTCTTAGTATTAGCCACACGGTGGCAGCAGA GCCGACCTTTTCAGGCACCTCACCCTCCCCTTTGTAGCCTCTTAGGTAAGAAAAAGGCCCCGCTGTCCCCTAGC CCCAGGCCCCCTG    SEQ ID NO: 15  ELF3‐2 enhancer forward (654bp)  TAGGGATGGGCCGAGGCTGGCACTGATGCTAGACTTCCGTGCACAGGGCAAGTATGGACAAGCCCCAAGTG GCTTTGTGAGGCCCACACAGTGAAGCTTGGGAAATGGGAAGTGGGGCTGCGCCCAGATTCTGGTATCTATGA CAACTAAGGCCGCTGCACATCCTCATGGCTCTCCCAGAGACCTCAGGTGAGGCCCTTCTGTGTTCCTCAAGCAC CCATGCCACCTGCGGGGTGGGGCAGGACCCTCCTACCCAGCCCTGGGCCTCCTGGGGAACCATGGGTGCACA GGGGTAACCTGAGCCAGCCTCTCTGGGGCATGGGCGGGCGTGGTGGCGTGGGCCTGGCGCCAGAGAGTGG AGCAGAATGTCAGCTCTGTGAGCCGCACCGGGTGCCAGCACTCTGCAAACAGACTCTAGTCACCAGATAGACT GGAGTCACGAACCTAACAAAGCGCTCAGCTGGGCAACTGTAACTGCAGAGGGCGGGGCCGCACAGTGCTGC CTAGTGCCTCCTGCCTTGATCTGGTCAGGGTCATGGGAGGAGACAGGTTCCCCAGGAGGGCCCATTAACAAG ATTAATTGGGGTGGCCTGGGGGACTCTGGGATGCTCACTGGAAACATATCCTGGGGGAATGAGGTAGGTGG GGAGCCAG    SEQ ID NO: 16  ELF3‐2 enhancer reverse (654bp)  CTGGCTCCCCACCTACCTCATTCCCCCAGGATATGTTTCCAGTGAGCATCCCAGAGTCCCCCAGGCCACCCCAA TTAATCTTGTTAATGGGCCCTCCTGGGGAACCTGTCTCCTCCCATGACCCTGACCAGATCAAGGCAGGAGGCA CTAGGCAGCACTGTGCGGCCCCGCCCTCTGCAGTTACAGTTGCCCAGCTGAGCGCTTTGTTAGGTTCGTGACT CCAGTCTATCTGGTGACTAGAGTCTGTTTGCAGAGTGCTGGCACCCGGTGCGGCTCACAGAGCTGACATTCTG CTCCACTCTCTGGCGCCAGGCCCACGCCACCACGCCCGCCCATGCCCCAGAGAGGCTGGCTCAGGTTACCCCT GTGCACCCATGGTTCCCCAGGAGGCCCAGGGCTGGGTAGGAGGGTCCTGCCCCACCCCGCAGGTGGCATGG GTGCTTGAGGAACACAGAAGGGCCTCACCTGAGGTCTCTGGGAGAGCCATGAGGATGTGCAGCGGCCTTAGT TGTCATAGATACCAGAATCTGGGCGCAGCCCCACTTCCCATTTCCCAAGCTTCACTGTGTGGGCCTCACAAAGC CACTTGGGGCTTGTCCATACTTGCCCTGTGCACGGAAGTCTAGCATCAGTGCCAGCCTCGGCCCATCCCTA    SEQ ID NO: 17  ELF3‐3 enhancer forward (681bp)  GGCAGGGCTTCTGGGGATTGGTGACACCCAGCTGACGTCAGGGAGGTGGAGAAGCCGCAGGTCCCTCTTATC CCCAGGGCAGTTCAGGGGCTTGACTGCATTTTAGCGATGATTGTAGTTACAGATTGCTCTCCGAAACACTGGG CAGGAAAAGCTTGCTGGTGTTCTGCTTTTGGCTCTGGGATTTGAACCCATGACTGGCTCAGGAGACTTGAGAT TACCAACGTGCTTTGTGTCCCCAACAAGTCACTTGTCCTCTCTGGGCCATCTCTAGAGCACAGCCTCCTAATTCA ATAAAATGAAGAGGCTGGAGGAGAGATTGCCAAGGACTCTTTCAAGAGGCCCAGATAGGAGAGCAGGAGGA TCGGGGGGTGGGGGGGTGTTCTGAGTTGGCCCTGTCATGGCCCTAATCTGTCCTTCCCCCATCCTTCACTCCCC CTCTTATCCCTGGTGGCCCCGGAGTGAGAACGTGACTCATCCAGCTCCAGGCACCCAGTTTCAGCCTTGCCCCA CCCCTGCCCCGGGCCTCTCATTTGCTGTTTACCCCCCAGGAGCAGGTGCACCTGGGTCACCTCACCGCAGGAA GGAGAGATATAAGGCTCTAGGGCACAGCCTGACTCCACACCCACAGATTGCCCAAGGCCAGCACAGGGGTTG GAGCTCTCTAGAAAGGTGAGGCAC    SEQ ID NO: 18  ELF3‐3 enhancer reverse (681bp)  GTGCCTCACCTTTCTAGAGAGCTCCAACCCCTGTGCTGGCCTTGGGCAATCTGTGGGTGTGGAGTCAGGCTGT GCCCTAGAGCCTTATATCTCTCCTTCCTGCGGTGAGGTGACCCAGGTGCACCTGCTCCTGGGGGGTAAACAGC AAATGAGAGGCCCGGGGCAGGGGTGGGGCAAGGCTGAAACTGGGTGCCTGGAGCTGGATGAGTCACGTTCT CACTCCGGGGCCACCAGGGATAAGAGGGGGAGTGAAGGATGGGGGAAGGACAGATTAGGGCCATGACAGG GCCAACTCAGAACACCCCCCCACCCCCCGATCCTCCTGCTCTCCTATCTGGGCCTCTTGAAAGAGTCCTTGGCA ATCTCTCCTCCAGCCTCTTCATTTTATTGAATTAGGAGGCTGTGCTCTAGAGATGGCCCAGAGAGGACAAGTG ACTTGTTGGGGACACAAAGCACGTTGGTAATCTCAAGTCTCCTGAGCCAGTCATGGGTTCAAATCCCAGAGCC AAAAGCAGAACACCAGCAAGCTTTTCCTGCCCAGTGTTTCGGAGAGCAATCTGTAACTACAATCATCGCTAAA ATGCAGTCAAGCCCCTGAACTGCCCTGGGGATAAGAGGGACCTGCGGCTTCTCCACCTCCCTGACGTCAGCTG GGTGTCACCAATCCCCAGAAGCCCTGCC    SEQ ID NO: 19  ELF3‐4 enhancer forward (571bp)  GTGGCTCAGGTGATAAGAGGAGACTCAGGACCAGTCCCTTAACAGGTTGGTGGAACTGTGAGTAGAAATTTT TTTTTGCTCCGCCCCTACCCAAAGTGGCTTCATTGAGATTCAGCCTATCTCCCACCCTTCCTGTGCTGTGATGAG GGACTCTGGGGACTGGGAGTGAATAGGCCCTGCTGAGTGCAGGGAAGACACACTCTGGCCTTCCTCACGCCA GGCTGGAGACCTGGCACTGGGCTCTGCTCCTAGCAGCCTGGAAAGCTGTTATTCTAGAACCTGATAACTGCTT GGGCTGGGCTGAAGATAAGAGGCAGCTGTTTGAAGTCCAGAGAGATAACAGCTTCAGAGAAATTTGATAAG AGGAGTGAGATGCACTGGGTTAAGCTCACCAGTCTCGCCTGTTTGCCTTTTGCCAGTTTATAATCATAACCAGC ACTGTACTCTCAAGTCAGACTTCCCAGCACCGTGAGGAACCTCAGTGTCACTGTCTCTTAATGGTAGTCTGTGC ATCTCTCCTGAGGTCCGCGCCTGGGCCTGCAGGCCTCTGTGCTTGTTCTATGAAATGACC    SEQ ID NO: 20  ELF3‐4 enhancer reverse (571bp)  GGTCATTTCATAGAACAAGCACAGAGGCCTGCAGGCCCAGGCGCGGACCTCAGGAGAGATGCACAGACTACC ATTAAGAGACAGTGACACTGAGGTTCCTCACGGTGCTGGGAAGTCTGACTTGAGAGTACAGTGCTGGTTATG ATTATAAACTGGCAAAAGGCAAACAGGCGAGACTGGTGAGCTTAACCCAGTGCATCTCACTCCTCTTATCAAA TTTCTCTGAAGCTGTTATCTCTCTGGACTTCAAACAGCTGCCTCTTATCTTCAGCCCAGCCCAAGCAGTTATCAG GTTCTAGAATAACAGCTTTCCAGGCTGCTAGGAGCAGAGCCCAGTGCCAGGTCTCCAGCCTGGCGTGAGGAA GGCCAGAGTGTGTCTTCCCTGCACTCAGCAGGGCCTATTCACTCCCAGTCCCCAGAGTCCCTCATCACAGCACA GGAAGGGTGGGAGATAGGCTGAATCTCAATGAAGCCACTTTGGGTAGGGGCGGAGCAAAAAAAAATTTCTA CTCACAGTTCCACCAACCTGTTAAGGGACTGGTCCTGAGTCTCCTCTTATCACCTGAGCCAC    SEQ ID NO: 21  actin enhancer forward (37bp)  CGCCTCCGACCAGTGTTTGCCTTTTATGGTAATAACG    SEQ ID NO: 22  actin enhancer reverse (37bp)  CGTTATTACCATAAAAGGCAAACACTGGTCGGAGGCG    SEQ ID NO: 23  LMO7‐1 enhancer forward (864bp)  ATTTGTTATAAATTCACTAATATCTTTTAAAGTGGGAGGAATGAGAGAATGACTTTAACCCATCTCTCCAAGCC ACCCAGCCTGAGGCCTTTGGTTTGTTGAATAAACGAAAACGTGGCCACTTCAGAGAAGAGCAAGGCTTCCAG CACCTTCCCACAGCCTGAATTCCATTAGCATGAGACTTTGAAACCACATCTGTTTTTCCTTATGAAATTCAAACA ATTGTCTAACCTCTTGGCTCTTAAGTCAAATGTTCAAGGGCTCATATGACAACACTTTGTTGTATATAAAAATTT GTTAAATTTAAATTGCTTTTGCAAAAAAGAGGAAAAGGGGGAATAAATAAATAGCTTTAAGGCATCTGGTTAG GATCTGCACAAGGTTGCATTCTTTCCATGTCTCCAAAAGGTTTGTTCTATCTGTCACCCATCACTTCACTGGAGT TCCTGCAGGCAGCAAAATCATGCAGAAGTTCTTTGTATAGAAATACAGTGTCTGGAGCCTAGTCTGACTTCCT GTTTGGCTGGAGCTGAGCTAGTCCATGGATAGGCAGAAAGCAGCTGTAGGAATCTGTGTTTAGAGTAGGGCT GTTTAGTAGACACAAGGGCAACCCACAGGGACCTGTGACACATCTCTATTGAAATATGGCATGATGTGCTGAT TTGTTTGTCCAAGATTTAATTATAGATGTTTAGCAGGGTTGCATAAACAACAGAGGGATAGGAGAGATAAGG AGGGAAGATTCAGGGAAAAAAATCAGTGGCTTATATTTTTGATCAGATATATTATGTGCCAGAATGAATCTCT ATCATGCTTTGAGATTTGCATGTGTGTGTGTGTGTGTGTGTAGACACATCAAAGGTC    SEQ ID NO: 24  LMO7‐1 enhancer reverse (864bp)  GACCTTTGATGTGTCTACACACACACACACACACACATGCAAATCTCAAAGCATGATAGAGATTCATTCTGGCA CATAATATATCTGATCAAAAATATAAGCCACTGATTTTTTTCCCTGAATCTTCCCTCCTTATCTCTCCTATCCCTCT GTTGTTTATGCAACCCTGCTAAACATCTATAATTAAATCTTGGACAAACAAATCAGCACATCATGCCATATTTCA ATAGAGATGTGTCACAGGTCCCTGTGGGTTGCCCTTGTGTCTACTAAACAGCCCTACTCTAAACACAGATTCCT ACAGCTGCTTTCTGCCTATCCATGGACTAGCTCAGCTCCAGCCAAACAGGAAGTCAGACTAGGCTCCAGACAC TGTATTTCTATACAAAGAACTTCTGCATGATTTTGCTGCCTGCAGGAACTCCAGTGAAGTGATGGGTGACAGA TAGAACAAACCTTTTGGAGACATGGAAAGAATGCAACCTTGTGCAGATCCTAACCAGATGCCTTAAAGCTATT TATTTATTCCCCCTTTTCCTCTTTTTTGCAAAAGCAATTTAAATTTAACAAATTTTTATATACAACAAAGTGTTGT CATATGAGCCCTTGAACATTTGACTTAAGAGCCAAGAGGTTAGACAATTGTTTGAATTTCATAAGGAAAAACA GATGTGGTTTCAAAGTCTCATGCTAATGGAATTCAGGCTGTGGGAAGGTGCTGGAAGCCTTGCTCTTCTCTGA AGTGGCCACGTTTTCGTTTATTCAACAAACCAAAGGCCTCAGGCTGGGTGGCTTGGAGAGATGGGTTAAAGTC ATTCTCTCATTCCTCCCACTTTAAAAGATATTAGTGAATTTATAACAAAT    SEQ ID NO: 25  LMO7‐2 enhancer reverse (675bp)  AAACACGTATTTGTTAAGGCCTTAGAGTTCAACACTCAAGATGGATTTTTGACCTCATAAATGTTAAAGTTATT GATGACAGCCACTTCAGTTATAAAATAACCTTTGGCCCTGGTGCCTGGCTTTTAAGGAGGCACGCTGACAAGT CAATTAAAATGAGTCACCTCCAACTTCCAAGATGACTGACTGAGGTGCTACAACTTCCACCAACGATTGTTTAA TTGGTATTTTATGTTAGCTTTTAAGCATTATCTGGTTAAATGGGAAATGCCTGATGAACACATTGGTTTTGATTT TAATAGAACTGATACAAAACATGTGATATAGTCACATACCTCTAACAGCTACCCCCAGTTTATTATAATGGAAT GGAACCACAACCTTAGTTTTATACACCATAAGGTCTTCAACTACCTCCTCTGGTTTTGAAACTGGTAACAGGAA ACAGCCTTTGATAAAGCATTCCTGGCATAGACACTGTACTAGGATTATTCAACCTGAGTCAGACTGTCATTAAA AGGACCAAAGGCTACAAAGACAGCCCTAGTACATAAGCCAAAGTCCCACATGGTTTTTTGTTTGTTTGTTTGTT TTTGAGACGGAGTTTTGCTCTTGTTGCCCAGGCTTGGCTCACCACAACCTCTGCCTCCCAGGTTCAAGCGACTC TCCTGCCTC    SEQ ID NO: 26  SFTPB enhancer forward (697bp)  GATGCTGATGTGACTGATTTGTAGGTGGCCAGAAGGAAAGGCGGGAAGAAAGATGTTGGTGGCAGCCAAGA CCGAGGAAGTCTGTGAGGCAGGAGGGGTGAGGGGTCCCTTGCCCTGTGACATAACCTTGTCGTGTTGGAGTT TCAGGCCCAGGTTACTTACTGAAGCATTGTCTTTTTATGTTGGACTTCTGTCTGGTCCCCCTGAGACGTTGTCTC TTCCGCAGGCCCTCGGGGGCTGGAGCACAGCTGTAGCCAACAGAACACAGGCTCTGTCTGCTGCCCTGAAGC ACAAAGGAAGTTGGTCTCTTGAGGTTTTTCTAGGAATGTTTTCCTGTGAGCAATCACAGGAGAAGCGGGAGTA AAACAAACAAAGGAGGGAAAAAGATGACACAAAAACATCATGGGAAGGCTGGCGGTTGGCAGAGGAGGCT CAGATTGAGGAAACAATCCCTCATGGAATGAATTCTTCTGTTAGTGGAGTGATGCATATTGACAGATCCAGGC ACATCTTCAGGCTGCCCTTCTGTCTGTCCCTCACCCATATATTCATTCTACAAATATTTGCTGAGAGCCTATTAT GTGTGAGGCATCATGCTAAGTGCATAACATAAATTACTTCATTTTATCTTCACAGTGTCCCTTTAAGGTAGGGT CTCTTATCCCCATATCTTACCCAAGGAACCCGAGGCTTAGAG  SEQ ID NO: 27  SFTPC enhancer forward (590bp)  CAAGTTCACCTCTCAGCTTCATTTTTGCTCATCTCTAAATGAGGAGAATGTCGTTTTGTGTGAGGTTTAGAGAC GATGTATGAAAGTCCCAAGTATACAGGGGCTGCCACCAAGCAGGTAGTGGTTGGCTGGAAACAGGAGCACA GCTAGACCAGGGGCCCTCCACCCCGAAGTTGCCATTCTGCCTGCGGAGGCTTCATTTCTAAAAGTCCAGGGGA GCCACAGTCTGGATCAGTTTCTGCCTGGAAGAAGAGCGCGCCAGGATGAGTGGGCACAAACTCGGAGGGCC CAGTGGGCGGGTCTTATCACCCACATCCTGGGGAGCTGTGGAGAGGAGAGGGTGGTGGGTGAGGTGGGGC TGGGCTGGTGGCTCAGATAAGGCAGGGACACACAGCTGGGGGAGGGTGGGGCTGAGATGGAGGGCGAGC GGCTGGCTGGCGTCCGACCAGGGCCAGGGGCAGCTGAGCTGTCCCTTCCGTCACCTGGACCTGCCTTCCTCTG TGGTGCTCTCTCTTCCTCCCTCCCTCCCCTCTTCCTGCTGCAGCTCGCCCTTTCTATCTCTTTGTGCACGGAGGTC GCTGGGACCTTGG    SEQ ID NO: 28  SFTPC enhancer reverse (590bp)  CCAAGGTCCCAGCGACCTCCGTGCACAAAGAGATAGAAAGGGCGAGCTGCAGCAGGAAGAGGGGAGGGAG GGAGGAAGAGAGAGCACCACAGAGGAAGGCAGGTCCAGGTGACGGAAGGGACAGCTCAGCTGCCCCTGGC CCTGGTCGGACGCCAGCCAGCCGCTCGCCCTCCATCTCAGCCCCACCCTCCCCCAGCTGTGTGTCCCTGCCTTA TCTGAGCCACCAGCCCAGCCCCACCTCACCCACCACCCTCTCCTCTCCACAGCTCCCCAGGATGTGGGTGATAA GACCCGCCCACTGGGCCCTCCGAGTTTGTGCCCACTCATCCTGGCGCGCTCTTCTTCCAGGCAGAAACTGATCC AGACTGTGGCTCCCCTGGACTTTTAGAAATGAAGCCTCCGCAGGCAGAATGGCAACTTCGGGGTGGAGGGCC CCTGGTCTAGCTGTGCTCCTGTTTCCAGCCAACCACTACCTGCTTGGTGGCAGCCCCTGTATACTTGGGACTTT CATACATCGTCTCTAAACCTCACACAAAACGACATTCTCCTCATTTAGAGATGAGCAAAAATGAAGCTGAGAG GTGAACTTG    SEQ ID NO: 29  SLC34A2‐1 enhancer forward (162bp)  GACTTTCCCATCAGTCTGAACTCCTGGGTTTTCCCAGGCCTGGTGACTCAGAGGGTAAGGCACCGAGCCGAGG AAGAGAAAGGAGAACCTGGACTCCTGGGTGCCTTTGTCTCACCACACAATCACCTCCCAGGAAGGCTGTCATC CAATCACTCCTGTGGA    SEQ ID NO: 30  SLC34A2‐1 enhancer reverse (162bp)  TCCACAGGAGTGATTGGATGACAGCCTTCCTGGGAGGTGATTGTGTGGTGAGACAAAGGCACCCAGGAGTCC AGGTTCTCCTTTCTCTTCCTCGGCTCGGTGCCTTACCCTCTGAGTCACCAGGCCTGGGAAAACCCAGGAGTTCA GACTGATGGGAAAGTC    SEQ ID NO: 31 SLC34A2‐2 Enhancer Forward (403bp)  AGCTAACCATTCAGACTGTCAGTTTAGATTATTGAAAGGAACAGAAGAGAAATCTTGCCTATAAAACACGGAA CTGAGAAAAAAGTGTATACCTCACTTAGAGCGTCATTAAGTTAGAGAACAACTCCCACACTCTGCTCTAAGAT GAGGCCAACATCCAGATTCTCACTGCAAACTTGGCTCCCCTAGAATTCTGTTCTTCTTTCCCCTGTCCCCGCTGC TTTCTAAACGTGAAATATCCACAGCTGCACCGTTTCTTTACTTTTTTTTTTTTTTCTGGAGAGGTTAACCTTCCTT CTTCAGTTGTATGTTGTGCAATACCCAAGAGGCCAACACTACCATTTGAGATATTTAAATATGACTTTGAGAAA TGAGTCTGTCTTAGATATAAAAGCCCACAGTC    SEQ ID NO: 32  SLC34A2‐2 enhancer reverse (403bp)  GACTGTGGGCTTTTATATCTAAGACAGACTCATTTCTCAAAGTCATATTTAAATATCTCAAATGGTAGTGTTGG CCTCTTGGGTATTGCACAACATACAACTGAAGAAGGAAGGTTAACCTCTCCAGAAAAAAAAAAAAAAGTAAA GAAACGGTGCAGCTGTGGATATTTCACGTTTAGAAAGCAGCGGGGACAGGGGAAAGAAGAACAGAATTCTA GGGGAGCCAAGTTTGCAGTGAGAATCTGGATGTTGGCCTCATCTTAGAGCAGAGTGTGGGAGTTGTTCTCTA ACTTAATGACGCTCTAAGTGAGGTATACACTTTTTTCTCAGTTCCGTGTTTTATAGGCAAGATTTCTCTTCTGTT CCTTTCAATAATCTAAACTGACAGTCTGAATGGTTAGCT    SEQ ID NO: 33  SLC34A2‐3 enhancer reverse (705bp)  TCAGGAGGCTGAGGTGGGAGGATCACTTGAGCCCAGGAGTTCGAGACTGCAGTGAGCTGTAATCACACTACT GCATTCCAGCTAGGGTGACAGTCTCATAAGGGAAAAAAAAAAAAGATTTAAAAAACCTCTCTCCAGCATTGAA ACTTCCTGTAGGCTAGTTGACCAAATCTCAACATTAGCTGCCAGGTTATGTGTGTCCAAGCAATGCAGGTTGTG GTTTTAATTAGCATGATGTCTCCAAGGATTGTGTCTGTCCTCCTGTGTATGGCGATTCCAATTTACTGGTTTTAA GATGTGGTGAACAAAAGTCTGTTCTTGTAATTGACCAAGGGGAAGTAGGCCACGAAGGTGTGAATCTTAAGT ACGTGACCCCCTAACTCACGCTGAAGCCATGCTCACTGCATCTGAGCAGGTGCCGCTTTCGCTTCTTGTTTTTTA TTTTTAAATTTCCTAAGTTTTGTTTGGTTTTTACCTTCCCCTGATTTGGCAGATGCGTTTGTACTTTTTGGAGAGA TATTTATACCATTCGTGTTAAAGGTATTTTTTTTTTAAATGGTGTCCTGTCAACAGTACATTGAGGATACAATTT AGGAACTCAAAATGTGTTTAAACACAACTTAGCAAATGTTATGTCTCATCATTTGGCATTTGGCCTCTGCAGAT TCATTTCATTTAATGCCATAACCAATATTCCTGATTTGA    SEQ ID NO: 34 SV40 forward enhancer (237bp)  CGATGGAGCGGAGAATGGGCGGAACTGGGCGGAGTTAGGGGCGGGATGGGCGGAGTTAGGGGCGGGACT ATGGTTGCTGACTAATTGAGATGCATGCTTTGCATACTTCTGCCTGCTGGGGAGCCTGGGGACTTTCCACACCT GGTTGCTGACTAATTGAGATGCATGCTTTGCATACTTCTGCCTGCTGGGGAGCCTGGGGACTTTCCACACCCTA ACTGACACACATTCCACAGC    Also referred to interchangeably as SV40 Forward in the Examples/Figures herein.    SEQ ID NO: 35 SV40 reverse enhancer (237bp)  GCTGTGGAATGTGTGTCAGTTAGGGTGTGGAAAGTCCCCAGGCTCCCCAGCAGGCAGAAGTATGCAAAGCAT GCATCTCAATTAGTCAGCAACCAGGTGTGGAAAGTCCCCAGGCTCCCCAGCAGGCAGAAGTATGCAAAGCAT GCATCTCAATTAGTCAGCAACCATAGTCCCGCCCCTAACTCCGCCCATCCCGCCCCTAACTCCGCCCAGTTCCGC CCATTCTCCGCTCCATCG    SEQ ID NO: 36  VEGFA‐1 enhancer forward (650bp)  AATTGGGTATGTAGATCATCAGGGTGGAAGGTGAGCTTCCTTTCTAGGTGGGGAAACAGGCTCAGAGAGGG GCCATGACTTACTGGTGTTTCACAATGAGCCAGTGGTGGAGCTCACAGATCCTGGCCCCAGACCAGGCACCCT CCTCCCCCTAACCCCCATCTCCAGCTCTCAATGCCAAGAGGCTGGAGCGGGCCCTGTGCACTCCAGGCAAGGC TGCGGTTTCTTCTGTGGGACTGTGGAGCCACCCAGGGTGTGTGGGTGCTGCGCACGCGCACCACAGTTGTGT AACAGGCTCTGGGGCGTGTGCACGCGCTCAGTAACCTGATGCAACAGGGAAGCTGTGGTCCCGTGAGCTGG GGTGTGGGCGCTTGGCCGGTGCCCAAACCACAAAGAGTATTTGGCTTGTGGCTAAGGAGGCCAGGCATGGG GGTCGCAGAGTTGTCTACGTCAGGGCTTGTCTATCCCTAGTTGCCCACATCCTGGTCCCTGTGGGCGCAGGGC TGGGAATGGGGCTGCCTGGCCAGGTGATACTGTGGAGGAGGGGTTGGGGCCTTGCCCCTTCCTCCTTCTGTTT GCCACGATGCCTATTGTGATGAGGATGGGGGGCAGGCGGGGCTTCCTCCCTGTGGTTCCATCTCTTCCTGGCT    SEQ ID NO: 37  VEGFA‐1 enhancer reverse (650bp)  AGCCAGGAAGAGATGGAACCACAGGGAGGAAGCCCCGCCTGCCCCCCATCCTCATCACAATAGGCATCGTGG CAAACAGAAGGAGGAAGGGGCAAGGCCCCAACCCCTCCTCCACAGTATCACCTGGCCAGGCAGCCCCATTCC CAGCCCTGCGCCCACAGGGACCAGGATGTGGGCAACTAGGGATAGACAAGCCCTGACGTAGACAACTCTGCG ACCCCCATGCCTGGCCTCCTTAGCCACAAGCCAAATACTCTTTGTGGTTTGGGCACCGGCCAAGCGCCCACACC CCAGCTCACGGGACCACAGCTTCCCTGTTGCATCAGGTTACTGAGCGCGTGCACACGCCCCAGAGCCTGTTAC ACAACTGTGGTGCGCGTGCGCAGCACCCACACACCCTGGGTGGCTCCACAGTCCCACAGAAGAAACCGCAGC CTTGCCTGGAGTGCACAGGGCCCGCTCCAGCCTCTTGGCATTGAGAGCTGGAGATGGGGGTTAGGGGGAGG AGGGTGCCTGGTCTGGGGCCAGGATCTGTGAGCTCCACCACTGGCTCATTGTGAAACACCAGTAAGTCATGG CCCCTCTCTGAGCCTGTTTCCCCACCTAGAAAGGAAGCTCACCTTCCACCCTGATGATCTACATACCCAATT    SEQ ID NO: 38 VEGFA‐2 forward enhancer (680bp)  GAGGCACAAAGCGATCCCCATCACTGCTCCACAATCATTCATTAGCTAACAAGACAGAGCAGCTCATAAAAAA AAAGCCGTTAAAAAAATTCCGGGGAAATGGAAAGCAGGAGGTGATGCAAGCCCTGGTTAACAAAGGCTGAG GGTTGGGGGGAGGCATGAGAGGGTGTGAGTGGAATAACCCAAGCCTGATAAGCCACAAAGCAGCGCCTGCT GCTCCCTCCTCCCTCTGCCGCCTGAGTCAGAGAAGCCGGGATGTGTTCAAAATCAAGCAATGTAATTAGCAGG TTGCATCATGCCTCTCGATTATAAATTACAATGGCCACAAAGAGGGTGGATGGAGGAGCAGGGATTAGATGG GGTAACCCAATCCCAGTGCCGGGAAGGGGGACAGAGGGCCCGGGAGCTGCCCCCCGCCACTAGGAGCCACT CGGAGTGGCCTCTGTCTACTGTGTCTAGGGACAGCAACAAGGCCTACTCAATACCTGCTTTGTGTAGCCTGGG ATAATGATTTACAAACTCAAGGAAATCAGGAGGTGTGAAAAGAAGATAGCCCTCTTAACCCAGCCATTGTCTA CCCTCTGCTGGAATGTCCCTCAGCTCCACCCACCTGGTGCACCTGTGCACACGGATGCTGTTTGAGGATGGTCC ACGTGAGGAGCCCTCATAGGGGCTGGACT    SEQ ID NO: 39  VEGFA‐3 enhancer forward (801bp)  GGAGGCTTGGCCTGCAGTTTGAGGATCAAACTGTCTCCTGGAAACTACCCCCATGCCCTGGACCCAGCCTCTG TGGCCCAGGACAGTAGTGGGGACAGGTTAAGGCTAGGGAGTGCCTTTGCCAGGCACTGGGAAAAGACACTG CAGGCCCCAAATTGTGGTTAATCTCCAAGGACACGCAAGAATGCAGCAGGAGAGACTAGGATTAGACAACAG GGAGAACTCGATGACAGCTAGAGCTCCGGGATTCATTCCCTGCAAGGGCAGGGAGGTTGAGGGGGCAGAGG AGGAAGGAATCTGCAAGGAGGTTGCTGACGTCCTCCTGGAAGGCTAGAAATGAAGCACAGGCAGCCGAGTC TTCTCCCACCTCCCTCCACCCCCAGCGATGAGTCAGGGATTCCGCAGAATGCTGAATTCCTGTCTTCCTTTGGG CAATGATACTTGACACTCTCCTAGGTTTTCTCTCCTGCCACAGAAGTAAAATCTTCAACCATTCTCAAGGACTTC TTTCTCCTCAGCTCCCCAGACAGGCCCCTGCTTTCCTGTGAGTCACTGTATCTTCAAAGCCCCTTCCTGCCTCAG CCTACCCCAATAGCTCTCAGGCTCCTGGCTTTGCACGTGTTGTTCTGTCTGCCTGGAACACTCTTCCCAGAACTA GCTTCTTCTCATCCACCAGAAACCAGTTTAGATGGCACCTGCTGAAACAGATCAGCTCACAAAGGCCCGGCAC AACCTGGCCCACAGAAGCATTCCTGAGTGCTGTAAATGAATGAACTGTGAATCAATGAATGCAGAAGATTC    SEQ ID NO: 40  VEGFA‐3 enhancer reverse (801bp)  GAATCTTCTGCATTCATTGATTCACAGTTCATTCATTTACAGCACTCAGGAATGCTTCTGTGGGCCAGGTTGTG CCGGGCCTTTGTGAGCTGATCTGTTTCAGCAGGTGCCATCTAAACTGGTTTCTGGTGGATGAGAAGAAGCTAG TTCTGGGAAGAGTGTTCCAGGCAGACAGAACAACACGTGCAAAGCCAGGAGCCTGAGAGCTATTGGGGTAG GCTGAGGCAGGAAGGGGCTTTGAAGATACAGTGACTCACAGGAAAGCAGGGGCCTGTCTGGGGAGCTGAG GAGAAAGAAGTCCTTGAGAATGGTTGAAGATTTTACTTCTGTGGCAGGAGAGAAAACCTAGGAGAGTGTCAA GTATCATTGCCCAAAGGAAGACAGGAATTCAGCATTCTGCGGAATCCCTGACTCATCGCTGGGGGTGGAGGG AGGTGGGAGAAGACTCGGCTGCCTGTGCTTCATTTCTAGCCTTCCAGGAGGACGTCAGCAACCTCCTTGCAGA TTCCTTCCTCCTCTGCCCCCTCAACCTCCCTGCCCTTGCAGGGAATGAATCCCGGAGCTCTAGCTGTCATCGAGT TCTCCCTGTTGTCTAATCCTAGTCTCTCCTGCTGCATTCTTGCGTGTCCTTGGAGATTAACCACAATTTGGGGCC TGCAGTGTCTTTTCCCAGTGCCTGGCAAAGGCACTCCCTAGCCTTAACCTGTCCCCACTACTGTCCTGGGCCAC AGAGGCTGGGTCCAGGGCATGGGGGTAGTTTCCAGGAGACAGTTTGATCCTCAAACTGCAGGCCAAGCCTCC    SEQ ID NO: 41 SLC34A2‐2 Enhancer & core SFTPB Promoter (1039bp)*  agctaaccat tcagactgtc agtttagatt attgaaagga acagaagaga aatcttgcct 60 ataaaacacg gaactgagaa aaaagtgtat acctcactta gagcgtcatt aagttagaga 120 acaactccca cactctgctc taagatgagg ccaacatcca gattctcact gcaaacttgg 180 ctcccctaga attctgttct tctttcccct gtccccgctg ctttctaaac gtgaaatatc 240 cacagctgca ccgtttcttt actttttttt ttttttctgg agaggttaac cttccttctt 300 cagttgtatg ttgtgcaata cccaagaggc caacactacc atttgagata tttaaatatg 360 actttgagaa atgagtctgt cttagatata aaagcccaca gtcagatctt atagggctgt 420 ctgggagcca ctccagggcc acagaaatct tgtctctgac tcagggtatt ttgttttctg 480 ttttgtgtaa atgctcttct gactaatgca aaccatgtgt ccatagaacc agaagatttt 540 tccaggggaa aaggtaagga ggtggtgaga gtgtcctggg tctgcccttc cagggcttgc 600 cctgggttaa gagccaggca ggaagctctc aagagcattg ctcaagagta gagggggcct 660 gggaggccca gggaggggat gggaggggaa cacccaggct gcccccaacc agatgccctc 720 caccctcctc aacctccctc ccacggcctg gagaggtggg accaggtatg gaggcttgag 780 agcccctggt tggaggaagc cacaagtcca ggaacatggg agtctgggca gggggcaaag 840 gaggcaggaa caggccatca gccaggacag gtggtaaggc aggcaggagt gttcctgctg 900 ggaaaaggtg ggatcaagca cctggagggc tcttcagagc aaagacaaac actgaggtcg 960 ctgccactcc tacagagccc ccacgccccg cccagctata aggggccatg cmccaagcag 1020 ggtacccagg ctgcagagg 1039   *wherein M is C or A, preferably M is C  Bases in bold and underlined represent a BglII restriction enzyme site created to operably link the  enhancer and promoter sequences.  For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in  which the underlined BglII restriction site is omitted (SEQ ID NO: 76).    SEQ ID NO: 42 VEGFA‐2 Enhancer & core SFTPB Promoter (1316bp)*  gaggcacaaa gcgatcccca tcactgctcc acaatcattc attagctaac aagacagagc 60 agctcataaa aaaaaagccg ttaaaaaaat tccggggaaa tggaaagcag gaggtgatgc 120 aagccctggt taacaaaggc tgagggttgg ggggaggcat gagagggtgt gagtggaata 180 acccaagcct gataagccac aaagcagcgc ctgctgctcc ctcctccctc tgccgcctga 240 gtcagagaag ccgggatgtg ttcaaaatca agcaatgtaa ttagcaggtt gcatcatgcc 300 tctcgattat aaattacaat ggccacaaag agggtggatg gaggagcagg gattagatgg 360 ggtaacccaa tcccagtgcc gggaaggggg acagagggcc cgggagctgc cccccgccac 420 taggagccac tcggagtggc ctctgtctac tgtgtctagg gacagcaaca aggcctactc 480 aatacctgct ttgtgtagcc tgggataatg atttacaaac tcaaggaaat caggaggtgt 540 gaaaagaaga tagccctctt aacccagcca ttgtctaccc tctgctggaa tgtccctcag 600 ctccacccac ctggtgcacc tgtgcacacg gatgctgttt gaggatggtc cacgtgagga 660 gccctcatag gggctggact agatcttata gggctgtctg ggagccactc cagggccaca 720 gaaatcttgt ctctgactca gggtattttg ttttctgttt tgtgtaaatg ctcttctgac 780 taatgcaaac catgtgtcca tagaaccaga agatttttcc aggggaaaag gtaaggaggt 840 ggtgagagtg tcctgggtct gcccttccag ggcttgccct gggttaagag ccaggcagga 900 agctctcaag agcattgctc aagagtagag ggggcctggg aggcccaggg aggggatggg 960 aggggaacac ccaggctgcc cccaaccaga tgccctccac cctcctcaac ctccctccca 1020 cggcctggag aggtgggacc aggtatggag gcttgagagc ccctggttgg aggaagccac 1080 aagtccagga acatgggagt ctgggcaggg ggcaaaggag gcaggaacag gccatcagcc 1140 aggacaggtg gtaaggcagg caggagtgtt cctgctggga aaaggtggga tcaagcacct 1200 ggagggctct tcagagcaaa gacaaacact gaggtcgctg ccactcctac agagccccca 1260 cgccccgccc agctataagg ggccatgcmc caagcagggt acccaggctg cagagg 1316   *wherein M is C or A, preferably M is C  Bases in bold and underlined represent a BglII restriction enzyme site created to operably link the  enhancer and promoter sequences.  For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in  which the underlined BglII restriction site is omitted (SEQ ID NO: 77).    SEQ ID NO: 43 CpG‐Free CMV Enhancer & core SFTPB Promoter (938bp)*  gttacataac ttatggtaaa tggcctgcct ggctgactgc ccaatgaccc ctgcccaatg 60 atgtcaataa tgatgtatgt tcccatgtaa tgccaatagg gactttccat tgatgtcaat 120 gggtggagta tttatggtaa ctgcccactt ggcagtacat caagtgtatc atatgccaag 180 tatgccccct attgatgtca atgatggtaa atggcctgcc tggcattatg cccagtacat 240 gaccttatgg gactttccta cttggcagta catctatgta ttagtcattg ctattaccat 300 ggagatctta tagggctgtc tgggagccac tccagggcca cagaaatctt gtctctgact 360 cagggtattt tgttttctgt tttgtgtaaa tgctcttctg actaatgcaa accatgtgtc 420 catagaacca gaagattttt ccaggggaaa aggtaaggag gtggtgagag tgtcctgggt 480 ctgcccttcc agggcttgcc ctgggttaag agccaggcag gaagctctca agagcattgc 540 tcaagagtag agggggcctg ggaggcccag ggaggggatg ggaggggaac acccaggctg 600 cccccaacca gatgccctcc accctcctca acctccctcc cacggcctgg agaggtggga 660 ccaggtatgg aggcttgaga gcccctggtt ggaggaagcc acaagtccag gaacatggga 720 gtctgggcag ggggcaaagg aggcaggaac aggccatcag ccaggacagg tggtaaggca 780 ggcaggagtg ttcctgctgg gaaaaggtgg gatcaagcac ctggagggct cttcagagca 840 aagacaaaca ctgaggtcgc tgccactcct acagagcccc cacgccccgc ccagctataa 900 ggggccatgc mccaagcagg gtacccaggc tgcagagg 938   *wherein M is C or A, preferably M is C  Bases in bold and underlined represent a BglII restriction enzyme site created to operably link the  enhancer and promoter sequences.  For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in  which the underlined BglII restriction site is omitted (SEQ ID NO: 78).    SEQ ID NO: 44 ELF3 Enhancer & core SFTPB Promoter (1290bp)*  tagggatggg ccgaggctgg cactgatgct agacttccgt gcacagggca agtatggaca 60 agccccaagt ggctttgtga ggcccacaca gtgaagcttg ggaaatggga agtggggctg 120 cgcccagatt ctggtatcta tgacaactaa ggccgctgca catcctcatg gctctcccag 180 agacctcagg tgaggccctt ctgtgttcct caagcaccca tgccacctgc ggggtggggc 240 aggaccctcc tacccagccc tgggcctcct ggggaaccat gggtgcacag gggtaacctg 300 agccagcctc tctggggcat gggcgggcgt ggtggcgtgg gcctggcgcc agagagtgga 360 gcagaatgtc agctctgtga gccgcaccgg gtgccagcac tctgcaaaca gactctagtc 420 accagataga ctggagtcac gaacctaaca aagcgctcag ctgggcaact gtaactgcag 480 agggcggggc cgcacagtgc tgcctagtgc ctcctgcctt gatctggtca gggtcatggg 540 aggagacagg ttccccagga gggcccatta acaagattaa ttggggtggc ctgggggact 600 ctgggatgct cactggaaac atatcctggg ggaatgaggt aggtggggag ccagagatct 660 tatagggctg tctgggagcc actccagggc cacagaaatc ttgtctctga ctcagggtat 720 tttgttttct gttttgtgta aatgctcttc tgactaatgc aaaccatgtg tccatagaac 780 cagaagattt ttccagggga aaaggtaagg aggtggtgag agtgtcctgg gtctgccctt 840 ccagggcttg ccctgggtta agagccaggc aggaagctct caagagcatt gctcaagagt 900 agagggggcc tgggaggccc agggagggga tgggagggga acacccaggc tgcccccaac 960 cagatgccct ccaccctcct caacctccct cccacggcct ggagaggtgg gaccaggtat 1020 ggaggcttga gagcccctgg ttggaggaag ccacaagtcc aggaacatgg gagtctgggc 1080 agggggcaaa ggaggcagga acaggccatc agccaggaca ggtggtaagg caggcaggag 1140 tgttcctgct gggaaaaggt gggatcaagc acctggaggg ctcttcagag caaagacaaa 1200 cactgaggtc gctgccactc ctacagagcc cccacgcccc gcccagctat aaggggccat 1260 gcmccaagca gggtacccag gctgcagagg 1290   *wherein M is C or A, preferably M is C  Bases in bold and underlined represent a BglII restriction enzyme site created to operably link the  enhancer and promoter sequences.  For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in  which the underlined BglII restriction site is omitted (SEQ ID NO: 79).    SEQ ID NO: 45  SV40 Enhancer & mSPB Promoter (873bp)*  cgatggagcg gagaatgggc ggaactgggc ggagttaggg gcgggatggg cggagttagg 60 ggcgggacta tggttgctga ctaattgaga tgcatgcttt gcatacttct gcctgctggg 120 gagcctgggg actttccaca cctggttgct gactaattga gatgcatgct ttgcatactt 180 ctgcctgctg gggagcctgg ggactttcca caccctaact gacacacatt ccacagcaga 240 tcttataggg ctgtctggga gccactccag ggccacagaa atcttgtctc tgactcaggg 300 tattttgttt tctgttttgt gtaaatgctc ttctgactaa tgcaaaccat gtgtccatag 360 aaccagaaga tttttccagg ggaaaaggta aggaggtggt gagagtgtcc tgggtctgcc 420 cttccagggc ttgccctggg ttaagagcca ggcaggaagc tctcaagagc attgctcaag 480 agtagagggg gcctgggagg cccagggagg ggatgggagg ggaacaccca ggctgccccc 540 aaccagatgc cctccaccct cctcaacctc cctcccacgg cctggagagg tgggaccagg 600 tatggaggct tgagagcccc tggttggagg aagccacaag tccaggaaca tgggagtctg 660 ggcagggggc aaaggaggca ggaacaggcc atcagccagg acaggtggta aggcaggcag 720 gagtgttcct gctgggaaaa ggtgggatca agcacctgga gggctcttca gagcaaagac 780 aaacactgag gtcgctgcca ctcctacaga gcccccacgc cccgcccagc tataaggggc 840 catgcmccaa gcagggtacc caggctgcag agg 873   *wherein M is C or A, preferably M is C  Bases in bold and underlined represent a BglII restriction enzyme created to operably link the  enhancer and promoter sequences.  For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in  which the underlined BglII restriction site is omitted (SEQ ID NO: 80).    SEQ ID NO: 46 Alv‐01 (CMV forward enhancer + mSFPB promoter) (955bp)  AGATCTGTTACATAACTTATGGTAAATGGCCTGCCTGGCTGACTGCCCAATGACCCCTGCCCAATGATGTCAAT AATGATGTATGTTCCCATGTAATGCCAATAGGGACTTTCCATTGATGTCAATGGGTGGAGTATTTATGGTAACT GCCCACTTGGCAGTACATCAAGTGTATCATATGCCAAGTATGCCCCCTATTGATGTCAATGATGGTAAATGGCC TGCCTGGCATTATGCCCAGTACATGACCTTATGGGACTTTCCTACTTGGCAGTACATCTATGTATTAGTCATTGC TATTACCATGGAGATCTTATAGGGCTGTCTGGGAGCCACTCCAGGGCCACAGAAATCTTGTCTCTGACTCAGG GTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTAATGCAAACCATGTGTCCATAGAACCAGAAGATTTT TCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCTGGGTCTGCCCTTCCAGGGCTTGCCCTGGGTTAAGA GCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGTAGAGGGGGCCTGGGAGGCCCAGGGAGGGGATGG GAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCTCCACCCTCCTCAACCTCCCTCCCACGGCCTGGAGAG GTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTGGAGGAAGCCACAAGTCCAGGAACATGGGAGTCTG GGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGCCAGGACAGGTGGTAAGGCAGGCAGGAGTGTTCCT GCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCTTCAGAGCAAAGACAAACACTGAGGTCGCTGCCAC TCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGGCCATGCMCCAAGCAGGGTACCCAGGCTGCAGAG GGTGCCGCTAGC    *wherein M is C or A, preferably M is C (C is present in the exemplified Alv‐01)    The bold and underlined 3’ sequence is an NheI restriction site.    For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in  which this Nhe1 restriction site is omitted (SEQ ID NO: 81).    SEQ ID NO: 47 Alv‐02 (SLC34A2‐2 forward enhancer + mSFPB promoter) (1056bp)  AGATCTAGCTAACCATTCAGACTGTCAGTTTAGATTATTGAAAGGAACAGAAGAGAAATCTTGCCTATAAAAC ACGGAACTGAGAAAAAAGTGTATACCTCACTTAGAGCGTCATTAAGTTAGAGAACAACTCCCACACTCTGCTC TAAGATGAGGCCAACATCCAGATTCTCACTGCAAACTTGGCTCCCCTAGAATTCTGTTCTTCTTTCCCCTGTCCC CGCTGCTTTCTAAACGTGAAATATCCACAGCTGCACCGTTTCTTTACTTTTTTTTTTTTTTCTGGAGAGGTTAACC TTCCTTCTTCAGTTGTATGTTGTGCAATACCCAAGAGGCCAACACTACCATTTGAGATATTTAAATATGACTTTG AGAAATGAGTCTGTCTTAGATATAAAAGCCCACAGTCAGATCTTATAGGGCTGTCTGGGAGCCACTCCAGGGC CACAGAAATCTTGTCTCTGACTCAGGGTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTAATGCAAACC ATGTGTCCATAGAACCAGAAGATTTTTCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCTGGGTCTGCC CTTCCAGGGCTTGCCCTGGGTTAAGAGCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGTAGAGGGGGC CTGGGAGGCCCAGGGAGGGGATGGGAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCTCCACCCTCCT CAACCTCCCTCCCACGGCCTGGAGAGGTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTGGAGGAAGC CACAAGTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGCCAGGACAG GTGGTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCTTCAGAGCA AAGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGGCCATGCMC CAAGCAGGGTACCCAGGCTGCAGAGGGTGCCGCTAGC    *wherein M is C or A, preferably M is C (C is present in the exemplified Alv‐02)    The bold and underlined 3’ sequence is an NheI restriction site.    For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in  which this Nhe1 restriction site is omitted (SEQ ID NO: 82).    SEQ ID NO: 48 Alv‐03 (VEGFA‐2 forward enhancer + mSFPB promoter) (1333bp)  AGATCTGAGGCACAAAGCGATCCCCATCACTGCTCCACAATCATTCATTAGCTAACAAGACAGAGCAGCTCAT AAAAAAAAAGCCGTTAAAAAAATTCCGGGGAAATGGAAAGCAGGAGGTGATGCAAGCCCTGGTTAACAAAG GCTGAGGGTTGGGGGGAGGCATGAGAGGGTGTGAGTGGAATAACCCAAGCCTGATAAGCCACAAAGCAGC GCCTGCTGCTCCCTCCTCCCTCTGCCGCCTGAGTCAGAGAAGCCGGGATGTGTTCAAAATCAAGCAATGTAATT AGCAGGTTGCATCATGCCTCTCGATTATAAATTACAATGGCCACAAAGAGGGTGGATGGAGGAGCAGGGATT AGATGGGGTAACCCAATCCCAGTGCCGGGAAGGGGGACAGAGGGCCCGGGAGCTGCCCCCCGCCACTAGGA GCCACTCGGAGTGGCCTCTGTCTACTGTGTCTAGGGACAGCAACAAGGCCTACTCAATACCTGCTTTGTGTAG CCTGGGATAATGATTTACAAACTCAAGGAAATCAGGAGGTGTGAAAAGAAGATAGCCCTCTTAACCCAGCCAT TGTCTACCCTCTGCTGGAATGTCCCTCAGCTCCACCCACCTGGTGCACCTGTGCACACGGATGCTGTTTGAGGA TGGTCCACGTGAGGAGCCCTCATAGGGGCTGGACTAGATCTTATAGGGCTGTCTGGGAGCCACTCCAGGGCC ACAGAAATCTTGTCTCTGACTCAGGGTATTTTGTTTTCTGTTTTGTGTAAATGCTCTTCTGACTAATGCAAACCA TGTGTCCATAGAACCAGAAGATTTTTCCAGGGGAAAAGGTAAGGAGGTGGTGAGAGTGTCCTGGGTCTGCCC TTCCAGGGCTTGCCCTGGGTTAAGAGCCAGGCAGGAAGCTCTCAAGAGCATTGCTCAAGAGTAGAGGGGGCC TGGGAGGCCCAGGGAGGGGATGGGAGGGGAACACCCAGGCTGCCCCCAACCAGATGCCCTCCACCCTCCTC AACCTCCCTCCCACGGCCTGGAGAGGTGGGACCAGGTATGGAGGCTTGAGAGCCCCTGGTTGGAGGAAGCC ACAAGTCCAGGAACATGGGAGTCTGGGCAGGGGGCAAAGGAGGCAGGAACAGGCCATCAGCCAGGACAGG TGGTAAGGCAGGCAGGAGTGTTCCTGCTGGGAAAAGGTGGGATCAAGCACCTGGAGGGCTCTTCAGAGCAA AGACAAACACTGAGGTCGCTGCCACTCCTACAGAGCCCCCACGCCCCGCCCAGCTATAAGGGGCCATGCMCC AAGCAGGGTACCCAGGCTGCAGAGGGTGCCGCTAGC    *wherein M is C or A, preferably M is C (C is present in the exemplified Alv‐03)    The bold and underlined 3’ sequence is an NheI restriction site.    For the avoidance of doubt, the invention also expressly encompasses a version of this sequence in  which this Nhe1 restriction site is omitted (SEQ ID NO: 83).    SEQ ID NO: 49  ATTTGAGCTCTTCTTTCTGCTGAACCATCG Underlined sequence is an SacI restriction enzyme site    SEQ ID NO: 50  SFTPB promoter reverse primer  TCTTAGATCTGTCAGACAGCTCTGGGTTCC Underlined sequence is a BgIII restriction enzyme site    SEQ ID NO: 51  Exemplary linker between SFTPB promoter fragment and enhancer   agatct  SEQ ID NO: 52 Exemplified WPRE component (mWPRE)  1 GGGCCCAATC AACCTCTGGA TTACAAAATT TGTGAAAGAT TGACTGGTAT TCTTAACTAT 61 GTTGCTCCTT TTACGCTATG TGGATACGCT GCTTTAATGC CTTTGTATCA TGCTATTGCT 121 TCCCGTATGG CTTTCATTTT CTCCTCCTTG TATAAATCCT GGTTGCTGTC TCTTTATGAG 181 GAGTTGTGGC CCGTTGTCAG GCAACGTGGC GTGGTGTGCA CTGTGTTTGC TGACGCAACC 241 CCCACTGGTT GGGGCATTGC CACCACCTGT CAGCTCCTTT CCGGGACTTT CGCTTTCCCC 301 CTCCCTATTG CCACGGCGGA ACTCATCGCC GCCTGCCTTG CCCGCTGCTG GACAGGGGCT 361 CGGCTGTTGG GCACTGACAA TTCCGTGGTG TTGTCGGGGA AATCATCGTC CTTTCCTTGG 421 CTGCTCGCCT GTGTTGCCAC CTGGATTCTG CGCGGGACGT CCTTCTGCTA CGTCCCTTCG 481 GCCCTCAATC CAGCGGACCT TCCTTCCCGC GGCCTGCTGC CGGCTCTGCG GCCTCTTCCG 541 CGTCTTCGCC TTCGCCCTCA GACGAGTCGG ATCTCCCTTT GGGCCGCCTC CCCGCAAGCT   SEQ ID NO: 53 Exemplified SFTPB transgene (NM_000542.5) (1146bp)  ATGGCTGAGTCACACCTGCTGCAGTGGCTGCTGCTGCTGCTGCCCACGCTCTGTGGCCCAGGCACTGCTGCCT GGACCACCTCATCCTTGGCCTGTGCCCAGGGCCCTGAGTTCTGGTGCCAAAGCCTGGAGCAAGCATTGCAGTG CAGAGCCCTAGGGCATTGCCTACAGGAAGTCTGGGGACATGTGGGAGCCGATGACCTATGCCAAGAGTGTG AGGACATCGTCCACATCCTTAACAAGATGGCCAAGGAGGCCATTTTCCAGGACACGATGAGGAAGTTCCTGG AGCAGGAGTGCAACGTCCTCCCCTTGAAGCTGCTCATGCCCCAGTGCAACCAAGTGCTTGACGACTACTTCCC CCTGGTCATCGACTACTTCCAGAACCAGACTGACTCAAACGGCATCTGTATGCACCTGGGCCTGTGCAAATCCC GGCAGCCAGAGCCAGAGCAGGAGCCAGGGATGTCAGACCCCCTGCCCAAACCTCTGCGGGACCCTCTGCCAG ACCCTCTGCTGGACAAGCTCGTCCTCCCTGTGCTGCCCGGGGCCCTCCAGGCGAGGCCTGGGCCTCACACACA GGATCTCTCCGAGCAGCAATTCCCCATTCCTCTCCCCTATTGCTGGCTCTGCAGGGCTCTGATCAAGCGGATCC AAGCCATGATTCCCAAGGGTGCGCTAGCTGTGGCAGTGGCCCAGGTGTGCCGCGTGGTACCTCTGGTGGCGG GCGGCATCTGCCAGTGCCTGGCTGAGCGCTACTCCGTCATCCTGCTCGACACGCTGCTGGGCCGCATGCTGCC CCAGCTGGTCTGCCGCCTCGTCCTCCGGTGCTCCATGGATGACAGCGCTGGCCCAAGGTCGCCGACAGGAGA ATGGCTGCCGCGAGACTCTGAGTGCCACCTCTGCATGTCCGTGACCACCCAGGCCGGGAACAGCAGCGAGCA GGCCATACCACAGGCAATGCTCCAGGCCTGTGTTGGCTCCTGGCTGGACAGGGAAAAGTGCAAGCAATTTGT GGAGCAGCACACGCCCCAGCTGCTGACCCTGGTGCCCAGGGGCTGGGATGCCCACACCACCTGCCAGGCCCT CGGGGTGTGTGGGACCATGTCCAGCCCTCTCCAGTGTATCCACAGCCCCGACCTTTGA    SEQ ID NO: 54 Exemplified human SFTPB (hSP‐B) transgene (1162bp)  GCTAGCCACCATGGCTGAGTCACACCTGCTGCAGTGGCTGCTGCTGCTGCTGCCCACGCTCTGTGGCCCAGGC ACTGCTGCCTGGACCACCTCATCCTTGGCCTGTGCCCAGGGACCTGAGTTCTGGTGCCAAAGCCTGGAGCAAG CATTGCAGTGCAGAGCCCTAGGGCATTGCCTACAGGAAGTCTGGGGACATGTGGGAGCCGATGACCTATGCC AAGAGTGTGAGGACATCGTCCACATCCTTAACAAGATGGCCAAGGAGGCCATTTTCCAGGACACGATGAGGA AGTTCCTGGAGCAGGAGTGCAACGTCCTCCCCTTGAAGCTGCTCATGCCCCAGTGCAACCAAGTGCTTGACGA CTACTTCCCCCTGGTCATCGACTACTTCCAGAACCAGACTGACTCAAACGGCATCTGTATGCACCTGGGCCTGT GCAAATCCCGGCAGCCAGAGCCAGAGCAGGAGCCAGGGATGTCAGACCCCCTGCCCAAACCTCTGCGGGACC CTCTGCCAGACCCTCTGCTGGACAAGCTCGTCCTCCCTGTGCTGCCCGGGGCTCTCCAGGCGAGGCCTGGGCC TCACACACAGGATCTCTCCGAGCAGCAATTCCCCATTCCTCTCCCCTATTGCTGGCTCTGCAGGGCTCTGATCA AGCGGATCCAAGCCATGATTCCCAAGGGTGCCCTAGCTGTGGCAGTGGCCCAGGTGTGCCGCGTGGTGCCTC TGGTGGCGGGCGGCATCTGCCAGTGCCTGGCTGAGCGCTACTCCGTCATCCTGCTCGACACGCTGCTGGGCC GCATGCTGCCCCAGCTGGTCTGCCGCCTCGTCCTCCGGTGCTCCATGGATGACAGCGCTGGCCCAAGGTCGCC GACAGGAGAATGGCTGCCGCGAGACTCTGAGTGCCACCTCTGCATGTCCGTGACCACCCAGGCCGGGAACAG CAGCGAGCAGGCCATACCACAGGCAATGCTCCAGGCCTGTGTTGGCTCCTGGCTGGACAGGGAAAAGTGCAA GCAATTTGTGGAGCAGCACACGCCCCAGCTGCTGACCCTGGTGCCCAGGGGCTGGGATGCCCACACCACCTG CCAGGCCCTCGGGGTGTGTGGGACCATGTCCAGCCCTCTCCAGTGTATCCACAGCCCCGACCTTTGAGGGCCC    SEQ ID NO: 55 Exemplified codon‐optimised human SFTPB (cohSP‐B) transgene (1162bp)  GCTAGCCACCATGGCCGAGAGTCACCTGCTTCAGTGGCTCTTGTTGCTCTTGCCAACGCTCTGTGGTCCAGGA ACAGCGGCCTGGACAACTTCATCACTGGCTTGTGCTCAGGGACCCGAGTTTTGGTGCCAATCCCTGGAGCAGG CTCTTCAGTGTCGGGCTCTGGGACACTGCTTGCAGGAGGTATGGGGACACGTCGGCGCCGATGACCTTTGCC AGGAATGTGAAGATATAGTCCACATTCTGAATAAGATGGCCAAGGAAGCTATTTTTCAGGATACGATGCGCA AGTTTTTGGAACAAGAGTGCAACGTGCTCCCTCTTAAACTCCTTATGCCGCAGTGTAATCAGGTCCTTGATGAC TACTTTCCCCTGGTTATCGACTATTTCCAGAATCAGACAGACTCCAATGGAATCTGCATGCACCTGGGTTTGTG CAAGTCTAGACAGCCAGAACCGGAACAAGAGCCTGGCATGTCCGACCCGCTGCCTAAGCCTCTCCGGGACCC TCTCCCGGACCCGCTGTTGGACAAATTGGTATTGCCAGTGTTGCCAGGAGCACTCCAAGCGCGACCAGGACCG CACACGCAGGATCTTTCCGAACAACAGTTTCCGATACCACTTCCATACTGCTGGTTGTGCCGAGCTTTGATAAA ACGGATTCAGGCGATGATTCCAAAGGGAGCGCTGGCTGTCGCGGTGGCGCAAGTCTGCCGCGTTGTACCATT GGTCGCTGGAGGTATCTGCCAATGCTTGGCGGAGCGCTATTCAGTGATTCTGCTCGACACCTTGCTTGGTCGC ATGCTCCCGCAACTCGTGTGCCGCCTTGTTCTCCGATGCAGTATGGACGACAGTGCGGGGCCACGGAGCCCCA CCGGTGAATGGTTGCCACGAGATTCAGAATGCCATCTGTGTATGTCTGTTACTACGCAGGCCGGAAACAGCAG CGAACAAGCGATCCCCCAAGCCATGTTGCAAGCATGCGTAGGCTCCTGGCTTGACAGGGAAAAATGCAAGCA GTTCGTGGAGCAACATACGCCGCAGTTGCTGACTCTTGTGCCACGGGGTTGGGACGCTCATACAACCTGTCAG GCATTGGGAGTGTGCGGAACAATGTCTTCACCGCTGCAATGCATTCATAGTCCCGATTTGTGAGGGCCC    SEQ ID NO: 56 Homo sapiens surfactant protein B (SFTPB), RefSeqGene on chromosome 2. (NCBI  Reference Sequence: NG_016967.1) (18425bp)  ttggctgtca ctctctacca agctccccaa cactctgaac ctctatttgt ggtgcagagt 60 ggtgaactct ccctgtacca gccccatgct ctctcttttc ctaggccttt gcatgaacta 120 tttatctgcc tggaaatcct cctcaatcaa tcccgcccca ccatttgacc aggccaattt 180 ctatgcctcc ccaagtctcc ccgtagctgc cccatctccc ccagcgtctc ctcttcctct 240 ccaggcctgg gtagtattcc actctgatta ctggtctgtg ccctcatcca tagatattaa 300 gtccaccagg gtggtgacca ttgtcatccc aactcctagc tcattaacca ttttttggtc 360 aaacaaatga attgtgaaag gtctttataa actgtaagtt tcattgccaa gggccctacc 420 tccctacccc ttcccacgga gaaattcctt cttttgaaag aaagacataa gatagagcta 480 tgctaggaaa ggctgctacc atgaggcctc ctgaggtgca ctggccttgt atttctcatg 540 ggttggttct accccagaaa cagaaactac actcaaataa attcaactac gaagagtggg 600 gataggagag gggaggaaag tcatgggtcc ttaggctcag gccccaagca gagtcagaag 660 aatgcagccc aaaatacaac aaaaggaact aaaatgaatt ctgctcttcc gtgcatctct 720 gcaaggacat aaaagagatt ttaatatcaa ctgtcccatt tggctgatct cttttttccc 780 agccccaggt gtctcaagtc ccacacttgg gagtctcaca ctttccacag tcaagaagaa 840 aaacaagact cttctccacc tgcagcctgg tttgtttctc tctcttcccc tttctttttc 900 aattgtgatg aaatatacaa atatagataa cataatactg ccttaaccat ttttaagtgt 960 acagctcggg gacgtcaagc ccacccacat tatgcaacca aaaccgtcgt tcatctccag 1020 aacttttcat cacaggatga aactttgttt ccattaacca gtacctccgc gttccctctt 1080 ccccagcccc tggtaacctc tgacctactt tctgtcccta cgagtttgac tactctgcat 1140 acttcctata aatggaatca cacaatatct gtccttttgt gactggatta tttcactgag 1200 cgtcatgttt tcaagattca tccatgttgt agcatgtgtc caaatttctt tccttttttt 1260 tcctaccccc acccacccca gacggagtat tgctctgttg cccaggctag agtgcagtgg 1320 cgcaatctca gttctctgta acctccgcct cccgggttca agcgattctc cagcctcagc 1380 ctcccgagta gctgggatta caggcatgca ccaccaggcc cggctaattt ttgtattttt 1440 agtagagacg gggtttcacc gtgttggtaa ggctggtctc gaactcctga cctcgtgatc 1500 tgcccacctt ggcctcccaa agtcctggga ttacaggcat gagccactgc gcctggctga 1560 atttccttcc taaaaaggct caataatatt cctttgagtg tgtacaacac attttgttta 1620 cccattcatc cgtggatgaa cacttgtttc taccttttgg ctattatgaa taatgctgca 1680 atgaaaactg acatacaaat atctgttccg gtacctgttt tccattctct tgaatacgta 1740 cgtaggagtg gaattgctgg gtctgaaagt aattcaatat tcaaactttt cagaaactgt 1800 tttgtaaaaa ctattttcca cacaggtttc cacagtgact gcaccatttc acattctcgg 1860 cagcaaagca cctagggttc caatttctcc acatcctctc aacatgtatt attttctggt 1920 ttttcataat agccgtccta gagggtgtga ggtatggttt tgatttgctc cattttgcat 1980 taagtggcca aggaccccct tggcaatgca ccattaactt gactttgatc cttcacccca 2040 caatcagctg acttggctct catctcccaa tttaaaaatc caattggttc agccaagcat 2100 gcagatggac tcttttgggg tctggtgtgt ctacctctgg tccagtcggc tctggccagg 2160 gggatgaagg gaggtggtcc atgaggcttc cctttccagg gatgtggtgc gtggctttct 2220 agacagtatt actggggcag gcaggggcag acctgtgtgg tgtagaccca tggacacaat 2280 gctaggcggg acatgtgtct tctatatttt gatgacaagg aactggctct cagaaacatg 2340 agtttgtcca cacaccaaga gaagaggcat ctcctgtgaa taaacttgag ttaatcaagt 2400 tctccatccc cagagtagct atcagaatgt tctttttttt tttttttttt tcccagaatg 2460 tttctttatt catttccagt gagaagcagg taggccagga gaagaaccgt cagtagaggc 2520 agatccatgc aggttccccc tgcagccacg cttcacagga agcccgattt tcatcagaga 2580 ctgtggtgag gatggcatca aagcaagggg tgccaggaga ctgcccatcc cagggaccag 2640 gagccaggac accctcagct cagtcaacag ctggcttttt gtgtctcctt tattggtgct 2700 cttggatggc agctctcgac agtcagatac caatccctct ttgaggttct gataaaagcc 2760 acggacccac tacctagaaa accccacata agcacctggt tttgcattca gcattatggc 2820 ttttagacaa tgctatgtgt ccaccaactc catctccctt ccctcctggg catgcaggcg 2880 cattgcactt cccagctttc cctggggtta gtgagggtgt gagactgaac tgagagggaa 2940 gaaatgacgg acaccactcc caggcctcgt ccctaaaacc tcctccacca tcttccactc 3000 tttctctctt ctccctgttg gccagatgca ggggacccag gggaggactc tgagccctcc 3060 aagcaagcag aagggggact ccgagcacct gaatgacggt acggacctga gtctcaccac 3120 ccggagccaa cctgtgctgc actgtaaccg agtaagaagg aaacatgtgg gccaggcgcg 3180 gtggctcaca cctataatcc cagcactttg ggaggccaag gcaggcggat cacctgagat 3240 caggagttcg aaaccaacct ggtcaacatg gtgaaacccc atctctacta aaaaaataca 3300 aaaattagct gggcgtgatc acgggtgcct gtactcccca gctattcagg aggctgaggc 3360 aggagaatcg cttgaacccg ggaggcagag gttgcagtga gccaagatcg taccactgca 3420 ctccagcctg ggcaacagag tgagactcca tctcaaaaaa aaaaagaagg aaacatgtgg 3480 ctaggcacag tggcttatgc ctgtaatccc agcactttgg gaggctgagg tgggtggatc 3540 acctgagctc aggagttcga aacgagcctg ggcaatatgg caaaaccctg tctctaccaa 3600 aaatatcata gaaaaaaatt agccagacgt ggtggcgtat gcctgtggtc cgagctactc 3660 tggagactta ggtgggagga tcacttgagc ctgggagaca gaggttgcag tgagccaaga 3720 tcacgccact gcactccagc ctgagcgcca gagtgagacc ccatgtcaaa aaacaaaaaa 3780 aaaaagaaaa agaaacttgc atgaagccca tagactggag ggctctctgt tatggtagtg 3840 ggaccaccct gaccaatacg cccctggagg cagccagcca caaatcccca ctctcttgcg 3900 cagcctgtcc tgattgtatt caaacgagat accactcgct gatttgttta atacatactg 3960 aataacaagc acgtttcagt tccaccttcc tgaggttcta ccaaggtcac tcagttccca 4020 gtctttggtt actcattttt acacagtctt ggtggactcc catgatggca gctgtgtttg 4080 ttgtactgac aaggacttga ggccaacaga aactacataa atccttgggt aaccttgagt 4140 gtggcattag ccaccttgta tctaattctt tttttttttt tttttttgag acagagtgtt 4200 acctggccaa catggtgtca ggagactgcc catcccaggg accaggagcc aggacaccct 4260 cagctcagtc acatgtttaa tggccacatg tttaatggcc tcacactccc attccagagt 4320 gcagcggtgc aatcatggct cactgcagcc gcgacctggt gggctcaagc gatcctcctg 4380 ccccagcttt ctgagtagct gggaccagag acatgcacca ccacacaagg ttaatttttt 4440 aatttttttg tagttatgga gtcttactat attggccagg ctggtctcaa actccagggc 4500 tcaaaggatc ctccctcctc ggcctcccaa agtgccagga ttacaggagt gagccaccac 4560 acccagcccc atctcttttc atcatggtac taattcctgc ccgtccaccc acaaaagcac 4620 tgtagtgctt cccgagtata gaggcctgtg agcctccact agggagaggg ctcctgcaga 4680 gatcagataa attgatcaca atggctgggg tggtggcaat gtgctaatgc tctctttctt 4740 ccactcaaga tatcctctgt ctccctcagc ctgtgagctt tttctccagt gtgctctgcc 4800 agtgggggcc ctgcctgaga gcccctgcag ctgcagagga cagtttcttt ctgctgaacc 4860 atcgcagcta tgccccagcc cctaccctgg aggggtcccc aggggccatg ggcagcacct 4920 cctgtatagg gctgtctggg agccactcca gggccacaga aatcttgtct ctgactcagg 4980 gtattttgtt ttctgttttg tgtaaatgct cttctgacta atgcaaacca tgtgtccata 5040 gaaccagaag atttttccag gggaaaaggt aaggaggtgg tgagagtgtc ctgggtctgc 5100 ccttccaggg cttgccctgg gttaagagcc aggcaggaag ctctcaagag cattgctcaa 5160 gagtagaggg ggcctgggag gcccagggag gggatgggag gggaacaccc aggctgcccc 5220 caaccagatg ccctccaccc tcctcaacct ccctcccacg gcctggagag gtgggaccag 5280 gtatggaggc ttgagagccc ctggttggag gaagccacaa gtccaggaac atgggagtct 5340 gggcaggggg caaaggaggc aggaacaggc catcagccag gacaggtggt aaggcaggca 5400 ggagtgttcc tgctgggaaa aggtgggatc aagcacctgg agggctcttc agagcaaaga 5460 caaacactga ggtcgctgcc actcctacag agcccccacg ccccgcccag ctataagggg 5520 ccatgcacca agcagggtac ccaggctgca gaggtgccat ggctgagtca cacctgctgc 5580 agtggctgct gctgctgctg cccacgctct gtggcccagg cactggtgag tctcccccag 5640 cctcccctct cctaggcagc tccaccactc actgagcact gctttgtgct aggcattaac 5700 ccaagtctgt cctcatttta aagacaaggc agctggggtt cagagagggt tcagagctta 5760 tccaaggtca cacagctggc gggtccagga gcaggtggaa cccagagctg tctgacgtcc 5820 acatgtttaa tggcctcaca ctcccagcaa aactgggtct agagggtggg tgaaatcatg 5880 atgccaggtg tgtagcctgg atcctgatta aggttgctct ggccccaaac cacagctgcc 5940 tggaccacct catccttggc ctgtgcccag ggccctgagt tctggtgcca aagcctggag 6000 caagcattgc agtgcagagc cctagggcat tgcctacagg aagtctgggg acatgtggga 6060 gccgtgagta ccaccaagga tgcatggcaa ctgggggtct gaaatgaagg gtgctgggtg 6120 ggctctggat gggcaggagg agagtggagc ccccataggg gatggatgag atgaaatggg 6180 atgagatgaa atgagatagg ataaaatgga atgggatgga tgcgatggga tacgatgaca 6240 tagaatagat ggagtcggat gaatgggatg ggatgggatg gatgggaggg gaagggatag 6300 gataggatga catagaataa agatggatgg gatgggatgg gatgggatgg gatgacacag 6360 aataaagatg gatggattgg gatggatgaa tagaagagat ggatgggata aattgatatg 6420 gatgagatgg gacaagttgg gctggtgggc agctgcatgt gccttggagt gctctgttgg 6480 cctcttccta agagaacctc cccattggag ctgggagcct cccccactca tgtgtcctcc 6540 accttggggc ccctccctcc ccaggatgac ctatgccaag agtgtgagga catcgtccac 6600 atccttaaca agatggccaa ggaggccatt ttccaggtaa tgatgcccag atcctggatg 6660 aaggttgggg cccaagagat gagggacaga gcagggaaga gctgagcccc ctaaaggggc 6720 catttccagg ctgaggagga ggcctgggtg cctgggaagt cccagctcct cctggctggg 6780 agcaggtcat ggccctgagc tcaatagcac agccagagat ggtcttccct gaggggaagg 6840 gcccctacat gtgcccaact acttaactcc ttggcactcg tgaactccag caccctgggg 6900 gattaggggt cagtctgccc tggtggggcc ttgtgtccag ggacttgggc ggggtagacc 6960 tcagagaggc ccagctgacg gccccctctg gcctcccagg acacgatgag gaagttcctg 7020 gagcaggagt gcaacgtcct ccccttgaag ctgctcatgc cccagtgcaa ccaagtgctt 7080 gacgactact tccccctggt catcgactac ttccagaacc agactgtgag ggctgcaagc 7140 tcacctcctg cctgcctccc cacgcaggcc cctgtgccca cccatgggga gccacacaca 7200 cagcacccca gccagccaga cacacacaca cacacacaca cacacacagc acccaagccg 7260 gccagacaca aacacacagc accccagcca gccggacaca cacacacaca cacacacaac 7320 accccagctg gccggacaca cacacacaca gtaccccagc tggccggaca cacacacaca 7380 cagcacccta tccagacaca tacacacaca cagtacccca gccagctgga aacacacaca 7440 cacacagcac tccatccaga cacataccca cacagtaccc cagccagcca gacacacaca 7500 cacacacaca cacacacaca cacacagcac acacacagca ccccagctgg ccacacacac 7560 acacacacac accctgtcca caaagggcct aggaaactac gtgcccttca gccatgcacc 7620 cgaccatggg cccccaggtt caggtgcaca cggtgggcct gtacgctcac acacccttac 7680 accctcactc tcacacacat gcttacacac ttattcattc tcacatatat gctcatgctc 7740 attcacacac aatcccggcc acctgcccta aagtccccac acagccctat ctttgccttt 7800 tgtcccccca catagagttc taaaccacag cacccccact aggcctgctt cctcccattc 7860 cagtggtccc tgagcccttg ggccggcctg aataggggtg ggcttccctc ccagacccta 7920 acactcccac cctgtgctgt gccccaggac tcaaacggca tctgtatgca cctgggcctg 7980 tgcaaatccc ggcagccaga gccagagcag gagccaggga tgtcagaccc cctgcccaaa 8040 cctctgcggg accctctgcc agaccctctg ctggacaagc tcgtcctccc tgtgctgccc 8100 ggggccctcc aggcgaggcc tgggcctcac acacaggtga gggaggcccc cacagccagt 8160 aaagtggaga tccagagggc tagagccacc tccgaagccc atgggcactg ggccctggga 8220 gaggcagagc cgggaaggtg ataggaagct ccaggcaggg cctaagggag gagggagaga 8280 aagggaggaa gagagagggg aggagagcct ggaggactct tctcccagca cccagcctgg 8340 cctccacctg attctttccc caggatctct ccgagcagca attccccatt cctctcccct 8400 attgctggct ctgcagggct ctgatcaagc ggatccaagc catgattccc aaggtgaggc 8460 atccagggcc tcacagagcc caggagcaca cgcatacctg tagctccctg cagctcccac 8520 ctctctccca actcacaccc ccgtcaggac ccagctggct gccagaagtt aggaggggag 8580 agagccgctt gtgcattgcc cccacccagg ggaccctggg gctcaggctc aggcctggta 8640 ggtgccaggc ctacagttca tgcaacaaac attaagcccc cactgtatgg aggtgccacg 8700 ccaggagcca aagtacaaaa acggacaaga cgcagctttg tcctccagca gctcaccatc 8760 tgatggagaa agatccccag aggtctctgt agaaaggttg ctttgatctt tcaagagggg 8820 aatttccaca gatagattcc ccatccttgc ctgagtccaa cttggagtct tccagacctg 8880 cagtggctat tgtccaatgg ccccgccagc ccagggctac cttgcccaaa ttggggccca 8940 aatgaggaaa ggccctgccc cctcagcctt tcccagatag ggttgcgtgg gccaccaggg 9000 gcacaaggca gcaggtgagg ttcctgctga ggcaggtggt tcacttgagc ccaggagttc 9060 aagaccagct tgggcaacat ggcgaaaccc cgtctctact aagaatacaa aaattagcca 9120 gatgtgacag gtgcctgtag tcccagctac tcgggaggct gaggcaggag aatcacttga 9180 acccaggagg cggaggttgc agtgagccga catcacgcca ctgtactcta gcctgggtga 9240 cagagcaaga ctctgtctca aaaaaaaaga aagaaggaaa gatcactgca gagattgcag 9300 tgagaggtga tgggacaggg acggagctga gggctggcct ggggatgcat ttgggaggtg 9360 ggcccactgc tattgggcat ggatgggcct ggagcgtgag gaccagggag gactccaaag 9420 tgacttttac acactggcca gagcaaccag ccctctgtaa tgccagcagc tgagatgggg 9480 agactaaaga agaaaacagg tttgagcaaa aaaacagaga gctccctcct ggccatgttg 9540 agttcaagat gcctgtgtga agtgcaggag aggagagtca ggcaagcagc tgaatcccaa 9600 gcattggggg aaggtcaggt ccaccatgtc agtctgagag tcactagctg tgggccagag 9660 cctttggggc cagacgtagg tctgaagctg gctcctacac tcagtgaccc tgtgtgagtc 9720 ccctgcatcc cctggactct ctgatcccca gtgtccttat ttgtgaatag ccttgccctc 9780 ccttctagaa gagaatgagg gaatgcgtag gaagtgccca gctgggtgct gggcagagag 9840 tggaggcttg ccaagtgaag gtcccatgct ggcctctctc cgcccccgcc ccagggtgcg 9900 ctagctgtgg cagtggccca ggtgtgccgc gtggtacctc tggtggcggg cggcatctgc 9960 cagtgcctgg ctgagcgcta ctccgtcatc ctgctcgaca cgctgctggg ccgcatgctg 10020 ccccagctgg tctgccgcct cgtcctccgg tgctccatgg atgacagcgc tggcccaagt 10080 gagcccactg cccactcctt agcccaatgc ctgctctcct cctcccccta ccctgccact 10140 gcatgaccct ctccctctgt ggtcccactg caatgcacca aggaggacag aaaccaaaca 10200 cctctgtagg gtggccttgc ctgctttccc cctaatgctc acatctccag ggtcgccgac 10260 aggagaatgg ctgccgcgag actctgagtg ccacctctgc atgtccgtga ccacccaggc 10320 cgggaacagc agcgagcagg ccataccaca ggcaatgctc caggcctgtg ttggctcctg 10380 gctggacagg gaaaaggtat gggctgggca catggggact catggtcagg gcccgttcaa 10440 ggcagaaggc tgagcccagg aaaggctttg cagccagaga cacctaggat gggccagaat 10500 ggagcacaga caggcagaca ggatgtgggg cagacaatgg tgggactgta agttagggca 10560 gagcctgcta aaggttagga gtcgcctctg gacaaagggc tgtgggctcc agaggaccag 10620 caggccctct tcacgggctg agtgagcacc aggcaagcct tcagaggcct ggttatctac 10680 caggagatga gtaatgctag ggccagttca agccaggaaa gggactagcc ttctctccag 10740 ggtcctgatc cctttactgc ccccacactc ctcaaggtgt gactcactca ggacaaaccc 10800 attggcaaaa ggagagggct ggacttgaag gtcctagggc ccttgccaat actcagtcaa 10860 tgacaggaaa ttcccttttt tttttttttt tttttttttt tgagatggag ttttgctctt 10920 gttgcccagg ctggagtgca atggcacaat cttggctcac tgcaacctct gcctccgggt 10980 tcaggcgatt ctcctgcctc agcctcttga gtagctggga ttacaggcat gtgctaccag 11040 gcccggctaa tttttgtatt tttagtagag acaaggtttc accatattgg tcaggctggt 11100 ctcgaacccc tgacctgaag tgatctgccc gccttggcct cccaaagtgc tgggattaca 11160 ggcataagcc actgcacccg gacaggaaat tcccttctta aagcgagatc ctgtcctgag 11220 gaaagccagc tgatgctctt cccaggaggc agctgtccac actgtgctcc ctgctcagca 11280 actcccaagc ctcccgactg cccatcacat ctggtctcaa ggaccagatg aacgttaagt 11340 ttccttctag aactgaaatg gaggtggagg gaggggaggg tggtggctga gattccaccc 11400 ctctgcctga gtcctccgtc tccagtgtcg cctgcttttc tgatggaagt cctccatttc 11460 agctggctcc agtttgttaa gggtttcaac tgcagccaga ggtgttccgt gagggctgat 11520 ggaggagtcg ggagggagcc ctagagtgat ccagagatgt ggagaggcca ggaccacacg 11580 acaggagagt cctgcaaagg gaccctccac agctgtgtgt ctccttcctc agtgcaagca 11640 atttgtggag cagcacacgc cccagctgct gaccctggtg cccaggggct gggatgccca 11700 caccacctgc caggtacacc cacccctccc agttggtcct aggacttccc ttggctccca 11760 gagcccccac cctttgggcc cgtgatcctc agaggcctca ctcccctggg tccaaggtgg 11820 tcccaggtgc acgggccagg gactgggagg cacccctctc tgtttcagtg taaaaaatca 11880 tgagagcatg gaaaaggggg atgggaaggg agggatggcc tgaggagtgc ggctggatgt 11940 ccattatagg atggggctgt gttccctggc cagtgtgtgc tggtggggtg ggggtacaaa 12000 gtgggtgttc tggagtgaac atctcacctc ctcaggctct aaaccctaag gcctgtggct 12060 cagggagtgg cccgaggggt ctacagagtc acactggtag cacccactag gcgggaggtg 12120 gagtgagtgc tgttctttcc cggaagagct gggtgtgggg agctgagggg gcccaggcct 12180 cagccctggt gctgtccctg tgacaggccc tcggggtgtg tgggaccatg tccagccctc 12240 tccagtgtat ccacagcccc gacctttgat gagaactcag ctgtccaggt gagtccaggc 12300 ccccagttgc ggggaggtaa gggggcaggt cctgaccatc agggcatggg aggcccttct 12360 gctccccaag caggaagagg cggccactcc tgccggctgc tccatcctcc ctctcaccgc 12420 acagctggag gctcctgagg gcttctggct ggccatcagg aaaacaccct ttccggaccc 12480 cgagcactgc cccgcccaga accccagtca ctgagtgccc aacccccagc ttccccccca 12540 accccccgcc ctgccctgtc ccaggcctcc ctctcagagc ttgccccagg gactctctgg 12600 ccctcagggt tcaatgtatt ctgaccaagg ccaagctttc ctggggctca gggaaaatca 12660 cactttgcta cccgaagctg tatcccctca gatgccagga aggccgtgat catctgactc 12720 caccctcctg agacacattc tctccctgac tgtcctgttc taagtcagcg gagcacctta 12780 ggatggaggg gtggaggcga ggccagatgc agcctctgtg aacaggtgcc tggaggctgg 12840 gaaatgaccc tgagagggca ggacacagca accgtgggct taaggtgacc ttgagagcaa 12900 gcttggccca ctttacaatt ctgttcagag ccagccccta acatggtggt catttattca 12960 tttgttccct cattttaaaa aatgtaaggc caggcatggt ggctcacgcc ggtaatccca 13020 gcactttggg aggccgaggc aggcagatca cctgaggtca ggagttcgag actagcctgg 13080 ccaacatggc gaaaccctgt ctctactaaa aatatttttt aaaaattagc tgagcatggt 13140 ggcaggtgcc tgtaatccca gctactcagg acgcttaggc aggagaatca cttgaacctg 13200 ggaggcgaag gttgcggtgt gccgagatcg tgccactgca ctctagccta ggcaacagag 13260 cacaactctg tctcaggaaa aaaaaaaaaa aaaaaaaggt atttctttgc tgggcgcagt 13320 ggctcacacc tgtaatccca gcactttggg agaccgaggc gagtggatca cttgaggtca 13380 ggagttcaag accagcctta ccaacatgat gaaaccccgt atctactaaa aaaaaaaaaa 13440 aaaaaaaaaa aaattagcca gatgtggtgg cacacacctg taatcccagc tacttgggag 13500 gctgaggagg agaattgctt gaacctggga ggcggagatt gcagcgagcc aagattgcgc 13560 ctctgcactc cagcctgggt gacagagtga gactccgtct caaaaaaaaa aaaaaaaaag 13620 tagtgggtgc ctgtggccag gccacatcct agggtagggg ctatggctga gccctgccct 13680 cctggagctc acagccaagt ccacttcttc catctgaggc ggggaagcca gccctgttcc 13740 tgaaaccctg catcacaagc ccctgtggga ggcagtgggg aggggaggtc ctcccccact 13800 cagacctgac ccacagggac cagtttaatg tgtccttgcc ccagtgatga cagctgggga 13860 tctgggggtg gggagtcacc caggacccgg gcagtcgcct ttccccagct cctagggctc 13920 ccggccttcc ctgctgaaac agcaagacca gtgggttggc gtgggaggcc tgggcttcaa 13980 accacctctg ctatcacctg gctgtgggtc cccaggcagg acatacacac agtccctctc 14040 tggccctcat cctcctcagc tgcaaaggaa aagccaagtg agacgggctc tgggaccatg 14100 gtgaccaggc tcttcccctg ctccctggcc ctcgccagct gccaggctga aaagaagcct 14160 cagctcccac accgccctcc tcaccgccct tcctcggcag tcacttccac tggtggacca 14220 cgggccccca gccctgtgtc ggccttgtct gtctcagctc aaccacagtc tgacaccaga 14280 gcccacttcc atcctctctg gtgtgaggca cagcgagggc agcatctgga ggagctctgc 14340 agcctccaca cctaccacga cctcccaggg ctgggctcag gaaaaaccag ccactgcttt 14400 acaggacagg gggttgaagc tgagccccgc ctcacaccca cccccatgca ctcaaagatt 14460 ggattttaca gctacttgca attcaaaatt cagaagaata aaaaatggga acatacagaa 14520 ctctaaaaga tagacatcag aaattgttaa gttaagcttt ttcaaaaaat cagcaattcc 14580 ccagcgtagt caagggtgga cactgcacgc tctggcatga tgggatggcg accgggcaag 14640 ctttcttcct cgagatgctc tgctgcttga gagctattgc tttgttaaga tataaaaagg 14700 ggtttctttt tgtctttctg taaggtggac ttccagcttt tgattgaaag tcctagggtg 14760 attctatttc tgctgtgatt tatctgctga aagctcagct ggggttgtgc aagctaggga 14820 cccattcctg tgtaatacaa tgtctgcacc aatgctaata aagtcctatt ctcttttatg 14880 agaaagaaaa agacaccgtc ctttaaagtg ctgcagtatg gccagacgtg gtggctcaca 14940 cctgcaatcc cagcacctta ggaggccgag gcaggaggat ccttgaggtc aggagttcga 15000 gaccagcctc gccaacatgg tgaaacccca tttctactaa aaatacaaaa aattagccaa 15060 gtgtggtggc atatgcctgt aatcccaact actcagaagg ccgaggcagg agaattactt 15120 gaacgcagga gaatcactgc agcccaggag gcagaggttg cagtgagccg agattgcacc 15180 actgcactcc agcctgggtg acagagcaag actccatctc agtaaataaa taaataaata 15240 aaaagcgctg cagtagctgt ggcctcaccc tgaagtcagc gggcccaggc ctacctcact 15300 ctctcccttg gcagagaagc agacgtccat agctcctctc cctcacaagc gctcccagcc 15360 tgccctccag ctgctgctct cccctcccag tctctactca ctgggatgag gttaggtcat 15420 gaggacacca aaaacctaaa aataaacaaa aagccaaaca agccttagct tttcttaaag 15480 actgaaatgc ctggaagtgt ccctttattt ataaaataac ttttgtcata tttcttatac 15540 atgtttcttg taagaaattc agaaactaca gacaaagaga gtggaaatta cccactgtca 15600 ggcctctgag cccaagctaa gccatcatat cccctgtgcc ctgcacgtat acacccagat 15660 ggcctgaagc aactgaagat ccacaaaaga agtgaaaata gccagttcct gccttaactg 15720 atgacattcc accattgtga tttgttcctg ccccacccta actgatcaat tgaccttgtg 15780 acaatacacc ttccccaccc ttgagaaggt gctttgtaat attctcccca cccaccccac 15840 gcccgcaccc ccgcaccctt aagaaggtat tttgtaatat tctctccgcc attgagaatg 15900 tgctttgtaa gatccacccc ctgcccacaa aaaattgctc ctaactccac cgcctatccc 15960 aaacctacaa gaactaatga taatcccacc accctttgct gactcttttt ggactcagcc 16020 cacctgcacc caggtgatta aaaagcttta ttgttcacac aaagcctgtt tggtagtctc 16080 ttcacaggga agcatgtgac acccacaatc ccacctagcc caggagagag ctacggcagg 16140 gtgtgtgttt tgacactgag cttggggctt tttccatctt ctccccacag cctctggctc 16200 cacacctcca ccgttcaagc gccagaaaga gctgtctatg cagcctgctc ttgggcctgg 16260 ggatgagaca cacaattcat tggctcctgg attttaagta gacatttgta aatctatagc 16320 taactactgt ccttaaagcc attgtttcca ttacaaaatc caactctctg agagaaaagg 16380 gtgttttaaa tttaaaaaaa taaaaacaaa aaagtttgat tgagacaatg aagagccaat 16440 gtttgaagaa ctggaggcat taatagttta tcccaaaaat atgagtggga atgagtgcct 16500 gataatgaaa aaggcggctt cattcaaaag gaggtcaaga attggcagac tttgggaggc 16560 tgaggtgagt gaatcacttg agcccaggag ttcaagacca gactgggcaa cgtggcaaaa 16620 caccatctct ataaaaaata taaaaatttg gctgggcatg gtggctcaca cctgtaatcc 16680 tagcactttt ggaggctgag gtgggtggat cacttgaggt taggagttcc agaccagcct 16740 aacatggtga aaccacatat ctactaaaaa aaatacaaaa ttagctggat gtggtagtgc 16800 atgcctgtta tcccagctac ttggaaggct gagacaggag aattgcttga acccaggagg 16860 cagaggctgc agtgagccag gattgcgcca ctgcactcca gcctgggcaa caagagcaaa 16920 actctgtcta aaagaaaaaa aaaaaagata caaacattag ctggccatgg tggtgtgtgt 16980 ctgtggtccc agctattcag gaggctgagg tgggaggatc gcctgaatcc aggtcaaggc 17040 tgcagtgagc cgtgatgatg ccactgcact ccaccctggg tgacacagag agatcctatt 17100 tcaaaaacaa ataaataaat aattttaaaa aagagttgac tagctggtag gaggacagaa 17160 ttggcaggaa gtggctacac ttgagactcc tagggtccct gagaagtgtc caacatcctg 17220 cacctctgcc cttaggaagt tgggggaaga gaagtcccag gagggattat agcccaagag 17280 gaggggactt attcttcact ttccagggtg tgggcttcta agagaagcct gcctctcaca 17340 ggacatcaga aagggtgagg acgctagcct gagggatgga gctaggcagc tggcctgagg 17400 tctgcctcac agagtgttcc cagtgtgatg tggttgcaga agagcaacct caacccacag 17460 accgagaggc agtcagaatt ggctctggac aagatggtgc acccgtaaga gcccacggtt 17520 ggggatgaat ctgtgacaga tttagtggat cttgagatga gcagaaatca gatgtgtagc 17580 agtgatggca caagaagatg atgatcaggc aagacgagag acactccaga gagacctgga 17640 tggccaggga atgcttagtc cccagtgcct ctcctcaccg cagggcaatc tgtaaatcat 17700 tcccagggaa aggggaaaga ggattcagtg gttagtcagc caagagaggc tgagtattaa 17760 gaagggaatg agttactccg aaatggaaat aatatgtttt gagggctcat gagccaaaat 17820 gagttcagtt atagagaaat aaagataatt acattatgtc cccctaatga ataactgaaa 17880 ttttaaatcc acgacataat cattattaac atgttattat ctataattcc aaagctttct 17940 gtgcacatat acaaaaacag atacagcaga tctacaaaat ggaaacatcc tatacatgct 18000 gctttgtaac ctgcactgtt catttaacaa taaccttttc aggtcaataa atgtaagtat 18060 ctactattat tgctaattgc cacatgacag ctctgtctcc caggctagag tgcagtggtg 18120 cgatcactct ggggtttaag caatcctcct gcctctgcct ccctaagtgc tgggactaca 18180 ggcgtcagat accatggctg gcccttactt tttttttttt ttttttttga gacaagagtc 18240 tcactctgtc atccaggctg gattgcagtg gcgtgatctc agctcactgc aagctccgca 18300 tcccaggttc atgccattct cctgcctcag cctcctgagt agctgggact acaggagccc 18360 gccaccacgc ccagctaatt tttttttttt tttttttttg tatttttact agacacgggg 18420 tttca 18425    SEQ ID NO: 57 Exemplified Human SFTPB polypeptide  MAESHLLQWLLLLLPTLCGPGTAAWTTSSLACAQGPEFWCQSLEQALQCRALGHCLQEVWGHVGADDLCQECE DIVHILNKMAKEAIFQDTMRKFLEQECNVLPLKLLMPQCNQVLDDYFPLVIDYFQNQTDSNGICMHLGLCKSRQPE PEQEPGMSDPLPKPLRDPLPDPLLDKLVLPVLPGALQARPGPHTQDLSEQQFPIPLPYCWLCRALIKRIQAMIPKGA LAVAVAQVCRVVPLVAGGICQCLAERYSVILLDTLLGRMLPQLVCRLVLRCSMDDSAGPRSPTGEWLPRDSECHLC MSVTTQAGNSSEQAIPQAMLQACVGSWLDREKCKQFVEQHTPQLLTLVPRGWDAHTTCQALGVCGTMSSPLQ CIHSPDL*    SEQ ID NO: 58 Exemplified ABACA3 (ABCA3) transgene (NM_000542.5) (5115bp)  ATGGCTGTGCTCAGGCAGCTGGCGCTCCTCCTCTGGAAGAACTACACCCTGCAGAAGCGGAAGGTCCT GGTGACGGTCCTGGAACTCTTCCTGCCATTGCTGTTTTCTGGGATCCTCATCTGGCTCCGCTTGAAGA TTCAGTCGGAAAATGTGCCCAACGCCACCATCTACCCGGGCCAGTCCATCCAGGAGCTGCCTCTGTTC TTCACCTTCCCTCCGCCAGGAGACACCTGGGAGCTTGCCTACATCCCTTCTCACAGTGACGCTGCCAA GACCGTCACTGAGACAGTGCGCAGGGCACTTGTGATCAACATGCGAGTGCGCGGCTTTCCCTCCGAGA AGGACTTTGAGGACTACATTAGGTACGACAACTGCTCGTCCAGCGTGCTGGCCGCCGTGGTCTTCGAG CACCCCTTCAACCACAGCAAGGAGCCCCTGCCGCTGGCGGTGAAATATCACCTACGGTTCAGTTACAC ACGGAGAAATTACATGTGGACCCAAACAGGCTCCTTTTTCCTGAAAGAGACAGAAGGCTGGCACACTA CTTCCCTTTTCCCGCTTTTCCCAAACCCAGGACCAAGGGAACCTACATCCCCTGATGGCGGAGAACCT GGGTACATCCGGGAAGGCTTCCTGGCCGTGCAGCATGCTGTGGACCGGGCCATCATGGAGTACCATGC CGATGCCGCCACACGCCAGCTGTTCCAGAGACTGACGGTGACCATCAAGAGGTTCCCGTACCCGCCGT TCATCGCAGACCCCTTCCTCGTGGCCATCCAGTACCAGCTGCCCCTGCTGCTGCTGCTCAGCTTCACC TACACCGCGCTCACCATTGCCCGTGCTGTCGTGCAGGAGAAGGAAAGGAGGCTGAAGGAGTACATGCG CATGATGGGGCTCAGCAGCTGGCTGCACTGGAGTGCCTGGTTCCTCTTGTTCTTCCTCTTCCTCCTCA TCGCCGCCTCCTTCATGACCCTGCTCTTCTGTGTCAAGGTGAAGCCAAATGTAGCCGTGCTGTCCCGC AGCGACCCCTCCCTGGTGCTCGCCTTCCTGCTGTGCTTCGCCATCTCTACCATCTCCTTCAGCTTCAT GGTCAGCACCTTCTTCAGCAAAGCCAACATGGCAGCAGCCTTCGGAGGCTTCCTCTACTTCTTCACCT ACATCCCCTACTTCTTCGTGGCCCCTCGGTACAACTGGATGACTCTGAGCCAGAAGCTCTGCTCCTGC CTCCTGTCTAATGTCGCCATGGCAATGGGAGCCCAGCTCATTGGGAAATTTGAGGCGAAAGGCATGGG CATCCAGTGGCGAGACCTCCTGAGTCCCGTCAACGTGGACGACGACTTCTGCTTCGGGCAGGTGCTGG GGATGCTGCTGCTGGACTCTGTGCTCTATGGCCTGGTGACCTGGTACATGGAGGCCGTCTTCCCAGGG CAGTTCGGCGTGCCTCAGCCCTGGTACTTCTTCATCATGCCCTCCTATTGGTGTGGGAAGCCAAGGGC GGTTGCAGGGAAGGAGGAAGAAGACAGTGACCCCGAGAAAGCACTCAGAAACGAGTACTTTGAAGCCG AGCCAGAGGACCTGGTGGCGGGGATCAAGATCAAGCACCTGTCCAAGGTGTTCAGGGTGGGAAATAAG GACAGGGCGGCCGTCAGAGACCTGAACCTCAACCTGTACGAGGGACAGATCACCGTCCTGCTGGGCCA CAACGGTGCCGGGAAGACCACCACCCTCTCCATGCTCACAGGTCTCTTTCCCCCCACCAGTGGACGGG CATACATCAGCGGGTATGAAATTTCCCAGGACATGGTTCAGATCCGGAAGAGCCTGGGCCTGTGCCCG CAGCACGACATCCTGTTTGACAACTTGACAGTCGCAGAGCACCTTTATTTCTACGCCCAGCTGAAGGG CCTGTCACGTCAGAAGTGCCCTGAAGAAGTCAAGCAGATGCTGCACATCATCGGCCTGGAGGACAAGT GGAACTCACGGAGCCGCTTCCTGAGCGGGGGCATGAGGCGCAAGCTCTCCATCGGCATCGCCCTCATC GCAGGCTCCAAGGTGCTGATACTGGACGAGCCCACCTCGGGCATGGACGCCATCTCCAGGAGGGCCAT CTGGGATCTTCTTCAGCGGCAGAAAAGTGACCGCACCATCGTGCTGACCACCCACTTCATGGACGAGG CTGACCTGCTGGGAGACCGCATCGCCATCATGGCCAAGGGGGAGCTGCAGTGCTGCGGGTCCTCGCTG TTCCTCAAGCAGAAATACGGTGCCGGCTATCACATGACGCTGGTGAAGGAGCCGCACTGCAACCCGGA AGACATCTCCCAGCTGGTCCACCACCACGTGCCCAACGCCACGCTGGAGAGCAGCGCTGGGGCCGAGC TGTCTTTCATCCTTCCCAGAGAGAGCACGCACAGGTTTGAAGGTCTCTTTGCTAAACTGGAGAAGAAG CAGAAAGAGCTGGGCATTGCCAGCTTTGGGGCATCCATCACCACCATGGAGGAAGTCTTCCTTCGGGT CGGGAAGCTGGTGGACAGCAGTATGGACATCCAGGCCATCCAGCTCCCTGCCCTGCAGTACCAGCACG AGAGGCGCGCCAGCGACTGGGCTGTGGACAGCAACCTCTGTGGGGCCATGGACCCCTCCGACGGCATT GGAGCCCTCATCGAGGAGGAGCGCACCGCTGTCAAGCTCAACACTGGGCTCGCCCTGCACTGCCAGCA ATTCTGGGCCATGTTCCTGAAGAAGGCCGCATACAGCTGGCGCGAGTGGAAAATGGTGGCGGCACAGG TCCTGGTGCCTCTGACCTGCGTCACCCTGGCCCTCCTGGCCATCAACTACTCCTCGGAGCTCTTCGAC GACCCCATGCTGAGGCTGACCTTGGGCGAGTACGGCAGAACCGTCGTGCCCTTCTCAGTTCCCGGGAC CTCCCAGCTGGGTCAGCAGCTGTCAGAGCATCTGAAAGACGCACTGCAGGCTGAGGGACAGGAGCCCC GCGAGGTGCTCGGTGACCTGGAGGAGTTCTTGATCTTCAGGGCTTCTGTGGAGGGGGGCGGCTTTAAT GAGCGGTGCCTTGTGGCAGCGTCCTTCAGAGATGTGGGAGAGCGCACGGTCGTCAACGCCTTGTTCAA CAACCAGGCGTACCACTCTCCAGCCACTGCCCTGGCCGTCGTGGACAACCTTCTGTTCAAGCTGCTGT GCGGGCCTCACGCCTCCATTGTGGTCTCCAACTTCCCCCAGCCCCGGAGCGCCCTGCAGGCTGCCAAG GACCAGTTTAACGAGGGCCGGAAGGGATTCGACATTGCCCTCAACCTGCTCTTCGCCATGGCATTCTT GGCCAGCACGTTCTCCATCCTGGCGGTCAGCGAGAGGGCCGTGCAGGCCAAGCATGTGCAGTTTGTGA GTGGAGTCCACGTGGCCAGTTTCTGGCTCTCTGCTCTGCTGTGGGACCTCATCTCCTTCCTCATCCCC AGTCTGCTGCTGCTGGTGGTGTTTAAGGCCTTCGACGTGCGTGCCTTCACGCGGGACGGCCACATGGC TGACACCCTGCTGCTGCTCCTGCTCTACGGCTGGGCCATCATCCCCCTCATGTACCTGATGAACTTCT TCTTCTTGGGGGCGGCCACTGCCTACACGAGGCTGACCATCTTCAACATCCTGTCAGGCATCGCCACC TTCCTGATGGTCACCATCATGCGCATCCCAGCTGTAAAACTGGAAGAACTTTCCAAAACCCTGGATCA CGTGTTCCTGGTGCTGCCCAACCACTGTCTGGGGATGGCAGTCAGCAGTTTCTACGAGAACTACGAGA CGCGGAGGTACTGCACCTCCTCCGAGGTCGCCGCCCACTACTGCAAGAAATATAACATCCAGTACCAG GAGAACTTCTATGCCTGGAGCGCCCCGGGGGTCGGCCGGTTTGTGGCCTCCATGGCCGCCTCAGGGTG CGCCTACCTCATCCTGCTCTTCCTCATCGAGACCAACCTGCTTCAGAGACTCAGGGGCATCCTCTGCG CCCTCCGGAGGAGGCGGACACTGACAGAATTATACACCCGGATGCCTGTGCTTCCTGAGGACCAAGAT GTAGCGGACGAGAGGACCCGCATCCTGGCCCCCAGTCCGGACTCCCTGCTCCACACACCTCTGATTAT CAAGGAGCTCTCCAAGGTGTACGAGCAGCGGGTGCCCCTCCTGGCCGTGGACAGGCTCTCCCTCGCGG TGCAGAAAGGGGAGTGCTTCGGCCTGCTGGGCTTCAATGGAGCCGGGAAGACCACGACTTTCAAAATG CTGACCGGGGAGGAGAGCCTCACTTCTGGGGATGCCTTTGTCGGGGGTCACAGAATCAGCTCTGATGT CGGAAAGGTGCGGCAGCGGATCGGCTACTGCCCGCAGTTTGATGCCTTGCTGGACCACATGACAGGCC GGGAGATGCTGGTCATGTACGCTCGGCTCCGGGGCATCCCTGAGCGCCACATCGGGGCCTGCGTGGAG AACACTCTGCGGGGCCTGCTGCTGGAGCCACATGCCAACAAGCTGGTCAGGACGTACAGTGGTGGTAA CAAGCGGAAGCTGAGCACCGGCATCGCCCTGATCGGAGAGCCTGCTGTCATCTTCCTGGACGAGCCGT CCACTGGCATGGACCCCGTGGCCCGGCGCCTGCTTTGGGACACCGTGGCACGAGCCCGAGAGTCTGGC AAGGCCATCATCATCACCTCCCACAGCATGGAGGAGTGTGAGGCCCTGTGCACCCGGCTGGCCATCAT GGTGCAGGGGCAGTTCAAGTGCCTGGGCAGCCCCCAGCACCTCAAGAGCAAGTTCGGCAGCGGCTACT CCCTGCGGGCCAAGGTGCAGAGTGAAGGGCAACAGGAGGCGCTGGAGGAGTTCAAGGCCTTCGTGGAC CTGACCTTTCCAGGCAGCGTCCTGGAAGATGAGCACCAAGGCATGGTCCATTACCACCTGCCGGGCCG TGACCTCAGCTGGGCGAAGGTTTTCGGTATTCTGGAGAAAGCCAAGGAAAAGTACGGCGTGGACGACT ACTCCGTGAGCCAGATCTCGCTGGAACAGGTCTTCCTGAGCTTCGCCCACCTGCAGCCGCCCACCGCA GAGGAGGGGCGATGA   SEQ ID NO: 59 Exemplified Human ABACA3 (hABCA3) transgene (5139bp)  GCTAGCCACCATGGCTGTGCTCAGGCAGCTGGCGCTCCTCCTCTGGAAGAACTACACCCTGCAGAAGCGGAA GGTCCTGGTGACGGTCCTGGAACTCTTCCTGCCATTGCTGTTTTCTGGGATCCTCATCTGGCTCCGCTTGAAGA TTCAGTCGGAAAATGTGCCCAACGCCACCATCTACCCGGGCCAGTCCATCCAGGAGCTGCCTCTGTTCTTCACC TTCCCTCCGCCAGGAGACACCTGGGAGCTTGCCTACATCCCTTCTCACAGTGACGCTGCCAAGACCGTCACTGA GACAGTGCGCAGGGCACTTGTGATCAACATGCGAGTGCGCGGCTTTCCCTCCGAGAAGGACTTTGAGGACTA CATTAGGTACGACAACTGCTCGTCCAGCGTGCTGGCCGCCGTGGTCTTCGAGCACCCCTTCAACCACAGCAAG GAGCCCCTGCCGCTGGCGGTGAAATATCACCTACGGTTCAGTTACACACGGAGAAATTACATGTGGACCCAAA CAGGCTCCTTTTTCCTGAAAGAGACAGAAGGCTGGCACACTACTTCCCTTTTCCCGCTTTTCCCAAACCCAGGA CCAAGGGAACCTACATCCCCTGATGGCGGAGAACCTGGGTACATCCGGGAAGGCTTCCTGGCCGTGCAGCAT GCTGTGGACCGGGCCATCATGGAGTACCATGCCGATGCCGCCACACGCCAGCTGTTCCAGAGACTGACGGTG ACCATCAAGAGGTTCCCGTACCCGCCGTTCATCGCAGACCCCTTCCTCGTGGCCATCCAGTACCAGCTGCCCCT GCTGCTGCTGCTCAGCTTCACCTACACCGCGCTCACCATTGCCCGTGCTGTCGTGCAGGAGAAGGAAAGGAG GCTGAAGGAGTACATGCGCATGATGGGGCTCAGCAGCTGGCTGCACTGGAGTGCCTGGTTCCTCTTGTTCTTC CTCTTCCTCCTCATCGCCGCCTCCTTCATGACCCTGCTCTTCTGTGTCAAGGTGAAGCCAAATGTAGCCGTGCTG TCCCGCAGCGACCCCTCCCTGGTGCTCGCCTTCCTGCTGTGCTTCGCCATCTCTACCATCTCCTTCAGCTTCATG GTCAGCACCTTCTTCAGCAAAGCCAACATGGCAGCAGCCTTCGGAGGCTTCCTCTACTTCTTCACCTACATCCC CTACTTCTTCGTGGCCCCTCGGTACAACTGGATGACTCTGAGCCAGAAGCTCTGCTCCTGCCTCCTGTCTAATG TCGCCATGGCAATGGGAGCCCAGCTCATTGGGAAATTTGAGGCGAAAGGCATGGGCATCCAGTGGCGAGAC CTCCTGAGTCCCGTCAACGTGGACGACGACTTCTGCTTCGGGCAGGTGCTGGGGATGCTGCTGCTGGACTCTG TGCTCTATGGCCTGGTGACCTGGTACATGGAGGCCGTCTTCCCAGGGCAGTTCGGCGTGCCTCAGCCCTGGTA CTTCTTCATCATGCCCTCCTATTGGTGTGGGAAGCCAAGGGCGGTTGCAGGGAAGGAGGAAGAAGACAGTGA CCCCGAGAAAGCACTCAGAAACGAGTACTTTGAAGCCGAGCCAGAGGACCTGGTGGCGGGGATCAAGATCA AGCACCTGTCCAAGGTGTTCAGGGTGGGAAATAAGGACAGGGCGGCCGTCAGAGACCTGAACCTCAACCTGT ACGAGGGACAGATCACCGTCCTGCTGGGCCACAACGGTGCCGGGAAGACCACCACCCTCTCCATGCTCACAG GTCTCTTTCCCCCCACCAGTGGACGGGCATACATCAGCGGGTATGAAATTTCCCAGGACATGGTTCAGATCCG GAAGAGCCTGGGCCTGTGCCCGCAGCACGACATCCTGTTTGACAACTTGACAGTCGCAGAGCACCTTTATTTC TACGCCCAGCTGAAGGGCCTGTCACGTCAGAAGTGCCCTGAAGAAGTCAAGCAGATGCTGCACATCATCGGC CTGGAGGACAAGTGGAACTCACGGAGCCGCTTCCTGAGCGGGGGCATGAGGCGCAAGCTCTCCATCGGCATC GCCCTCATCGCAGGCTCCAAGGTGCTGATACTGGACGAGCCCACCTCGGGCATGGACGCCATCTCCAGGAGG GCCATCTGGGATCTTCTTCAGCGGCAGAAAAGTGACCGCACCATCGTGCTGACCACCCACTTCATGGACGAGG CTGACCTGCTGGGAGACCGCATCGCCATCATGGCCAAGGGGGAGCTGCAGTGCTGCGGGTCCTCGCTGTTCC TCAAGCAGAAATACGGTGCCGGCTATCACATGACGCTGGTGAAGGAGCCGCACTGCAACCCGGAAGACATCT CCCAGCTGGTCCACCACCACGTGCCCAACGCCACGCTGGAGAGCAGCGCTGGGGCCGAGCTGTCTTTCATCCT TCCCAGAGAGAGCACGCACAGGTTTGAAGGTCTCTTTGCTAAACTGGAGAAGAAGCAGAAAGAGCTGGGCAT TGCCAGCTTTGGGGCATCCATCACCACCATGGAGGAAGTCTTCCTTCGGGTCGGGAAGCTGGTGGACAGCAG TATGGACATCCAGGCCATCCAGCTCCCTGCCCTGCAGTACCAGCACGAGAGGCGCGCCAGCGACTGGGCTGT GGACAGCAACCTCTGTGGGGCCATGGACCCCTCCGACGGCATTGGAGCCCTCATCGAGGAGGAGCGCACCGC TGTCAAGCTCAACACTGGGCTCGCCCTGCACTGCCAGCAATTCTGGGCCATGTTCCTGAAGAAGGCCGCATAC AGCTGGCGCGAGTGGAAAATGGTGGCGGCACAGGTCCTGGTGCCTCTGACCTGCGTCACCCTGGCCCTCCTG GCCATCAACTACTCCTCGGAGCTCTTCGACGACCCCATGCTGAGGCTGACCTTGGGCGAGTACGGCAGAACCG TCGTGCCCTTCTCAGTTCCCGGGACCTCCCAGCTGGGTCAGCAGCTGTCAGAGCATCTGAAAGACGCACTGCA GGCTGAGGGACAGGAGCCCCGCGAGGTGCTCGGTGACCTGGAGGAGTTCTTGATCTTCAGGGCTTCTGTGGA GGGGGGCGGCTTTAATGAGCGGTGCCTTGTGGCAGCGTCCTTCAGAGATGTGGGAGAGCGCACGGTCGTCA ACGCCTTGTTCAACAACCAGGCGTACCACTCTCCAGCCACTGCCCTGGCCGTCGTGGACAACCTTCTGTTCAAG CTGCTGTGCGGGCCTCACGCCTCCATTGTGGTCTCCAACTTCCCCCAGCCCCGGAGCGCCCTGCAGGCTGCCA AGGACCAGTTTAACGAGGGCCGGAAGGGATTCGACATTGCCCTCAACCTGCTCTTCGCCATGGCATTCTTGGC CAGCACGTTCTCCATCCTGGCGGTCAGCGAGAGGGCCGTGCAGGCCAAGCATGTGCAGTTTGTGAGTGGAGT CCACGTGGCCAGTTTCTGGCTCTCTGCTCTGCTGTGGGACCTCATCTCCTTCCTCATCCCCAGTCTGCTGCTGCT GGTGGTGTTTAAGGCCTTCGACGTGCGTGCCTTCACGCGGGACGGCCACATGGCTGACACCCTGCTGCTGCTC CTGCTCTACGGCTGGGCCATCATCCCCCTCATGTACCTGATGAACTTCTTCTTCTTGGGGGCGGCCACTGCCTA CACGAGGCTGACCATCTTCAACATCCTGTCAGGCATCGCCACCTTCCTGATGGTCACCATCATGCGCATCCCAG CTGTAAAACTGGAAGAACTTTCCAAAACCCTGGATCACGTGTTCCTGGTGCTGCCCAACCACTGTCTGGGGAT GGCAGTCAGCAGTTTCTACGAGAACTACGAGACGCGGAGGTACTGCACCTCCTCCGAGGTCGCCGCCCACTA CTGCAAGAAATATAACATCCAGTACCAGGAGAACTTCTATGCCTGGAGCGCCCCGGGGGTCGGCCGGTTTGT GGCCTCCATGGCCGCCTCAGGGTGCGCCTACCTCATCCTGCTCTTCCTCATCGAGACCAACCTGCTTCAGAGAC TCAGGGGCATCCTCTGCGCCCTCCGGAGGAGGCGGACACTGACAGAATTATACACCCGGATGCCTGTGCTTCC TGAGGACCAAGATGTAGCGGACGAGAGGACCCGCATCCTGGCCCCCAGCCCGGACTCCCTGCTCCACACACC TCTGATTATCAAGGAGCTCTCCAAGGTGTACGAGCAGCGGGTGCCCCTCCTGGCCGTGGACAGGCTCTCCCTC GCGGTGCAGAAAGGGGAGTGCTTCGGCCTGCTGGGCTTCAATGGAGCCGGGAAGACCACGACTTTCAAAAT GCTGACCGGGGAGGAGAGCCTCACTTCTGGGGATGCCTTTGTCGGGGGTCACAGAATCAGCTCTGATGTCGG AAAGGTGCGGCAGCGGATCGGCTACTGCCCGCAGTTTGATGCCTTGCTGGACCACATGACAGGCCGGGAGAT GCTGGTCATGTACGCTCGGCTCCGGGGCATCCCTGAGCGCCACATCGGGGCCTGCGTGGAGAACACTCTGCG GGGCCTGCTGCTGGAGCCACATGCCAACAAGCTGGTCAGGACGTACAGTGGTGGTAACAAGCGGAAGCTGA GCACCGGCATCGCCCTGATCGGAGAGCCTGCTGTCATCTTCCTGGACGAGCCGTCCACTGGCATGGACCCCGT GGCCCGGCGCCTGCTTTGGGACACCGTGGCACGAGCCCGAGAGTCTGGCAAGGCCATCATCATCACCTCCCA CAGCATGGAGGAGTGTGAGGCCCTGTGCACCCGGCTGGCCATCATGGTGCAGGGGCAGTTCAAGTGCCTGG GCAGCCCCCAGCACCTCAAGAGCAAGTTCGGCAGCGGCTACTCCCTGCGGGCCAAGGTGCAGAGTGAAGGG CAACAGGAGGCGCTGGAGGAGTTCAAGGCCTTCGTGGACCTGACCTTTCCAGGCAGCGTCCTGGAAGATGAG CACCAAGGCATGGTCCATTACCACCTGCCGGGCCGTGACCTCAGCTGGGCGAAGGTTTTCGGTATTCTGGAGA AAGCCAAGGAAAAGTACGGCGTGGACGACTACTCCGTGAGCCAGATCTCGCTGGAACAGGTCTTCCTGAGCT TCGCCCACCTGCAGCCGCCCACCGCAGAGGAGGGGCGATGAGCGGCCGCGGGCCC    SEQ ID NO: 60 Exemplified codon‐optimised human ABACA3 (cohABCA3) transgene (5131bp)  GCTAGCCACCATGGCCGTGCTGCGCCAGCTGGCCCTGCTGCTGTGGAAGAACTACACCCTGCAGAAGCGCAA GGTGCTGGTGACCGTGCTGGAGCTGTTCCTGCCCCTGCTGTTCAGCGGCATCCTGATCTGGCTGCGCCTGAAG ATCCAGAGCGAGAACGTGCCCAACGCCACCATCTACCCCGGCCAGAGCATCCAGGAGCTGCCCCTGTTCTTCA CCTTCCCCCCCCCCGGCGACACCTGGGAGCTGGCCTACATCCCCAGCCACAGCGACGCCGCCAAGACCGTGAC CGAGACCGTGCGCCGCGCCCTGGTGATCAACATGCGCGTGCGCGGCTTCCCCAGCGAGAAGGACTTCGAGGA CTACATCCGCTACGACAACTGCAGCAGCAGCGTGCTGGCCGCCGTGGTGTTCGAGCACCCCTTCAACCACAGC AAGGAGCCCCTGCCCCTGGCCGTGAAGTACCACCTGCGCTTCAGCTACACCCGCCGCAACTACATGTGGACCC AGACCGGCAGCTTCTTCCTGAAGGAGACCGAGGGCTGGCACACCACCAGCCTGTTCCCCCTGTTCCCCAACCC CGGCCCCCGCGAGCCCACCAGCCCCGACGGCGGCGAGCCCGGCTACATCCGCGAGGGCTTCCTGGCCGTGCA GCACGCCGTGGACCGCGCCATCATGGAGTACCACGCCGACGCCGCCACCCGCCAGCTGTTCCAGCGCCTGAC CGTGACCATCAAGCGCTTCCCCTACCCCCCCTTCATCGCCGACCCCTTCCTGGTGGCCATCCAGTACCAGCTGC CCCTGCTGCTGCTGCTGAGCTTCACCTACACCGCCCTGACCATCGCCCGCGCCGTGGTGCAGGAGAAGGAGCG CCGCCTGAAGGAGTACATGCGCATGATGGGCCTGAGCAGCTGGCTGCACTGGAGCGCCTGGTTCCTGCTGTT CTTCCTGTTCCTGCTGATCGCCGCCAGCTTCATGACCCTGCTGTTCTGCGTGAAGGTGAAGCCCAACGTGGCCG TGCTGAGCCGCAGCGACCCCAGCCTGGTGCTGGCCTTCCTGCTGTGCTTCGCCATCAGCACCATCAGCTTCAG CTTCATGGTGAGCACCTTCTTCAGCAAGGCCAACATGGCCGCCGCCTTCGGCGGCTTCCTGTACTTCTTCACCT ACATCCCCTACTTCTTCGTGGCCCCCCGCTACAACTGGATGACCCTGAGCCAGAAGCTGTGCAGCTGCCTGCTG AGCAACGTGGCCATGGCCATGGGCGCCCAGCTGATCGGCAAGTTCGAGGCCAAGGGCATGGGCATCCAGTG GCGCGACCTGCTGAGCCCCGTGAACGTGGACGACGACTTCTGCTTCGGCCAGGTGCTGGGCATGCTGCTGCT GGACAGCGTGCTGTACGGCCTGGTGACCTGGTACATGGAGGCCGTGTTCCCCGGCCAGTTCGGCGTGCCCCA GCCCTGGTACTTCTTCATCATGCCCAGCTACTGGTGCGGCAAGCCCCGCGCCGTGGCCGGCAAGGAGGAGGA GGACAGCGACCCCGAGAAGGCCCTGCGCAACGAGTACTTCGAGGCCGAGCCCGAGGACCTGGTGGCCGGCA TCAAGATCAAGCACCTGAGCAAGGTGTTCCGCGTGGGCAACAAGGACCGCGCCGCCGTGCGCGACCTGAACC TGAACCTGTACGAGGGCCAGATCACCGTGCTGCTGGGCCACAACGGCGCCGGCAAGACCACCACCCTGAGCA TGCTGACCGGCCTGTTCCCCCCCACCAGCGGCAGGGCCTACATCAGCGGCTACGAGATCAGCCAGGACATGG TGCAGATCCGCAAGAGCCTGGGCCTGTGCCCCCAGCACGACATCCTGTTCGACAACCTGACCGTGGCCGAGC ACCTGTACTTCTACGCCCAGCTGAAGGGCCTGAGCCGCCAGAAGTGCCCCGAGGAGGTGAAGCAGATGCTGC ACATCATCGGCCTGGAGGACAAGTGGAACAGCCGCAGCCGCTTCCTGAGCGGCGGCATGCGCCGCAAGCTG AGCATCGGCATCGCCCTGATCGCCGGCAGCAAGGTGCTGATCCTGGACGAGCCCACCAGCGGCATGGACGCC ATCAGCCGCCGCGCCATCTGGGACCTGCTGCAGCGCCAGAAGAGCGACCGCACCATCGTGCTGACCACCCAC TTCATGGACGAGGCCGACCTGCTGGGCGACCGCATCGCCATCATGGCCAAGGGCGAGCTGCAGTGCTGCGGC AGCAGCCTGTTCCTGAAGCAGAAGTACGGCGCCGGCTACCACATGACCCTGGTGAAGGAGCCCCACTGCAAC CCCGAGGACATCAGCCAGCTGGTGCACCACCACGTGCCCAACGCCACCCTGGAGAGCAGCGCCGGCGCCGAG CTGAGCTTCATCCTGCCCCGCGAGAGCACCCACCGCTTCGAGGGCCTGTTCGCCAAGCTGGAGAAGAAGCAG AAGGAGCTGGGCATCGCCAGCTTCGGCGCCAGCATCACCACCATGGAGGAGGTGTTCCTGCGCGTGGGCAA GCTGGTGGACAGCAGCATGGACATCCAGGCCATCCAGCTGCCCGCCCTGCAGTACCAGCACGAGCGCCGCGC CAGCGACTGGGCCGTGGACAGCAACCTGTGCGGCGCCATGGACCCCAGCGACGGCATCGGCGCCCTGATCG AGGAGGAGCGCACCGCCGTGAAGCTGAACACCGGCCTGGCCCTGCACTGCCAGCAGTTCTGGGCCATGTTCC TGAAGAAGGCCGCCTACAGCTGGCGCGAGTGGAAGATGGTGGCCGCCCAGGTGCTGGTGCCCCTGACCTGC GTGACCCTGGCCCTGCTGGCCATCAACTACAGCAGCGAGCTGTTCGACGACCCCATGCTGCGCCTGACCCTGG GCGAGTACGGCCGCACCGTGGTGCCCTTCAGCGTGCCCGGCACCAGCCAGCTGGGCCAGCAGCTGAGCGAG CACCTGAAGGACGCCCTGCAGGCCGAGGGCCAGGAGCCCCGCGAGGTGCTGGGCGACCTGGAGGAGTTCCT GATCTTCCGCGCCAGCGTGGAGGGCGGCGGCTTCAACGAGCGCTGCCTGGTGGCCGCCAGCTTCCGCGACGT GGGCGAGCGCACCGTGGTGAACGCCCTGTTCAACAACCAGGCCTACCACAGCCCCGCCACCGCCCTGGCCGT GGTGGACAACCTGCTGTTCAAGCTGCTGTGCGGCCCCCACGCCAGCATCGTGGTGAGCAACTTCCCCCAGCCC CGCAGCGCCCTGCAGGCCGCCAAGGACCAGTTCAACGAGGGCCGCAAGGGCTTCGACATCGCCCTGAACCTG CTGTTCGCCATGGCCTTCCTGGCCAGCACCTTCAGCATCCTGGCCGTGAGCGAGAGGGCCGTGCAGGCCAAG CACGTGCAGTTCGTGAGCGGCGTGCACGTGGCCAGCTTCTGGCTGAGCGCCCTGCTGTGGGACCTGATCAGC TTCCTGATCCCCAGCCTGCTGCTGCTGGTGGTGTTCAAGGCCTTCGACGTGAGGGCCTTCACCCGCGACGGCC ACATGGCCGACACCCTGCTGCTGCTGCTGCTGTACGGCTGGGCCATCATCCCCCTGATGTACCTGATGAACTTC TTCTTCCTGGGCGCCGCCACCGCCTACACCCGCCTGACCATCTTCAACATCCTGAGCGGCATCGCCACCTTCCT GATGGTGACCATCATGCGCATCCCCGCCGTGAAGCTGGAGGAGCTGAGCAAGACCCTGGACCACGTGTTCCT GGTGCTGCCCAACCACTGCCTGGGCATGGCCGTGAGCAGCTTCTACGAGAACTACGAGACCCGCCGCTACTG CACCAGCAGCGAGGTGGCCGCCCACTACTGCAAGAAGTACAACATCCAGTACCAGGAGAACTTCTACGCCTG GAGCGCCCCCGGCGTGGGCCGCTTCGTGGCCAGCATGGCCGCCAGCGGCTGCGCCTACCTGATCCTGCTGTT CCTGATCGAGACCAACCTGCTGCAGCGCCTGCGCGGCATCCTGTGCGCCCTGCGCCGCCGCCGCACCCTGACC GAGCTGTACACCCGCATGCCCGTGCTGCCCGAGGACCAGGACGTGGCCGACGAGCGCACCCGCATCCTGGCC CCCAGCCCCGACAGCCTGCTGCACACCCCCCTGATCATCAAGGAGCTGAGCAAGGTGTACGAGCAGCGCGTG CCCCTGCTGGCCGTGGACCGCCTGAGCCTGGCCGTGCAGAAGGGCGAGTGCTTCGGCCTGCTGGGCTTCAAC GGCGCCGGCAAGACCACCACCTTCAAGATGCTGACCGGCGAGGAGAGCCTGACCAGCGGCGACGCCTTCGT GGGCGGCCACCGCATCAGCAGCGACGTGGGCAAGGTGCGCCAGCGCATCGGCTACTGCCCCCAGTTCGACGC CCTGCTGGACCACATGACCGGCCGCGAGATGCTGGTGATGTACGCCCGCCTGCGCGGCATCCCCGAGCGCCA CATCGGCGCCTGCGTGGAGAACACCCTGCGCGGCCTGCTGCTGGAGCCCCACGCCAACAAGCTGGTGCGCAC CTACAGCGGCGGCAACAAGCGCAAGCTGAGCACCGGCATCGCCCTGATCGGCGAGCCCGCCGTGATCTTCCT GGACGAGCCCAGCACCGGCATGGACCCCGTGGCCCGCCGCCTGCTGTGGGACACCGTGGCCCGCGCCCGCG AGAGCGGCAAGGCCATCATCATCACCAGCCACAGCATGGAGGAGTGCGAGGCCCTGTGCACCCGCCTGGCCA TCATGGTGCAGGGCCAGTTCAAGTGCCTGGGCAGCCCCCAGCACCTGAAGAGCAAGTTCGGCAGCGGCTACA GCCTGAGGGCCAAGGTGCAGAGCGAGGGCCAGCAGGAGGCCCTGGAGGAGTTCAAGGCCTTCGTGGACCTG ACCTTCCCCGGCAGCGTGCTGGAGGACGAGCACCAGGGCATGGTGCACTACCACCTGCCCGGCCGCGACCTG AGCTGGGCCAAGGTGTTCGGCATCCTGGAGAAGGCCAAGGAGAAGTACGGCGTGGACGACTACAGCGTGAG CCAGATCAGCCTGGAGCAGGTGTTCCTGAGCTTCGCCCACCTGCAGCCCCCCACCGCCGAGGAGGGCCGCTA AGGGCCC    SEQ ID NO: 61 Exemplified Human ABCA3 polypeptide  MAVLRQLALLLWKNYTLQKRKVLVTVLELFLPLLFSGILIWLRLKIQSENVPNATIYPGQSIQELPLFFTFPPPGDTWE LAYIPSHSDAAKTVTETVRRALVINMRVRGFPSEKDFEDYIRYDNCSSSVLAAVVFEHPFNHSKEPLPLAVKYHLRFSY TRRNYMWTQTGSFFLKETEGWHTTSLFPLFPNPGPREPTSPDGGEPGYIREGFLAVQHAVDRAIMEYHADAATR QLFQRLTVTIKRFPYPPFIADPFLVAIQYQLPLLLLLSFTYTALTIARAVVQEKERRLKEYMRMMGLSSWLHWSAWFL LFFLFLLIAASFMTLLFCVKVKPNVAVLSRSDPSLVLAFLLCFAISTISFSFMVSTFFSKANMAAAFGGFLYFFTYIPYFF VAPRYNWMTLSQKLCSCLLSNVAMAMGAQLIGKFEAKGMGIQWRDLLSPVNVDDDFCFGQVLGMLLLDSVLYG LVTWYMEAVFPGQFGVPQPWYFFIMPSYWCGKPRAVAGKEEEDSDPEKALRNEYFEAEPEDLVAGIKIKHLSKVF RVGNKDRAAVRDLNLNLYEGQITVLLGHNGAGKTTTLSMLTGLFPPTSGRAYISGYEISQDMVQIRKSLGLCPQHDI LFDNLTVAEHLYFYAQLKGLSRQKCPEEVKQMLHIIGLEDKWNSRSRFLSGGMRRKLSIGIALIAGSKVLILDEPTSG MDAISRRAIWDLLQRQKSDRTIVLTTHFMDEADLLGDRIAIMAKGELQCCGSSLFLKQKYGAGYHMTLVKEPHCN PEDISQLVHHHVPNATLESSAGAELSFILPRESTHRFEGLFAKLEKKQKELGIASFGASITTMEEVFLRVGKLVDSSMD IQAIQLPALQYQHERRASDWAVDSNLCGAMDPSDGIGALIEEERTAVKLNTGLALHCQQFWAMFLKKAAYSWRE WKMVAAQVLVPLTCVTLALLAINYSSELFDDPMLRLTLGEYGRTVVPFSVPGTSQLGQQLSEHLKDALQAEGQEPR EVLGDLEEFLIFRASVEGGGFNERCLVAASFRDVGERTVVNALFNNQAYHSPATALAVVDNLLFKLLCGPHASIVVS NFPQPRSALQAAKDQFNEGRKGFDIALNLLFAMAFLASTFSILAVSERAVQAKHVQFVSGVHVASFWLSALLWDLI SFLIPSLLLLVVFKAFDVRAFTRDGHMADTLLLLLLYGWAIIPLMYLMNFFFLGAATAYTRLTIFNILSGIATFLMVTIM RIPAVKLEELSKTLDHVFLVLPNHCLGMAVSSFYENYETRRYCTSSEVAAHYCKKYNIQYQENFYAWSAPGVGRFVA SMAASGCAYLILLFLIETNLLQRLRGILCALRRRRTLTELYTRMPVLPEDQDVADERTRILAPSPDSLLHTPLIIKELSKV YEQRVPLLAVDRLSLAVQKGECFGLLGFNGAGKTTTFKMLTGEESLTSGDAFVGGHRISSDVGKVRQRIGYCPQFD ALLDHMTGREMLVMYARLRGIPERHIGACVENTLRGLLLEPHANKLVRTYSGGNKRKLSTGIALIGEPAVIFLDEPST GMDPVARRLLWDTVARARESGKAIIITSHSMEECEALCTRLAIMVQGQFKCLGSPQHLKSKFGSGYSLRAKVQSEG QQEALEEFKAFVDLTFPGSVLEDEHQGMVHYHLPGRDLSWAKVFGILEKAKEKYGVDDYSVSQISLEQVFLSFAHL QPPTAEEGR*    SEQ ID NO: 62 Exemplified Human SFTPC transgene  GCTAGCCACCATGGATGTGGGCAGCAAAGAGGTCCTGATGGAGAGCCCGCCGGACTACTCCGCAGCTCCCCG GGGCCGATTTGGCATTCCCTGCTGCCCAGTGCACCTGAAACGCCTTCTTATCGTGGTGGTGGTGGTGGTCCTC ATCGTCGTGGTGATTGTGGGAGCCCTGCTCATGGGTCTCCACATGAGCCAGAAACACACGGAGATGGTTCTG GAGATGAGCATTGGGGCGCCGGAAGCCCAGCAACGCCTGGCCCTGAGTGAGCACCTGGTTACCACTGCCACC TTCTCCATCGGCTCCACTGGCCTCGTGGTGTATGACTACCAGCAGCTGCTGATCGCCTACAAGCCAGCCCCTGG CACCTGCTGCTACATCATGAAGATAGCTCCAGAGAGCATCCCCAGTCTTGAGGCTCTCACTAGAAAAGTCCAC AACTTCCAGGCCAAGCCCGCAGTGCCTACGTCTAAGCTGGGCCAGGCAGAGGGGCGAGATGCAGGCTCAGC ACCCTCCGGAGGGGACCCGGCCTTCCTGGGCATGGCCGTGAGCACCCTGTGTGGCGAGGTGCCGCTCTACTA CATCTAGGGGCCC    Underlined sequence = combined NheI and Kozak sequence  Bold and dashed‐underlined sequence = ApaI site  Either or both of the underlined and/or bold+dashed‐underlined sequences may be omitted and  such a sequence still expressly falls within the scope of the invention (SEQ ID NOs: 84‐86).    SEQ ID NO: 63 Exemplified Human SFTPC polypeptide  MDVGSKEVLMESPPDYSAAPRGRFGIPCCPVHLKRLLIVVVVVVLIVVVIVGALLMGLHMSQKHTEMVLEMSIGAP EAQQRLALSEHLVTTATFSIGSTGLVVYDYQQLLIAYKPAPGTCCYIMKIAPESIPSLEALTRKVHNFQAKPAVPTSKL GQAEGRDAGSAPSGGDPAFLGMAVSTLCGEVPLYYI*    SEQ ID NO: 64 Exemplified Human GM‐CSF (CSF2) transgene  GCTAGCCACCATGTGGCTGCAGAGCCTGCTGCTCTTGGGCACTGTGGCCTGCAGCATCTCTGCACCCGCCCGC TCGCCCAGCCCCAGCACGCAGCCCTGGGAGCATGTGAATGCCATCCAGGAGGCCCGGCGTCTCCTGAACCTG AGTAGAGACACTGCTGCTGAGATGAATGAAACAGTAGAAGTCATCTCAGAAATGTTTGACCTCCAGGAGCCG ACCTGCCTACAGACCCGCCTGGAGCTGTACAAGCAGGGCCTGCGGGGCAGCCTCACCAAGCTCAAGGGCCCC TTGACCATGATGGCCAGCCACTACAAGCAGCACTGCCCTCCAACCCCGGAAACTTCCTGTGCAACCCAGATTAT CACCTTTGAAAGTTTCAAAGAGAACCTGAAGGACTTTCTGCTTGTCATCCCCTTTGACTGCTGGGAGCCAGTCC AGGAGTGAGGGCCC    SEQ ID NO: 65  Exemplified Human GM‐CSF polypeptide  MWLQSLLLLGTVACSISAPARSPSPSTQPWEHVNAIQEARRLLNLSRDTAAEMNETVEVISEMFDLQEPTCLQTRL ELYKQGLRGSLTKLKGPLTMMASHYKQHCPPTPETSCATQIITFESFKENLKDFLLVIPFDCWEPVQE*    SEQ ID NO: 66 Exemplified Human DCN (Decorin) transgene  GCTAGCCACCATGAAGGCCACTATCATCCTCCTTCTGCTTGCACAAGTTTCCTGGGCTGGACCGTTTCAACAGA GAGGCTTATTTGACTTTATGCTAGAAGATGAGGCTTCTGGGATAGGCCCAGAAGTTCCTGATGACCGCGACTT CGAGCCCTCCCTAGGCCCAGTGTGCCCCTTCCGCTGTCAATGCCATCTTCGAGTGGTCCAGTGTTCTGATTTGG GTCTGGACAAAGTGCCAAAGGATCTTCCCCCTGACACAACTCTGCTAGACCTGCAAAACAACAAAATAACCGA AATCAAAGATGGAGACTTTAAGAACCTGAAGAACCTTCACGCATTGATTCTTGTCAACAATAAAATTAGCAAA GTTAGTCCTGGAGCATTTACACCTTTGGTGAAGTTGGAACGACTTTATCTGTCCAAGAATCAGCTGAAGGAAT TGCCAGAAAAAATGCCCAAAACTCTTCAGGAGCTGCGTGCCCATGAGAATGAGATCACCAAAGTGCGAAAAG TTACTTTCAATGGACTGAACCAGATGATTGTCATAGAACTGGGCACCAATCCGCTGAAGAGCTCAGGAATTGA AAATGGGGCTTTCCAGGGAATGAAGAAGCTCTCCTACATCCGCATTGCTGATACCAATATCACCAGCATTCCTC AAGGTCTTCCTCCTTCCCTTACGGAATTACATCTTGATGGCAACAAAATCAGCAGAGTTGATGCAGCAAGCCTG AAAGGACTGAATAATTTGGCTAAGTTGGGATTGAGTTTCAACAGCATCTCTGCTGTTGACAATGGCTCTCTGG CCAACACGCCTCATCTGAGGGAGCTTCACTTGGACAACAACAAGCTTACCAGAGTACCTGGTGGGCTGGCAG AGCATAAGTACATCCAGGTTGTCTACCTTCATAACAACAATATCTCTGTAGTTGGATCAAGTGACTTCTGCCCA CCTGGACACAACACCAAAAAGGCTTCTTATTCGGGTGTGAGTCTTTTCAGCAACCCGGTCCAGTACTGGGAGA TACAGCCATCCACCTTCAGATGTGTCTACGTGCGCTCTGCCATTCAACTCGGAAACTATAAGTAAGGGCCCA    SEQ ID NO: 67 Exemplified Human Decorin polypeptide  MKATIILLLLAQVSWAGPFQQRGLFDFMLEDEASGIGPEVPDDRDFEPSLGPVCPFRCQCHLRVVQCSDLGLDKVP KDLPPDTTLLDLQNNKITEIKDGDFKNLKNLHALILVNNKISKVSPGAFTPLVKLERLYLSKNQLKELPEKMPKTLQEL RAHENEITKVRKVTFNGLNQMIVIELGTNPLKSSGIENGAFQGMKKLSYIRIADTNITSIPQGLPPSLTELHLDGNKIS RVDAASLKGLNNLAKLGLSFNSISAVDNGSLANTPHLRELHLDNNKLTRVPGGLAEHKYIQVVYLHNNNISVVGSSD FCPPGHNTKKASYSGVSLFSNPVQYWEIQPSTFRCVYVRSAIQLGNYK*    SEQ ID NO: 68 Exemplified Human TRIM72 transgene  GCTAGCCACCATGTCGGCTGCGCCCGGCCTCCTGCACCAGGAGCTGTCCTGCCCGCTGTGCCTGCAGCTGTTC GACGCGCCCGTGACAGCCGAGTGCGGCCACAGTTTCTGCCGCGCCTGCCTAGGCCGCGTGGCCGGGGAGCC GGCGGCGGATGGCACCGTTCTCTGCCCCTGCTGCCAGGCCCCCACGCGGCCGCAGGCACTCAGCACCAACCT GCAGCTGGCGCGCCTGGTGGAGGGGCTGGCCCAGGTGCCGCAGGGCCACTGCGAGGAGCACCTGGACCCGC TGAGCATCTACTGCGAGCAGGACCGCGCGCTGGTGTGCGGAGTGTGCGCCTCACTCGGCTCGCACCGCGGTC ATCGCCTCCTGCCTGCCGCCGAGGCCCACGCACGCCTCAAGACACAGCTGCCACAGCAGAAACTGCAGCTGCA GGAGGCATGCATGCGCAAGGAGAAGAGTGTGGCTGTGCTGGAGCATCAGCTGGTGGAGGTGGAGGAGACA GTGCGTCAGTTCCGGGGGGCCGTGGGGGAGCAGCTGGGCAAGATGCGGGTGTTCCTGGCTGCACTGGAGG GCTCCTTGGACCGCGAGGCAGAGCGTGTACGGGGTGAGGCAGGGGTCGCCTTGCGCCGGGAGCTGGGGAG CCTGAACTCTTACCTGGAGCAGCTGCGGCAGATGGAGAAGGTCCTGGAGGAGGTGGCGGACAAGCCGCAGA CTGAGTTCCTCATGAAATACTGCCTGGTGACCAGCAGGCTGCAGAAGATCCTGGCAGAGTCTCCCCCACCCGC CCGTCTGGACATCCAGCTGCCAATTATCTCAGATGACTTCAAATTCCAGGTGTGGAGGAAGATGTTCCGGGCT CTGATGCCAGCGCTGGAGGAGCTGACCTTTGACCCGAGCTCTGCGCACCCGAGCCTGGTGGTGTCTTCCTCTG GCCGCCGCGTGGAGTGCTCGGAGCAGAAGGCGCCGCCGGCCGGGGAGGACCCGCGCCAGTTCGACAAGGC GGTGGCGGTGGTGGCGCACCAGCAGCTCTCCGAGGGCGAGCACTACTGGGAGGTGGATGTTGGCGACAAGC CGCGCTGGGCGCTGGGCGTGATCGCGGCCGAGGCCCCCCGCCGCGGGCGCCTGCACGCGGTGCCCTCGCAG GGCCTGTGGCTGCTGGGGCTGCGCGAGGGCAAGATCCTGGAGGCACACGTGGAGGCCAAGGAGCCGCGCG CTCTGCGCAGCCCCGAGAGGCGGCCCACGCGCATTGGCCTTTACCTGAGCTTCGGCGACGGCGTCCTCTCCTT CTACGATGCCAGCGACGCCGACGCGCTCGTGCCGCTTTTTGCCTTCCACGAGCGCCTGCCCAGGCCCGTGTAC CCCTTCTTCGACGTGTGCTGGCACGACAAGGGCAAGAATGCCCAGCCGCTGCTGCTCGTGGGTCCCGAAGGC GCCGAGGCCTGAGGGCCC    SEQ ID NO: 69 Exemplified Human TRIM72 polypeptide  MSAAPGLLHQELSCPLCLQLFDAPVTAECGHSFCRACLGRVAGEPAADGTVLCPCCQAPTRPQALSTNLQLARLVE GLAQVPQGHCEEHLDPLSIYCEQDRALVCGVCASLGSHRGHRLLPAAEAHARLKTQLPQQKLQLQEACMRKEKSV AVLEHQLVEVEETVRQFRGAVGEQLGKMRVFLAALEGSLDREAERVRGEAGVALRRELGSLNSYLEQLRQMEKVL EEVADKPQTEFLMKYCLVTSRLQKILAESPPPARLDIQLPIISDDFKFQVWRKMFRALMPALEELTFDPSSAHPSLVV SSSGRRVECSEQKAPPAGEDPRQFDKAVAVVAHQQLSEGEHYWEVDVGDKPRWALGVIAAEAPRRGRLHAVPS QGLWLLGLREGKILEAHVEAKEPRALRSPERRPTRIGLYLSFGDGVLSFYDASDADALVPLFAFHERLPRPVYPFFDV CWHDKGKNAQPLLLVGPEGAEA    SEQ ID NO: 70 Exemplified SERPINA1 (AAT) transgene  ATGCCCAGCTCTGTGTCCTGGGGCATTCTGCTGCTGGCTGGCCTGTGCTGTCTGGTGCCTGTGTCCCTGG CTGAGGACCCTCAGGGGGATGCTGCCCAGAAAACAGACACCTCCCACCATGACCAGGACCACCCCACCTT CAACAAGATCACCCCCAACCTGGCAGAGTTTGCCTTCAGCCTGTACAGACAGCTGGCCCACCAGAGCAAC AGCACCAACATCTTTTTCAGCCCTGTGTCCATTGCCACAGCCTTTGCCATGCTGAGCCTGGGCACCAAGG CTGACACCCATGATGAGATCCTGGAAGGCCTGAACTTCAACCTGACAGAGATCCCTGAGGCCCAGATCCA TGAGGGCTTCCAGGAACTGCTGAGAACCCTGAACCAGCCAGACAGCCAGCTGCAGCTGACAACAGGCAAT GGGCTGTTCCTGTCTGAGGGCCTGAAGCTGGTGGACAAGTTTCTGGAAGATGTGAAGAAGCTGTACCACT CTGAGGCCTTCACAGTGAACTTTGGGGACACAGAAGAGGCCAAGAAACAGATCAATGACTATGTGGAAAA GGGCACCCAGGGCAAGATTGTGGACCTTGTGAAAGAGCTGGACAGGGACACTGTGTTTGCCCTTGTGAAC TACATCTTCTTCAAGGGCAAGTGGGAGAGGCCCTTTGAAGTGAAGGACACTGAGGAAGAGGACTTCCATG TGGACCAAGTGACCACAGTGAAGGTGCCAATGATGAAGAGACTGGGGATGTTCAATATCCAGCACTGCAA GAAACTGAGCAGCTGGGTGCTGCTGATGAAGTACCTGGGCAATGCTACAGCCATATTCTTTCTGCCTGAT GAGGGCAAGCTGCAGCACCTGGAAAATGAGCTGACCCATGACATCATCACCAAATTTCTGGAAAATGAGG ACAGAAGATCTGCCAGCCTGCATCTGCCCAAGCTGAGCATCACAGGCACATATGACCTGAAGTCTGTGCT GGGACAGCTGGGAATCACCAAGGTGTTCAGCAATGGGGCAGACCTGAGTGGAGTGACAGAGGAAGCCCCT CTGAAGCTGTCCAAGGCTGTGCACAAGGCAGTGCTGACCATTGATGAGAAGGGCACAGAGGCTGCTGGGG CCATGTTTCTGGAAGCCATCCCCATGTCCATCCCCCCAGAAGTGAAGTTCAACAAGCCCTTTGTGTTCCT GATGATTGAGCAGAACACCAAGAGCCCCCTGTTCATGGGCAAGGTTGTGAACCCCACCCAGAAATGA SEQ ID NO: 71 Exemplified AAT polypeptide  MPSSVSWGILLLAGLCCLVPVSLAEDPQGDAAQKTDTSHHDQDHPTFAEDPQGDAAQKTDTSHHDQDH PTFNKITPNLAEFAFSLYRQLAHQSNSTNIFFSPVSIATAFAMLSLGTKADTHDEILEGLNFNLTEIP EAQIHEGFQELLRTLNQPDSQLQLTTGNGLFLSEGLKLVDKFLEDVKKLYHSEAFTVNFGDTEEAKKQ INDYVEKGTQGKIVDLVKELDRDTVFALVNYIFFKGKWERPFEVKDTEEEDFHVDQVTTVKVPMMKRL GMFNIQHCKKLSSWVLLMKYLGNATAIFFLPDEGKLQHLENELTHDIITKFLENEDRRSASLHLPKLS ITGTYDLKSVLGQLGITKVFSNGADLSGVTEEAPLKLSKAVHKAVLTIDEKGTEAAGAMFLEAIPMSI PPEVKFNKPFVFLMIEQNTKSPLFMGKVVNPTQK   SEQ ID NO: 72 Exemplified FVIII transgene (N6)  ATGCAGATTGAGCTGAGCACCTGCTTCTTCCTGTGCCTGCTGAGGTTCTGCTTCTCTGCCACCAGGAGAT  ACTACCTGGGGGCTGTGGAGCTGAGCTGGGACTACATGCAGTCTGACCTGGGGGAGCTGCCTGTGGATGC  CAGGTTCCCCCCCAGAGTGCCCAAGAGCTTCCCCTTCAACACCTCTGTGGTGTACAAGAAGACCCTGTTT  GTGGAGTTCACTGACCACCTGTTCAACATTGCCAAGCCCAGGCCCCCCTGGATGGGCCTGCTGGGCCCCA  CCATCCAGGCTGAGGTGTATGACACTGTGGTGATCACCCTGAAGAACATGGCCAGCCACCCTGTGAGCCT  GCATGCTGTGGGGGTGAGCTACTGGAAGGCCTCTGAGGGGGCTGAGTATGATGACCAGACCAGCCAGAGG  GAGAAGGAGGATGACAAGGTGTTCCCTGGGGGCAGCCACACCTATGTGTGGCAGGTGCTGAAGGAGAATG  GCCCCATGGCCTCTGACCCCCTGTGCCTGACCTACAGCTACCTGAGCCATGTGGACCTGGTGAAGGACCT  GAACTCTGGCCTGATTGGGGCCCTGCTGGTGTGCAGGGAGGGCAGCCTGGCCAAGGAGAAGACCCAGACC  CTGCACAAGTTCATCCTGCTGTTTGCTGTGTTTGATGAGGGCAAGAGCTGGCACTCTGAAACCAAGAACA  GCCTGATGCAGGACAGGGATGCTGCCTCTGCCAGGGCCTGGCCCAAGATGCACACTGTGAATGGCTATGT  GAACAGGAGCCTGCCTGGCCTGATTGGCTGCCACAGGAAGTCTGTGTACTGGCATGTGATTGGCATGGGC  ACCACCCCTGAGGTGCACAGCATCTTCCTGGAGGGCCACACCTTCCTGGTCAGGAACCACAGGCAGGCCA  GCCTGGAGATCAGCCCCATCACCTTCCTGACTGCCCAGACCCTGCTGATGGACCTGGGCCAGTTCCTGCT  GTTCTGCCACATCAGCAGCCACCAGCATGATGGCATGGAGGCCTATGTGAAGGTGGACAGCTGCCCTGAG  GAGCCCCAGCTGAGGATGAAGAACAATGAGGAGGCTGAGGACTATGATGATGACCTGACTGACTCTGAGA  TGGATGTGGTGAGGTTTGATGATGACAACAGCCCCAGCTTCATCCAGATCAGGTCTGTGGCCAAGAAGCA  CCCCAAGACCTGGGTGCACTACATTGCTGCTGAGGAGGAGGACTGGGACTATGCCCCCCTGGTGCTGGCC  CCTGATGACAGGAGCTACAAGAGCCAGTACCTGAACAATGGCCCCCAGAGGATTGGCAGGAAGTACAAGA  AGGTCAGGTTCATGGCCTACACTGATGAAACCTTCAAGACCAGGGAGGCCATCCAGCATGAGTCTGGCAT  CCTGGGCCCCCTGCTGTATGGGGAGGTGGGGGACACCCTGCTGATCATCTTCAAGAACCAGGCCAGCAGG  CCCTACAACATCTACCCCCATGGCATCACTGATGTGAGGCCCCTGTACAGCAGGAGGCTGCCCAAGGGGG  TGAAGCACCTGAAGGACTTCCCCATCCTGCCTGGGGAGATCTTCAAGTACAAGTGGACTGTGACTGTGGA  GGATGGCCCCACCAAGTCTGACCCCAGGTGCCTGACCAGATACTACAGCAGCTTTGTGAACATGGAGAGG  GACCTGGCCTCTGGCCTGATTGGCCCCCTGCTGATCTGCTACAAGGAGTCTGTGGACCAGAGGGGCAACC  AGATCATGTCTGACAAGAGGAATGTGATCCTGTTCTCTGTGTTTGATGAGAACAGGAGCTGGTACCTGAC  TGAGAACATCCAGAGGTTCCTGCCCAACCCTGCTGGGGTGCAGCTGGAGGACCCTGAGTTCCAGGCCAGC  AACATCATGCACAGCATCAATGGCTATGTGTTTGACAGCCTGCAGCTGTCTGTGTGCCTGCATGAGGTGG  CCTACTGGTACATCCTGAGCATTGGGGCCCAGACTGACTTCCTGTCTGTGTTCTTCTCTGGCTACACCTT  CAAGCACAAGATGGTGTATGAGGACACCCTGACCCTGTTCCCCTTCTCTGGGGAGACTGTGTTCATGAGC  ATGGAGAACCCTGGCCTGTGGATTCTGGGCTGCCACAACTCTGACTTCAGGAACAGGGGCATGACTGCCC  TGCTGAAAGTCTCCAGCTGTGACAAGAACACTGGGGACTACTATGAGGACAGCTATGAGGACATCTCTGC  CTACCTGCTGAGCAAGAACAATGCCATTGAGCCCAGGAGCTTCAGCCAGAACAGCAGGCACCCCAGCACC  AGGCAGAAGCAGTTCAATGCCACCACCATCCCTGAGAATGACATAGAGAAGACAGACCCATGGTTTGCCC  ACCGGACCCCCATGCCCAAGATCCAGAATGTGAGCAGCTCTGACCTGCTGATGCTGCTGAGGCAGAGCCC  CACCCCCCATGGCCTGAGCCTGTCTGACCTGCAGGAGGCCAAGTATGAAACCTTCTCTGATGACCCCAGC  CCTGGGGCCATTGACAGCAACAACAGCCTGTCTGAGATGACCCACTTCAGGCCCCAGCTGCACCACTCTG  GGGACATGGTGTTCACCCCTGAGTCTGGCCTGCAGCTGAGGCTGAATGAGAAGCTGGGCACCACTGCTGC  CACTGAGCTGAAGAAGCTGGACTTCAAAGTCTCCAGCACCAGCAACAACCTGATCAGCACCATCCCCTCT  GACAACCTGGCTGCTGGCACTGACAACACCAGCAGCCTGGGCCCCCCCAGCATGCCTGTGCACTATGACA  GCCAGCTGGACACCACCCTGTTTGGCAAGAAGAGCAGCCCCCTGACTGAGTCTGGGGGCCCCCTGAGCCT  GTCTGAGGAGAACAATGACAGCAAGCTGCTGGAGTCTGGCCTGATGAACAGCCAGGAGAGCAGCTGGGGC  AAGAATGTGAGCAGCAGGGAGATCACCAGGACCACCCTGCAGTCTGACCAGGAGGAGATTGACTATGATG  ACACCATCTCTGTGGAGATGAAGAAGGAGGACTTTGACATCTACGACGAGGACGAGAACCAGAGCCCCAG  GAGCTTCCAGAAGAAGACCAGGCACTACTTCATTGCTGCTGTGGAGAGGCTGTGGGACTATGGCATGAGC  AGCAGCCCCCATGTGCTGAGGAACAGGGCCCAGTCTGGCTCTGTGCCCCAGTTCAAGAAGGTGGTGTTCC  AGGAGTTCACTGATGGCAGCTTCACCCAGCCCCTGTACAGAGGGGAGCTGAATGAGCACCTGGGCCTGCT  GGGCCCCTACATCAGGGCTGAGGTGGAGGACAACATCATGGTGACCTTCAGGAACCAGGCCAGCAGGCCC  TACAGCTTCTACAGCAGCCTGATCAGCTATGAGGAGGACCAGAGGCAGGGGGCTGAGCCCAGGAAGAACT  TTGTGAAGCCCAATGAAACCAAGACCTACTTCTGGAAGGTGCAGCACCACATGGCCCCCACCAAGGATGA  GTTTGACTGCAAGGCCTGGGCCTACTTCTCTGATGTGGACCTGGAGAAGGATGTGCACTCTGGCCTGATT  GGCCCCCTGCTGGTGTGCCACACCAACACCCTGAACCCTGCCCATGGCAGGCAGGTGACTGTGCAGGAGT  TTGCCCTGTTCTTCACCATCTTTGATGAAACCAAGAGCTGGTACTTCACTGAGAACATGGAGAGGAACTG  CAGGGCCCCCTGCAACATCCAGATGGAGGACCCCACCTTCAAGGAGAACTACAGGTTCCATGCCATCAAT  GGCTACATCATGGACACCCTGCCTGGCCTGGTGATGGCCCAGGACCAGAGGATCAGGTGGTACCTGCTGA  GCATGGGCAGCAATGAGAACATCCACAGCATCCACTTCTCTGGCCATGTGTTCACTGTGAGGAAGAAGGA  GGAGTACAAGATGGCCCTGTACAACCTGTACCCTGGGGTGTTTGAGACTGTGGAGATGCTGCCCAGCAAG  GCTGGCATCTGGAGGGTGGAGTGCCTGATTGGGGAGCACCTGCATGCTGGCATGAGCACCCTGTTCCTGG  TGTACAGCAACAAGTGCCAGACCCCCCTGGGCATGGCCTCTGGCCACATCAGGGACTTCCAGATCACTGC  CTCTGGCCAGTATGGCCAGTGGGCCCCCAAGCTGGCCAGGCTGCACTACTCTGGCAGCATCAATGCCTGG  AGCACCAAGGAGCCCTTCAGCTGGATCAAGGTGGACCTGCTGGCCCCCATGATCATCCATGGCATCAAGA  CCCAGGGGGCCAGGCAGAAGTTCAGCAGCCTGTACATCAGCCAGTTCATCATCATGTACAGCCTGGATGG  CAAGAAGTGGCAGACCTACAGGGGCAACAGCACTGGCACCCTGATGGTGTTCTTTGGCAATGTGGACAGC  TCTGGCATCAAGCACAACATCTTCAACCCCCCCATCATTGCCAGATACATCAGGCTGCACCCCACCCACT  ACAGCATCAGGAGCACCCTGAGGATGGAGCTGATGGGCTGTGACCTGAACAGCTGCAGCATGCCCCTGGG  CATGGAGAGCAAGGCCATCTCTGATGCCCAGATCACTGCCAGCAGCTACTTCACCAACATGTTTGCCACC  TGGAGCCCCAGCAAGGCCAGGCTGCACCTGCAGGGCAGGAGCAATGCCTGGAGGCCCCAGGTCAACAACC  CCAAGGAGTGGCTGCAGGTGGACTTCCAGAAGACCATGAAGGTGACTGGGGTGACCACCCAGGGGGTGAA  GAGCCTGCTGACCAGCATGTATGTGAAGGAGTTCCTGATCAGCAGCAGCCAGGATGGCCACCAGTGGACC  CTGTTCTTCCAGAATGGCAAGGTGAAGGTGTTCCAGGGCAACCAGGACAGCTTCACCCCTGTGGTGAACA  GCCTGGACCCCCCCCTGCTGACCAGATACCTGAGGATTCACCCCCAGAGCTGGGTGCACCAGATTGCCCT  GAGGATGGAGGTGCTGGGCTGTGAGGCCCAGGACCTGTACTGA    SEQ ID NO: 73 Exemplified FVIII transgene (V3)  ATGCAGATTGAGCTGAGCACCTGCTTCTTCCTGTGCCTGCTGAGGTTCTGCTTCTCTGCCACCAGGAGAT  ACTACCTGGGGGCTGTGGAGCTGAGCTGGGACTACATGCAGTCTGACCTGGGGGAGCTGCCTGTGGATGC  CAGGTTCCCCCCCAGAGTGCCCAAGAGCTTCCCCTTCAACACCTCTGTGGTGTACAAGAAGACCCTGTTT  GTGGAGTTCACTGACCACCTGTTCAACATTGCCAAGCCCAGGCCCCCCTGGATGGGCCTGCTGGGCCCCA  CCATCCAGGCTGAGGTGTATGACACTGTGGTGATCACCCTGAAGAACATGGCCAGCCACCCTGTGAGCCT  GCATGCTGTGGGGGTGAGCTACTGGAAGGCCTCTGAGGGGGCTGAGTATGATGACCAGACCAGCCAGAGG  GAGAAGGAGGATGACAAGGTGTTCCCTGGGGGCAGCCACACCTATGTGTGGCAGGTGCTGAAGGAGAATG  GCCCCATGGCCTCTGACCCCCTGTGCCTGACCTACAGCTACCTGAGCCATGTGGACCTGGTGAAGGACCT  GAACTCTGGCCTGATTGGGGCCCTGCTGGTGTGCAGGGAGGGCAGCCTGGCCAAGGAGAAGACCCAGACC  CTGCACAAGTTCATCCTGCTGTTTGCTGTGTTTGATGAGGGCAAGAGCTGGCACTCTGAAACCAAGAACA  GCCTGATGCAGGACAGGGATGCTGCCTCTGCCAGGGCCTGGCCCAAGATGCACACTGTGAATGGCTATGT  GAACAGGAGCCTGCCTGGCCTGATTGGCTGCCACAGGAAGTCTGTGTACTGGCATGTGATTGGCATGGGC  ACCACCCCTGAGGTGCACAGCATCTTCCTGGAGGGCCACACCTTCCTGGTCAGGAACCACAGGCAGGCCA  GCCTGGAGATCAGCCCCATCACCTTCCTGACTGCCCAGACCCTGCTGATGGACCTGGGCCAGTTCCTGCT  GTTCTGCCACATCAGCAGCCACCAGCATGATGGCATGGAGGCCTATGTGAAGGTGGACAGCTGCCCTGAG  GAGCCCCAGCTGAGGATGAAGAACAATGAGGAGGCTGAGGACTATGATGATGACCTGACTGACTCTGAGA  TGGATGTGGTGAGGTTTGATGATGACAACAGCCCCAGCTTCATCCAGATCAGGTCTGTGGCCAAGAAGCA  CCCCAAGACCTGGGTGCACTACATTGCTGCTGAGGAGGAGGACTGGGACTATGCCCCCCTGGTGCTGGCC  CCTGATGACAGGAGCTACAAGAGCCAGTACCTGAACAATGGCCCCCAGAGGATTGGCAGGAAGTACAAGA  AGGTCAGGTTCATGGCCTACACTGATGAAACCTTCAAGACCAGGGAGGCCATCCAGCATGAGTCTGGCAT  CCTGGGCCCCCTGCTGTATGGGGAGGTGGGGGACACCCTGCTGATCATCTTCAAGAACCAGGCCAGCAGG  CCCTACAACATCTACCCCCATGGCATCACTGATGTGAGGCCCCTGTACAGCAGGAGGCTGCCCAAGGGGG  TGAAGCACCTGAAGGACTTCCCCATCCTGCCTGGGGAGATCTTCAAGTACAAGTGGACTGTGACTGTGGA  GGATGGCCCCACCAAGTCTGACCCCAGGTGCCTGACCAGATACTACAGCAGCTTTGTGAACATGGAGAGG  GACCTGGCCTCTGGCCTGATTGGCCCCCTGCTGATCTGCTACAAGGAGTCTGTGGACCAGAGGGGCAACC  AGATCATGTCTGACAAGAGGAATGTGATCCTGTTCTCTGTGTTTGATGAGAACAGGAGCTGGTACCTGAC  TGAGAACATCCAGAGGTTCCTGCCCAACCCTGCTGGGGTGCAGCTGGAGGACCCTGAGTTCCAGGCCAGC  AACATCATGCACAGCATCAATGGCTATGTGTTTGACAGCCTGCAGCTGTCTGTGTGCCTGCATGAGGTGG  CCTACTGGTACATCCTGAGCATTGGGGCCCAGACTGACTTCCTGTCTGTGTTCTTCTCTGGCTACACCTT  CAAGCACAAGATGGTGTATGAGGACACCCTGACCCTGTTCCCCTTCTCTGGGGAGACTGTGTTCATGAGC  ATGGAGAACCCTGGCCTGTGGATTCTGGGCTGCCACAACTCTGACTTCAGGAACAGGGGCATGACTGCCC  TGCTGAAAGTCTCCAGCTGTGACAAGAACACTGGGGACTACTATGAGGACAGCTATGAGGACATCTCTGC  CTACCTGCTGAGCAAGAACAATGCCATTGAGCCCAGGAGCTTCAGCCAGAATGCCACTAATGTGTCTAAC  AACAGCAACACCAGCAATGACAGCAATGTGTCTCCCCCAGTGCTGAAGAGGCACCAGAGGGAGATCACCA  GGACCACCCTGCAGTCTGACCAGGAGGAGATTGACTATGATGACACCATCTCTGTGGAGATGAAGAAGGA  GGACTTTGACATCTACGACGAGGACGAGAACCAGAGCCCCAGGAGCTTCCAGAAGAAGACCAGGCACTAC  TTCATTGCTGCTGTGGAGAGGCTGTGGGACTATGGCATGAGCAGCAGCCCCCATGTGCTGAGGAACAGGG  CCCAGTCTGGCTCTGTGCCCCAGTTCAAGAAGGTGGTGTTCCAGGAGTTCACTGATGGCAGCTTCACCCA  GCCCCTGTACAGAGGGGAGCTGAATGAGCACCTGGGCCTGCTGGGCCCCTACATCAGGGCTGAGGTGGAG  GACAACATCATGGTGACCTTCAGGAACCAGGCCAGCAGGCCCTACAGCTTCTACAGCAGCCTGATCAGCT  ATGAGGAGGACCAGAGGCAGGGGGCTGAGCCCAGGAAGAACTTTGTGAAGCCCAATGAAACCAAGACCTA  CTTCTGGAAGGTGCAGCACCACATGGCCCCCACCAAGGATGAGTTTGACTGCAAGGCCTGGGCCTACTTC  TCTGATGTGGACCTGGAGAAGGATGTGCACTCTGGCCTGATTGGCCCCCTGCTGGTGTGCCACACCAACA  CCCTGAACCCTGCCCATGGCAGGCAGGTGACTGTGCAGGAGTTTGCCCTGTTCTTCACCATCTTTGATGA  AACCAAGAGCTGGTACTTCACTGAGAACATGGAGAGGAACTGCAGGGCCCCCTGCAACATCCAGATGGAG  GACCCCACCTTCAAGGAGAACTACAGGTTCCATGCCATCAATGGCTACATCATGGACACCCTGCCTGGCC  TGGTGATGGCCCAGGACCAGAGGATCAGGTGGTACCTGCTGAGCATGGGCAGCAATGAGAACATCCACAG  CATCCACTTCTCTGGCCATGTGTTCACTGTGAGGAAGAAGGAGGAGTACAAGATGGCCCTGTACAACCTG  TACCCTGGGGTGTTTGAGACTGTGGAGATGCTGCCCAGCAAGGCTGGCATCTGGAGGGTGGAGTGCCTGA  TTGGGGAGCACCTGCATGCTGGCATGAGCACCCTGTTCCTGGTGTACAGCAACAAGTGCCAGACCCCCCT  GGGCATGGCCTCTGGCCACATCAGGGACTTCCAGATCACTGCCTCTGGCCAGTATGGCCAGTGGGCCCCC  AAGCTGGCCAGGCTGCACTACTCTGGCAGCATCAATGCCTGGAGCACCAAGGAGCCCTTCAGCTGGATCA  AGGTGGACCTGCTGGCCCCCATGATCATCCATGGCATCAAGACCCAGGGGGCCAGGCAGAAGTTCAGCAG  CCTGTACATCAGCCAGTTCATCATCATGTACAGCCTGGATGGCAAGAAGTGGCAGACCTACAGGGGCAAC  AGCACTGGCACCCTGATGGTGTTCTTTGGCAATGTGGACAGCTCTGGCATCAAGCACAACATCTTCAACC  CCCCCATCATTGCCAGATACATCAGGCTGCACCCCACCCACTACAGCATCAGGAGCACCCTGAGGATGGA  GCTGATGGGCTGTGACCTGAACAGCTGCAGCATGCCCCTGGGCATGGAGAGCAAGGCCATCTCTGATGCC  CAGATCACTGCCAGCAGCTACTTCACCAACATGTTTGCCACCTGGAGCCCCAGCAAGGCCAGGCTGCACC  TGCAGGGCAGGAGCAATGCCTGGAGGCCCCAGGTCAACAACCCCAAGGAGTGGCTGCAGGTGGACTTCCA  GAAGACCATGAAGGTGACTGGGGTGACCACCCAGGGGGTGAAGAGCCTGCTGACCAGCATGTATGTGAAG  GAGTTCCTGATCAGCAGCAGCCAGGATGGCCACCAGTGGACCCTGTTCTTCCAGAATGGCAAGGTGAAGG  TGTTCCAGGGCAACCAGGACAGCTTCACCCCTGTGGTGAACAGCCTGGACCCCCCCCTGCTGACCAGATA  CCTGAGGATTCACCCCCAGAGCTGGGTGCACCAGATTGCCCTGAGGATGGAGGTGCTGGGCTGTGAGGCC  CAGGACCTGTACTGA   SEQ ID NO: 74 Exemplified FVIII polypeptide (N6)  MQIELSTCFFLCLLRFCFSATRRYYLGAVELSWDYMQSDLGELPVDARFPPRVPKSFPFNTSVVYKKT LFVEFTDHLFNIAKPRPPWMGLLGPTIQAEVYDTVVITLKNMASHPVSLHAVGVSYWKASEGAEYDDQ TSQREKEDDKVFPGGSHTYVWQVLKENGPMASDPLCLTYSYLSHVDLVKDLNSGLIGALLVCREGSLA KEKTQTLHKFILLFAVFDEGKSWHSETKNSLMQDRDAASARAWPKMHTVNGYVNRSLPGLIGCHRKSV YWHVIGMGTTPEVHSIFLEGHTFLVRNHRQASLEISPITFLTAQTLLMDLGQFLLFCHISSHQHDGME AYVKVDSCPEEPQLRMKNNEEAEDYDDDLTDSEMDVVRFDDDNSPSFIQIRSVAKKHPKTWVHYIAAE EEDWDYAPLVLAPDDRSYKSQYLNNGPQRIGRKYKKVRFMAYTDETFKTREAIQHESGILGPLLYGEV GDTLLIIFKNQASRPYNIYPHGITDVRPLYSRRLPKGVKHLKDFPILPGEIFKYKWTVTVEDGPTKSD PRCLTRYYSSFVNMERDLASGLIGPLLICYKESVDQRGNQIMSDKRNVILFSVFDENRSWYLTENIQR FLPNPAGVQLEDPEFQASNIMHSINGYVFDSLQLSVCLHEVAYWYILSIGAQTDFLSVFFSGYTFKHK MVYEDTLTLFPFSGETVFMSMENPGLWILGCHNSDFRNRGMTALLKVSSCDKNTGDYYEDSYEDISAY LLSKNNAIEPRSFSQNSRHPSTRQKQFNATTIPENDIEKTDPWFAHRTPMPKIQNVSSSDLLMLLRQS PTPHGLSLSDLQEAKYETFSDDPSPGAIDSNNSLSEMTHFRPQLHHSGDMVFTPESGLQLRLNEKLGT TAATELKKLDFKVSSTSNNLISTIPSDNLAAGTDNTSSLGPPSMPVHYDSQLDTTLFGKKSSPLTESG GPLSLSEENNDSKLLESGLMNSQESSWGKNVSSREITRTTLQSDQEEIDYDDTISVEMKKEDFDIYDE DENQSPRSFQKKTRHYFIAAVERLWDYGMSSSPHVLRNRAQSGSVPQFKKVVFQEFTDGSFTQPLYRG ELNEHLGLLGPYIRAEVEDNIMVTFRNQASRPYSFYSSLISYEEDQRQGAEPRKNFVKPNETKTYFWK VQHHMAPTKDEFDCKAWAYFSDVDLEKDVHSGLIGPLLVCHTNTLNPAHGRQVTVQEFALFFTIFDET KSWYFTENMERNCRAPCNIQMEDPTFKENYRFHAINGYIMDTLPGLVMAQDQRIRWYLLSMGSNENIH SIHFSGHVFTVRKKEEYKMALYNLYPGVFETVEMLPSKAGIWRVECLIGEHLHAGMSTLFLVYSNKCQ TPLGMASGHIRDFQITASGQYGQWAPKLARLHYSGSINAWSTKEPFSWIKVDLLAPMIIHGIKTQGAR QKFSSLYISQFIIMYSLDGKKWQTYRGNSTGTLMVFFGNVDSSGIKHNIFNPPIIARYIRLHPTHYSI RSTLRMELMGCDLNSCSMPLGMESKAISDAQITASSYFTNMFATWSPSKARLHLQGRSNAWRPQVNNP KEWLQVDFQKTMKVTGVTTQGVKSLLTSMYVKEFLISSSQDGHQWTLFFQNGKVKVFQGNQDSFTPVV NSLDPPLLTRYLRIHPQSWVHQIALRMEVLGCEAQDLY   SEQ ID NO: 75 Exemplified FVIII polypeptide (V3)      MQIELSTCFFLCLLRFCFSATRRYYLGAVELSWDYMQSDLGELPVDARFPPRVPKSFPFNTSVVYKKTLF VEFTDHLFNIAKPRPPWMGLLGPTIQAEVYDTVVITLKNMASHPVSLHAVGVSYWKASEGAEYDDQTSQR EKEDDKVFPGGSHTYVWQVLKENGPMASDPLCLTYSYLSHVDLVKDLNSGLIGALLVCREGSLAKEKTQT LHKFILLFAVFDEGKSWHSETKNSLMQDRDAASARAWPKMHTVNGYVNRSLPGLIGCHRKSVYWHVIGMG TTPEVHSIFLEGHTFLVRNHRQASLEISPITFLTAQTLLMDLGQFLLFCHISSHQHDGMEAYVKVDSCPE EPQLRMKNNEEAEDYDDDLTDSEMDVVRFDDDNSPSFIQIRSVAKKHPKTWVHYIAAEEEDWDYAPLVLA PDDRSYKSQYLNNGPQRIGRKYKKVRFMAYTDETFKTREAIQHESGILGPLLYGEVGDTLLIIFKNQASR PYNIYPHGITDVRPLYSRRLPKGVKHLKDFPILPGEIFKYKWTVTVEDGPTKSDPRCLTRYYSSFVNMER DLASGLIGPLLICYKESVDQRGNQIMSDKRNVILFSVFDENRSWYLTENIQRFLPNPAGVQLEDPEFQAS NIMHSINGYVFDSLQLSVCLHEVAYWYILSIGAQTDFLSVFFSGYTFKHKMVYEDTLTLFPFSGETVFMS MENPGLWILGCHNSDFRNRGMTALLKVSSCDKNTGDYYEDSYEDISAYLLSKNNAIEPRSFSQNATNVSN NSNTSNDSNVSPPVLKRHQREITRTTLQSDQEEIDYDDTISVEMKKEDFDIYDEDENQSPRSFQKKTRHY FIAAVERLWDYGMSSSPHVLRNRAQSGSVPQFKKVVFQEFTDGSFTQPLYRGELNEHLGLLGPYIRAEVE DNIMVTFRNQASRPYSFYSSLISYEEDQRQGAEPRKNFVKPNETKTYFWKVQHHMAPTKDEFDCKAWAYF SDVDLEKDVHSGLIGPLLVCHTNTLNPAHGRQVTVQEFALFFTIFDETKSWYFTENMERNCRAPCNIQME DPTFKENYRFHAINGYIMDTLPGLVMAQDQRIRWYLLSMGSNENIHSIHFSGHVFTVRKKEEYKMALYNL YPGVFETVEMLPSKAGIWRVECLIGEHLHAGMSTLFLVYSNKCQTPLGMASGHIRDFQITASGQYGQWAP KLARLHYSGSINAWSTKEPFSWIKVDLLAPMIIHGIKTQGARQKFSSLYISQFIIMYSLDGKKWQTYRGN STGTLMVFFGNVDSSGIKHNIFNPPIIARYIRLHPTHYSIRSTLRMELMGCDLNSCSMPLGMESKAISDA QITASSYFTNMFATWSPSKARLHLQGRSNAWRPQVNNPKEWLQVDFQKTMKVTGVTTQGVKSLLTSMYVK EFLISSSQDGHQWTLFFQNGKVKVFQGNQDSFTPVVNSLDPPLLTRYLRIHPQSWVHQIALRMEVLGCEA QDLY   EXAMPLES    The invention is now described with reference to the Examples below. These are not limiting  on  the  scope of  the  invention,  and  a person  skilled  in  the  art would be  appreciate  that  suitable  equivalents  could be used within  the  scope of  the present  invention. Thus,  the Examples may be  considered component parts of the  invention, and the  individual aspects described therein may be  considered as disclosed independently, or in any combination.    Example 1 – Design and in vitro validation of a cell‐specific core SFPB (mSPB) promoter sequence  The SFTPB promoter fragment of SEQ ID NO: 1 was designed following careful analysis of the   SFTPB genomic sequence. The expression driven by the core SFTPB promoter was compared to the  full length SFTPB promoter (the 972bp fragment of SEQ ID NO: 2, fSPB) using a human surfactant air‐ liquid interface (SALI) culture model; an in vitro cell culture model that robustly recapitulates human  ATII  cells  in  primary  cell  culture  (Munis  et  al.  (2021)  Molecular  Therapy:  Methods  &  Clinical  Development 20: 237‐246). H441 cells, when grown under SALI culture conditions, successfully mimic  key characteristics of primary ATII cells.  Briefly,  SALI  cultures  were  established  by  culturing  cells  in  12‐well  Transwell  inserts.  Approximately 1 × 105 cells/well in base media (5 × 104 each of H441 and A549 cells in the case of co‐ culture) were seeded  into Transwells and allowed to attach and proliferate for 48 h. On day 3, the  medium on the apical side of the Transwell chamber was removed to air‐lift the cells, and the medium  on  the basolateral side of  the chamber was  replaced with either “base” medium or “polarization”  medium. The polarization medium comprised either RPMI‐1640 (H441s cells and co‐culture) or F12‐K  (A549s) supplemented with 2 mM l‐glutamine, 50 U/mL penicillin, 50 mg/mL streptomycin, 1% insulin‐ transferrin‐selenium (GIBCO), 4% FCS, and 1 μM dexamethasone (Sigma). Media were subsequently  changed three times per week throughout the described experiments. Cells were grown under SALI  conditions for 14 days after air‐lift prior to experimentation unless otherwise stated.  SALI cultures were transduced with titre‐matched (1x106 TU) lentiviral vector (LV) expressing  EGFP from range of promoters (CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), EF1aS (SEQ ID NO: 7), PGK  (SEQ ID NO: 8), the core SFTPB fragment (mSPB) (SEQ ID NO: 1), full‐length SPB (fSPB) (SEQ ID NO: 4),  minimal SPC (mSPC) (SEQ ID NO: 9) and full‐length SPC (fSPC) (SEQ ID NO: 10)).   After 48 hours cells were subject to FACS analysis to quantify the % of cells expressing EGFP,  as  indicated  in the respective FACS plot shown Figure 1. The negative control sample  (Naïve; non‐ transduced cells H441 cells) shows a background  level of 0.4%. The % EGFP‐positive cells observed  with the widely‐used, non‐specific CMV [SEQ ID NO: 6], EF1aS [SEQ ID NO; 7], and PGK [SEQ ID NO: 8]  promoters is relatively high (17‐22%) in line with expectations. The % EGFP‐positive cells observed for  the hCEF promoter [SEQ ID NO: 5], which has been used previously in the lungs of patients, is 11%.  Promoter sequences taken from  lung surfactant genes (mSPB [SEQ  ID NO: 1], fSPB [SEQ  ID NO: 4],  mSPC [SEQ ID NO: 9], fSPC [SEQ ID NO: 10]) led to relatively lower % EGFP‐positive cells (0.7% to 9%)  suggesting that expression from these sequences have potential utility in a rationally designed new  promoter specific for ATII cells.  In particular, as shown in Figure 1, the mSPB promoter achieves expression levels (9%) similar  to hCEF (11%) in hSALI cultures. Notably, the mSPB promoter also achieves higher expression levels  compared to the full‐length SPB (fSPB) promoter sequence.   By way of comparison, human HEK293T cells were transduced with recombinant SIV lentiviral  vectors pseudotyped with VSV‐G and expressing  the EGFP  transgene  from  these  same promoters  (CMV (SEQ ID NO: 6), hCEF (SEQ ID NO: 5), EF1aS (SEQ ID NO: 7), PGK (SEQ ID NO: 8), the core SFTPB  fragment (mSPB) (SEQ ID NO: 1), full‐length SPB (fSPB) (SEQ ID NO: 4), minimal SPC (mSPC) (SEQ ID  NO: 9) and full‐length SPC (fSPC) (SEQ ID NO: 10)) at a multiplicity of infection (moi) of 0.5 or 1.0. After  48 hours, cells were subject to FACS analysis to quantify the % of cells expressing EGFP.  As shown in  Figure 2,  the % EGFP‐positive  cells  is  low  for mSPB,  fSPB, mSPC and  fSPC promoter  sequences  in  generic HEK293T cells, indicating that the activity of these may be cell‐specific. In particular, the mSPB  promoter drives very low level expression in HEK293T cells (see, Figure 2). The  expression driven by  the mSPB promoter is likely to be specific to lung parenchymal cells, particularly ATII cells.     Example 2 – The core SFPB (mSPB) promoter sequence drives cell‐specific gene expression in vivo  The specificity of the mSPB promoter was further investigated using an in vivo mouse model.   As shown  in Figure 3A, on day 0 mice were dosed by nasal  instillation with SIV vector. 7 days after  dosing, lung tissue was harvested, fixed, frozen and cryosectioned. The sections were then analysed  for expression of EGFP. As shown in Figure 3B, on day 0 female BALB/c mice (4‐6 weeks; n=3 per group)  were dosed with 1x106 TU (in 100µL TSSM buffer) of rSIV.F/HN.mSPB, an SIV vector comprising EGFP  under the control of an mSPB promoter (SEQ ID NO: 1). As controls, mice were alternatively dosed  with 100 µl of vehicle only (TSSM) or 1x106 TU  (in 100µL TSSM buffer) of  rSIV.F/HN.fSPB, an SIV vector  comprising EGFP under the control of an fSPB promoter (SEQ ID NO: 4).  Representative images are shown in Figure 4. As expected, no EGFP expression was seen in  the tissue obtained from the native mice. Foci corresponding to EGFP expression were seen  in the  sections obtained from mice injected with either rSIV.F/HN.mSPB or rSIV.F/HN.fSPB. However, there  are very few EGFP positive cells in the airways, indicating that the SFTPB promoters drive specific gene  expression in cells of the lung parenchyma, rather than airway cells.   Signficantly, approximately 5‐fold more punctate  fluorescent  signal was observed  in  lung  sections  from the mSPB group compared with the fSPB group of mice. These data suggest that the core SFTPB  promoter fragment (mSPB) is capable of achieving significantly greater transgene expression than the  full‐length SFTPB promoter of fSPB.    Example 3 – Design and production of improved mSPB promoters  The inventors then sought to further increase transgene expression by and/or activity of (i.e.   the number of  lung parenchyma cells  in which  the promoter  is active)  the mSPB promoter, whilst  retaining specificity for the lung parenchyma.  Specifically, the inventors generated a panel of improved mSPB promoters, each comprising  mSPB and an enhancer. Different enhancer  sequences were  selected,  including ubiquitously used  strong enhancers (i.e., CMV, hB‐actin or SV40), as well as tissue specific enhancers.    Lung‐specific  enhancers were  selected  using  the ATII  cell  gene  expression  database  from  LungGENS and a tissue expression database. To generate candidate sequences to boost the level of  transgene expression from the mSPB (SEQ ID NO: 1) promoter sequence in ATII cells, tissue expression  databases  including  https://research.cchmc.org/pbge/lunggens/default.html  and  https://tissues.jensenlab.org/Search were  interrogated  to  identify  genes with high mean  levels of  expression in ATII cells.   The resulting panel of ATII specific genes (and enhancers) is shown in Figure 5. Of these, SFTPC,  SFTPB, SLC34A2, GOLGA8B, LMO7, VEGFA and ELF3 were found to be most highly expressed in ATII  cells. SFTPC was found to be highly specific for the lungs. Although SLC34A2 was found to be highly  expressed  in the  lungs, high  levels of expression was also found  in the gut.   Further, GOLGA8B was  found to be most highly expressed in the brain.   Enhancer  regions  within  the  above‐mentioned  genes  were  identified  using  Genecards:  https://www.genecards.org/Guide/GeneCard.  Specific  sequences  in  these  enhancer  regions were  selected (and transcription factor binding sites identified) using University of California at Santa Cruz  (UCSC) genome browser https://genome.ucsc.edu.   The  identified enhancer  sequences were cloned  into a construct comprising mSPB and an  operably  linked  EGFP2ALux  transgene.  To  determine whether  expression  from  the  enhancers  is  affected by orientation, enhancers were added to the constructs  in the forward (F) and reverse (R)  direction.  Schematics  of  exemplary  constructs  are  shown  in  Figure  6,  with  the mSPB  promoter  sequence of SEQ ID NO: 1 and enhancer sequences as per SEQ ID NOs: 13 to 20, 23 to 33 and 36 to 40.      Similar reporter transgene expression constructs were also generated  incorporating the commonly  used enhancers hB‐Actin (SEQ ID NOs: 21 and 22), SV40 (SEQ ID NOs: 34 and 35) and CMV (SEQ ID  NOs:  11  and  12).  Using  transient  transfection  with  standard  lentiviral  producer  plasmids,  the  expression constructs were  incorporated  into  recombinant  lentiviral vectors. The various  lentiviral  vectors  were  prepared  for  analysis  of  EGFP  or  Lux  reporter  transgene  expression  from  each  enhancer/promoter combination.  Human HEK293T cells, murine LA‐4 cells and human SALI cells were transfected with the mSPB  promoter constructs to determine whether the addition or the enhancer (or the orientation) increased  gene expression.   Recombinant  SIV  lentiviral  vectors  pseudotyped  with  VSV‐G  and  expressing  EGFP2ALux  transgene from each of the enhancer/promoter combinations, were produced at small scale. These  were used  to  transduce  human HEK293T  cells  in  a  24‐well plate  (seeded @1x105  cells/well)  at  a  multiplicity of infection (moi) of approximately 0.1. Luciferase levels (Relative Light Units (RLU)) were  determined  in n=4  independent biological  replicates. Naïve  (non‐transduced) cells were used as a  control to determine background RLU level. The results for human HEK293T cells are shown in Figure  7. In HEK293T cells, mSPB did not significantly increase gene expression relative to the naïve control.  Differences in expression driven by mSPB compared with mSPB + the different enhancers was also not  significant.   In contrast, significant differences in gene expression are seen in murine LA‐4 cells (see, Figure  8) and human SALI  cells  (see, Figure 9).  In particular,  in murine LA‐4 cells,  the addition of a CMV  enhancer (forwards or reverse), ELF3 enhancer (forwards), SV40 enhancer (forwards), and two of the  VEGFA  enhancers  –  VEG1  (forwards  or  reverse)  and  VEG2  (forwards)  to  the  mSPB  promoter  significantly increased expression relative to the mSPB promoter alone. Similarly, in human SALI cells,  the addition of a CMV enhancer  (forwards or  reverse), Elf3  (forwards or  reverse) enhancer,  SLC3  (reverse) enhancer,  SV40 (forwards or reverse) enhancer, and two of the VEGFA enhancers – VEG1  (forwards or reverse) and VEG3  (reverse)  to  the mSPB promoter significantly  increased expression  relative to the mSPB promoter alone.  These data demonstrate that it is possible to further increase transgene expression in a cell‐ specific manner, even beyond the  level obtained using the mSPB promoter (which  itself provides a  significant  increase  in  expression  compared  with  the  fSPB  promoter),  using  specific  enhancer  sequences.    Example 4 – Further characterisation of improved mSPB promoters  Using the results of Example 3, the following seven candidate enhancers were selected for  further screening: SV40, CMV, SLC34A2‐1, SLC34A2‐2, VEGFA‐2, ELF3‐1 and ELF3‐2.   Recombinant  SIV  lentiviral  vectors  pseudotyped  with  F/HN  and  expressing  EGFP2ALux  reporter  transgene,  were  used  to  transduce  human  SALI  cultures  (n=6  replicates)  in  a  repeat  secondary screening experiment, which included the following promoter sequences: hCEF (SEQ ID NO:  5),  CMV  (SEQ  ID NO:  6),  and mSPB  (SEQ  ID NO:  1),  as well  as  the  selected  enhancer/promoter  combinations SV40 (f) (SEQ ID NO: 34), CMVenh (f (SEQ ID NO: 11)), Elf1 (f) (SEQ ID NO: 13), Elf2 (f)  (SEQ ID NO: 15), Slc1 (f) (SEQ  ID NO: 29), Slc2 (f) (SEQ ID NO: 31), and Veg2 (f) (SEQ ID NO: 38))  in  combination with mSPB (SEQ ID NO: 1).  In a secondary screen, expression driven by the mSPB promoter combined with the CMV, SV40  and VEGFA‐2 enhancer in SALI cells was found to be comparable to the level of expression driven by  the strong promoter hCEF (Figure 10).    Recombinant SIV  lentiviral vectors pseudotyped with  the F/HN and expressing EGFP2ALux  reporter transgene, were then used to transduce human HEK293T cell cultures (n=6 replicates) in a  repeat secondary screening experiment, which  included the following selected enhancer/promoter  sequences: hCEF (SEQ ID NO: 5), CMV (SEQ ID NO: 6), and mSPB (SEQ ID NO: 1), as well as the selected  enhancer/promoter combinations SV40 (f) (SEQ ID NO: 34), CMVenh (f (SEQ ID NO: 11)), Elf1 (f) (SEQ  ID NO: 13), Elf2 (f) (SEQ ID NO: 15), Slc1 (f) (SEQ ID NO: 29), Slc2 (f) (SEQ ID NO: 31), and Veg2 (f) (SEQ  ID  NO:  38))  in  combination with mSPB  (SEQ  ID  NO:  1).  These  data  demonstrate  that  transgene  expression using the hCEF promoter was not cell‐specific, with significantly greater expression with  hCEF seen in HEK293T cells compared with expression using the mSPB/enhancers (Figure 11).   The inventors next assessed whether the mSPB/enhancer improved promoters were able to  drive expression of luciferase in vivo. Lentiviral vectors comprising different promoter constructs were  administered  to mice  via  the  nose  and  the  lungs.  The mice were  then monitored  for  luciferase  expression. Female BALB/c mice (6 weeks; n=5 per group) were dosed with recombinant SIV lentiviral  vectors pseudotyped with F/HN and expressing EGFP2ALux  reporter  transgene  from  the  following  promoter sequences: hCEF (SEQ ID NO: 5), CMV (SEQ ID NO: 6), and mSPB (SEQ ID NO: 1), as well as  the selected enhancer/promoter combinations SV40 (f) (SEQ ID NO: 34), CMVenh (f) (SEQ ID NO: 11),  Elf1 (f) (SEQ ID NO: 13), Elf2 (f) (SEQ ID NO: 15), Slc1 (f) (SEQ ID NO: 29), Slc2 (f) (SEQ ID NO: 31), and  Veg2 (f) (SEQ ID NO: 38)) in combination with mSPB (SEQ ID NO: 1). Lentivirus was dosed intranasally  (1x107  TU  per mouse  in  100uL  TSSM  Buffer).  On  days  2,  7  and  28  post‐dosing,  all  mice  were  anaesthetised and imaged for luciferase expression.   Figure 12 shows the expression of luciferase as a heat map. Expression driven by hCEF or CMV  promoters  is shown as a reference. The  inventors selected promoter which drive expression  in the  lungs, but not the nose. Localised expression in the lungs is important, because the nose has cells that  are  similar  to  (airway epithelial  cells  ciliated/non‐ciliated epithelial)  to  those  in  the  lungs,  so  it  is  important to determine that the transgene will not be expressed in the nose.   The expression of luciferase is quantified in Figure 13A.  The ratio of expression in the lungs  and the nose is shown in Figure 13B. The level of luciferase signal (Figure 13A) was greatest with the  CMVenh group (SEQ ID NO: 11) (**** compared to mSPB (SEQ ID NO: 1) although the specificity for  lung expression (Figure 13B), determined by the ratio of signal in the lung and nose (L‐to‐N), was not  significantly different from CMV (SEQ ID NO: 6)or hCEF (SEQ ID NO: 5) at day 28.  The luciferase signal  (Figure 13A) in the slc2 (SEQ ID NO: 31) and Veg2 (SEQ ID NO: 38) groups was significantly increased  compared with mSPB (SEQ ID NO: 1) (*and ** respectively, compared to hCEF group; # NS compared  to mSPB; Kruskal Wallis) and also showed good overall specificity of lung expression (Figure 13A). Thus,  the results indicate that the mSPB promoter combined with an SLC2, VEGF2 or CMV enhancer drives  high levels of expression in the lungs, but not the nose.   Human  SALI  cultures  (n=8) were  transfected with plasmid DNA  (2µg plasmid per  culture)  complexed with linear polyethyleneimine (PEIPro; 100µL per culture) and luciferase activity in Relative  Light Units (RLU) determined. A naïve group (n=8; non‐transfected) was included as a negative control.  Similar results to those shown in Figure 13 using a SIV vector were observed when the promoters were  used  to  drive  expression  of  luciferase  in  a  non‐viral  (plasmid)  vector,  with  the  mSPB  promoter/enhancer combinations driving  increased  transgene expression  in SALI cells  (Figure 14A)  and  in  vivo  (Figure  14B)  compared with  expression  in HEK293T  cells  (Figure  14C).    In  particular,  luciferase activity in the veg2 (SEQ ID NO: 38) and CMVenh (SEQ ID NO: 11) groups were significantly  different from mSPB (SEQ ID NO: 1).  Overall, the levels of transgene expression (luciferase) from these  constructs delivered as a non‐viral  (plasmid)  formulation were similar  to  those obtained  following  delivery with viral (recombinant lentiviral) vectors.    Example 5 – Improved mSPB promoters drive targeted expression in cells of the lung parenchyma  Further analysis was conducted to confirm that the mSPB promoter/enhancer combinations  result in transgene expression in the target cells, particularly that expression is focussed in the cells of  the lung parenchyma, rather than airway cells.  This was investigated using immunohistochemistry.   As described  in Example 4, Figure 12,  female BALB/c mice  (6 weeks; n=5 per group) were  dosed with  recombinant  SIV  lentiviral  vectors pseudotyped with  F/HN  and  expressing  EGFP2ALux  reporter transgene from different enhancer/promoter sequences. Lentivirus was dosed intranasally  (1x107 TU per mouse in 100µL TSSM Buffer) and culled on day 28 post‐dosing. The enhancer sequences  CMVenh (SEQ ID NO: 11), Slc2 (SEQ ID NO: 31) and VEGFA‐2 (SEQ ID NO: 38) combined with the mSPB  (SEQ  ID  NO:  1)  promoter  sequence  were  assigned  the  nomenclature  Alv‐01,  Alv‐02  and  Alv‐03  respectively, as shown  in SEQ ID NOs: 46, 47 and 48 respectively. Mouse  lungs were processed for  cryosections and imaging. Cryosections (7µM) were subject to immunohistochemistry using primary  antibodies to detect colocalization of EGFP and Pro/Mature Surfactant Protein‐B, the  latter being a  marker  for ATII  cells. Cryosections were also  stained with DAPI  for ease of visualisation and  then  imaged (at least n=3 sections per group) using a confocal microscope.  As shown in Figure 15, Alv‐01,  Alv‐02 and Alv‐03 were able  to drive EGFP expression which  colocalises with  the ATII  cell‐specific  marker  SP‐B.    These  results  demonstrate  that  the  newly  identified  enhancer/mSPB  promoter  constructs can facilitate transgene expression in ATII cells in the lung parenchyma. In particular, this  experiment  confirmed  that  the  mSPB  promoter/enhancer  combinations  tested  drove  transgene  expression primarily in cells of the parenchyma, making them attractive candidates for targeted gene  therapy of diseases resulting from or associated with deficiency of expression and/or expression of  defective proteins in lung parenchymal cells, such as surfactant deficiencies.  Repetition of this experiment using a different ATII cell marker, Surfactant Protein‐C yielded  the  same  pattern  of  expression,  with  Alv‐01,  Alv‐02  and  Alv‐03  driving  EGFP  expression  which  colocalises with the ATII cell‐specific marker SP‐C (data not shown).  Therefore, Alv‐01, Alv‐02 and Alv‐ 03 have been reproducibly shown to drive expression in the lung parenchyma.    Example 6 – Improved mSPB promoters drive long‐term expression in vivo  The in vivo expression of luciferase driven by the mSPB/enhancer  improved promoters Alv‐ 01, Alv‐02 and Alv‐03 as described in Example 5 above was further characterised.   Lentiviral  vectors  comprising  mSPB,  Alv‐01,  Alv‐02  and  Alv‐03  were  administered  to  intranasally. The mice were then monitored for luciferase expression. Female BALB/c mice (n=10 per  group) were dosed with recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing a  firefly  luciferase  reporter  transgene  from mSPB,  Alv‐01,  Alv‐02  and  Alv‐03.  SIV  lentiviral  vectors  pseudotyped with F/HN and expressing a firefly  luciferase reporter transgene from a CMV or hCEF  promoter were used as a control.  Lentivirus was dosed intranasally (2.5x108 TU per mouse in TSSM Buffer). On days 7, 14, 28,  91 and 186 post‐dosing, mice were anaesthetised and imaged for luciferase expression.   Figure 16 shows the expression of luciferase as a heat map. Expression driven by hCEF or CMV  promoters is shown as a reference.   The expression of luciferase in the lungs over the time course of the experiment is quantified  in Figure 16B.  The area under the curve is shown in Figure 16C and the ratio of expression in the lungs  and the nose is shown in Figure 16D. All of hCEF, mSPB, Alv‐01, Alv‐02 and Alv‐03 gave significantly  higher expression of the luciferase reporter compared with the CMV promoter (Figure 16A, B and C)  The level of luciferase signal (Figure 16A, B and C) was greatest with Alv‐01.  The CMV promoter and  hCEF promoter also drove significant expression in the nose compared with the mSPB, Alv‐02 and Alv‐ 03 promoters (Figure 16A).   Specificity for  lung expression (Figure 16D), determined by the ratio of  signal  in  the  lung  and  nose  (L‐to‐N), was  significantly  increased  by  the mSPB, Alv‐02  and  Alv‐03  promoters  compared with  the  CMV  or  hCEF  promoters  at  day  28.    In  this  experiment,  the  lung  specificity of Alv‐01 was low (although further experiments below demonstrate specificity with Alv‐ 01).   The specificity of luciferase signal in the lungs driven by the Alv‐02 and Alv‐03 promoters was  observably increased compared with the mSPB promoter(Figure 16D).   Furthermore,  the  high  level  of  luciferase  expression  and  high  degree  of  specificity  was  maintained over  the 6‐month  time course of  the experiment  (Figure 16A and B). Thus,  the results  indicate that the Alv‐01, Alv‐02 and Alv‐03 promoter drives high levels of expression in the lungs over  at least a six‐month period.  Further, mSPB, Alv‐02 and Alv‐03 drive high levels of expression in the  lungs, but not the nose.     Example 7 – Improved mSPB promoters drive targeted expression in cells of the lung parenchyma  Further analysis was conducted to confirm that Alv‐01, Alv‐02 and Alv‐03 result in transgene  expression  in  the  target  cells,  particularly  that  expression  is  focussed  in  the  cells  of  the  lung  parenchyma, rather than airway cells.  This was investigated using immunohistochemistry.   As  described  in  Example  5,  female  BALB/c  mice  (n=5‐10  per  group)  were  dosed  with  recombinant SIV lentiviral vectors pseudotyped with F/HN and expressing EGFP from mSPB, Alv‐01,  Alv‐02 and Alv‐03. Lentivirus was dosed intranasally  and culled on day 14 post‐dosing. Cryosections  (7µM) were subject to immunohistochemistry using primary antibodies to detect EGFP in the airway  and parenchymal cells. Cryosections were also stained with DAPI for ease of visualisation and then  imaged (at least n=3 sections per group) using a confocal microscope.  As shown in Figure 17A, the  hCEF promoter mainly drove expression in the cells of the airway, whereas as shown in Figure 17B,  the mSPB promoter drove expression primarily in the parenchymal cells.  The % of EGFP expression in  the parenchymal cells for hCEF, mSPB, Alv‐01, Alv‐02 and Alv‐03 was quantified.  As shown in Figure  17C , Alv‐01, Alv‐02 and Alv‐03, particularly Alv‐02 and Alv‐03 were able to drive EGFP expression in  the parenchyma compared with the hCEF promoter.    These results demonstrate that Alv‐01, Alv‐02 and Alv‐03, particularly Alv‐02 and Alv‐03 can  facilitate transgene expression in the lung parenchyma. In particular, this experiment confirmed that  the mSPB  promoter,  Alv‐01,  Alv‐02  and  Alv‐03,  particularly  Alv‐02  and  Alv‐03,  drove  transgene  expression primarily in cells of the parenchyma, making them attractive candidates for targeted gene  therapy of diseases resulting from or associated with deficiency of expression and/or expression of  defective proteins in lung parenchymal cells, such as surfactant deficiencies.    Example 8 – Expression of SP‐B restores transepithelial electrical resistance (TEER) in SFTPB knock  out lung cells  To investigate whether the mSPB, Alv‐01, Alv‐02 and Alv‐03 promoters could perform well in  the context of a human lung cell model, a  Surfactant Air Liquid Interface (SALI) model was used with  the human H441 lung cell line and the H441 SP‐B KO cell line where the SFTPB gene had been ablated  (Munis et al. (2021) Mo. Ther. Methods Clin. Dev 20:20:237‐246, herein incorporated by reference).  The ablation of the gene encoding SP‐B leads to an observed phenotype – a reduction in transepithelial  electrical resistance (TEER) ‐ that can be corrected by expression of human SP‐B.   SALI  cultures  were  generated  following  airlift  and  transduced  with  rSIV.F/HN  vectors  expressing human SP‐B under the control of one of the following promoters: hCEF, mSP‐B, Alv‐1, Alv‐ 2, or Alv‐3. Transduction with a rSIV.F/HN vector expressing EGFP under the control of hCEF was used  as a negative control (mock). Before transduction, and up to 30 days thereafter, TEER was measured.   As previously published, (Figure 18 A and B) TEER values from SALI cultures generated from  the H441 SP‐B KO cells were lower than the parental H441 cell line at 14 days post airlift (i.e. 4 days  prior to transductions with viral vectors) constituting a phenotypic defect.   At 5 days post transduction, rSIV.F/HN expressing EGFP control vector did not  increase the  TEER in H441 SP‐B KO cells SALI cultures (Figure 18C), however transduction with any/all of the vectors  expressing SP‐B showed correction of the TEER towards normal, calculated as a percentage of  the  mock‐transduced parental H441 cells (Figure 18D).   These data show that expression of human SP‐B under control of mSP‐B, or Alv‐1‐3 promoters  can also phenotypically correct the TEER defect offering alternative sequences for gene expression in  ATII cells.   These data  show  that  switching  from  the  strong hCEF promoter  to  the  lung parenchymal  specific mSP‐B, Alv‐1, Alv‐2, or Alv‐3 promoters does not have a negative effect on TEER restoration.   Thus,  these  promoters  advantageously  allow  parenchymal‐specific  expression  of  SFTPB  whilst  achieving the same clinically desirable restoration of TEER.    Example 9 – Pulmonary surfactant has no effect on HEK293/T cell transduction with rSIV.F/HN  vectors  To investigate whether complexing rSIV.F/HN vectors with synthetic surfactant (such as BLES  or Curosurf) affects cell transduction, rSIV.F/HN vector encoding EGFP under CMV promoter control,  was  mixed  1:1  with  TSSM  (vehicle  control)  or  with  BLES  or  Curosurf  and  incubated  at  room  temperature for 30mins. The mixtures were then diluted in OptiMEM‐I (supplemented with polybrene  for final 8μg/mL working concentration) and used to transduce HEK293T cells.   At 72h post‐transduction, cells were analysed by flow cytometry, and transduction efficiencies  were corrected by  subtracting  the background  fluorescence  from cells  treated with  the Sham  (no  vector) control mixed with respective TSSM, BLES, or Curosurf. As shown in Figure 19, there was no  statistically significant difference in transduction efficiency between vectors mixed 1:1 with BLES or  Curosurf, compared with control vector mixed with TSSM buffer. These data show that pulmonary  surfactants  (BLES  and  Curosurf)  can  be  complexed with  rSIV.F/HN  vector without  compromising  transduction of human HEK293T cells. It was concluded that, despite the rSIV.F/HN lentiviral vectors  having a lipid envelope, this is not affected by the presence of the pulmonary surfactant, and thus the  pulmonary surfactants do not compromise the integrity of rSIV.F/HN lentiviral vectors.    Example 10 – Murine lung transduction with rSIV.F/HN vectors is at least as effective in the  presence of a pulmonary surfactant  Having  surprisingly  shown  that pulmonary  surfactants do not  compromise  the  integrity of  rSIV.F/HN  lentiviral vectors  in an  in vitro  setting,  the effect of pulmonary  surfactants on  rSIV.FHN  transduction of the murine lung was then investigated.    rSIV.F/HN vector expressing firefly luciferase from the hCEF promoter (rSIV.F/HN hCEF Flux)  was mixed 1:1 with TSSM (vehicle control) or BLES or Curosurf (to generate 2E8 TU in 100uL per dose).  BALB/c mice (n=5) were dosed by intranasal administration. TSSM diluent served as vehicle control.  Mice were subjected to in vivo bioluminescent imaging to measure Firefly luciferase expression in the  murine lungs. The time course of luciferase expression in the lungs of each mouse on days 7, 14, 28,  112 days post‐dosing is shown in Figure 20A, from which it can be seen that the luciferase expression  was stable across the duration of the experiment.   When  the  total  luciferase was quantified as area under  the  curve  (AUC)  in Figure 20B,   a  statistically significant  increase  in transformation was achieved  in the presence of either surfactant  compared with the vehicle control.  When this was compared with the TSSM vehicle control (Figure  Figure 20C) a ∼2‐3‐fold improved expression was observed when the vector was mixed with either   pulmonary surfactant BLES and Curosurf.       

Claims

CLAIMS    1. An SFTPB promoter fragment which comprises or consists of at  least part of exon 1 of the  SFTPB gene, but does not comprise the SFTPB gene start codon.    
2. The SFTPB promoter fragment of claim 1, which is less than 800 bases in length, preferably  less than 700 bases in length.    
3. The SFTPB promoter fragment of claim 1 or 2, which comprises or consists of:     (a) SEQ ID NO: 1  or a sequence with at least 80%  identity to SEQ ID NO: 1; or  (b) bases 81‐710 of SEQ ID NO: 2, or a sequence with at least 80% identity to bases 81‐710 of  SEQ ID NO: 2;     wherein optionally:    (i) said SFTPB promoter fragment further comprises up to 20 bases at the 5’ end, which may  optionally correspond to up to 20 bases 5’ to base 81 of SEQ ID NO: 2; and/or  (ii) said SFTPB promoter fragment further comprises up to 4 bases at the 3’ end, which may  optionally correspond to up to 4 bases 3’ to base 710 of SEQ ID NO: 2.   
4. The SFTPB promoter fragment of any one of claims 1‐3, which comprises or consists of SEQ ID  NO: 1 or a sequence with at least 90% identity to SEQ ID NO: 1.     
5. The SFTPB promoter fragment of any one of the preceding claims, which further comprises a  5’ enhancer.    
6. The SFTPB promoter fragment of claim 5, wherein the enhancer is in (i) the forward, or (ii) the  reverse, orientation; preferably wherein the enhancer is in the forward orientation.    
7. The  SFTPB  promoter  fragment  of  claim  5  or  6, wherein  the  enhancer  is  selected  from  a  SLC34A2 enhancer, a VEGFA enhancer, a CMV enhancer, an SV40 enhancer, an ELF3 enhancer,  an actin enhancer, an LMO7 enhancer, a SFTPC enhancer or a SFTPB enhancer,   wherein preferably the enhancer is selected from a SLC34A2 enhancer, a VEGFA enhancer or  a CMV enhancer.   
8. The SFTPB promoter fragment of claim 7, wherein:    (a) the SLC34A2 enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to any one of SEQ ID NOs: 29‐33, preferably SEQ ID NO: 31;  (b) the VEGFA enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to any one of SEQ ID NOs: 36‐40, preferably SEQ ID NO: 38;  (c) the CMV enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to SEQ ID NO: 11 or 12, preferably SEQ ID NO: 11;   (d) the SV40 enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to SEQ ID NO: 34 or 35;   (e) the ELF3 enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to any one of SEQ ID NOs: 13‐20;  (f) the actin enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to SEQ ID NO: 21 or 22;  (g) the LMO7 enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to any one of SEQ ID NOs: 23—25;  (h) the SFTPC enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to SEQ ID NO:  27 or 28; or  (i) the SFTPB enhancer comprises or consists of a nucleotide sequence with at least 90%  identity to SEQ ID NO: 26.   
9. The SFTPB promoter fragment of any one of the preceding claims, which comprises or consists  of a nucleic acid sequence of any one of SEQ  ID NOs: 41, 42, 43, 44, 45, 46, 47 or 49, or a  sequence with at least 90% identity to any one of SEQ ID NOs: 41, 42, 43, 44, 45, 46, 47 or 48,  preferably wherein said promoter fragment comprises or consists of a nucleic acid sequence  of any one of SEQ ID NOs: 46, 47 or 48, or  a sequence with at least 90% identity to any one of  SEQ ID NOs: 46, 47 or 48.   
10. A nucleic acid cassette comprising:     (a) an SFTPB promoter fragment as defined in any one of the preceding claims; and  (b) a transgene.   
11. The  nucleic  acid  cassette  of  claim  10, wherein  transgene  encodes  a  therapeutic  protein  selected from:    (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP‐B), Surfactant  Protein  C  (SP‐C),  AAT,  Factor  VIII,  Factor  VII,  Factor  IX,  Factor  X,  Factor  XI,  van  Willebrand  Factor,  Granulocyte‐Macrophage  Colony‐Stimulating  Factor  (GM‐CSF),  decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody, an  anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or  (b) ATP‐binding cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.   
12. The nucleic acid cassette of claim 10 or 11, wherein the SFTPB promoter fragment increases  expression of the transgene by lung parenchymal cells, wherein optionally:    (a) expression of the transgene by the SFTPB promoter fragment is increased compared  with expression of the transgene by the full‐length SFTPB promoter;  (b)  expression of the transgene by the SFTPB promoter fragment is increased by at least  2‐fold, preferably at  least 5‐fold compared expression of the  transgene by the  full‐ length SFTPB promoter.   
13. The nucleic acid cassette of any one of claims 10 to 12, wherein the lung parenchymal cells  comprise one or more cell type selected from: alveolar type  I epithelial (ATI) cells, alveolar  type II epithelial cells (ATII), and/or club cells, preferably ATII and/or ATI cells.   
14. The nucleic acid cassette of any one of claims 11 to 13, wherein:     (a) expression of the transgene is specific to lung parenchymal cells; and/or  (b) the ratio of lung expression: nose expression of the transgene by the SFTPB promoter  fragment is at least 2:1.    
15. A  gene  therapy  vector,  comprising  a  nucleic  acid  cassette  as  defined  in  any  one  of  the  preceding claims.   
16. The gene therapy vector of claim 15, which is a non‐viral vector, wherein optionally:    (a) the non‐viral vector is a plasmid; and/or  (b) the non‐viral vector is comprised in a cationic liposome, which preferably comprises  GL67A.   
17. The gene therapy vector of claim 15, which is a viral vector, optionally selected from:    (a) a lentiviral vector;  (b) an AAV vector; and  (c) an adenoviral vector.    
18. The gene  therapy vector of claim 17, which  is a  lentiviral vector  that  is pseudotyped with  haemagglutinin‐neuraminidase (HN) and fusion (F) proteins from a respiratory paramyxovirus,  wherein optionally the respiratory paramyxovirus is a Sendai virus.   
19. The gene therapy vector of claim 17 or 18, wherein the lentiviral vector is selected from the  group consisting of a Human immunodeficiency virus (HIV) vector, a Simian immunodeficiency  virus (SIV) vector, a Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia  virus (EIAV) vector, and a Visna/maedi virus vector, preferably a SIV vector.   
20. A method of expressing a therapeutic protein in a target cell, comprising delivering a nucleic  acid cassette as defined in any one of claims 1‐14 or a gene therapy vector as defined in any  one of claims 15‐19 into the target cells.   
21. The method  of  claim  20, wherein  said  delivering  comprises  integrating  said  nucleic  acid  cassette or gene therapy vector into said target cell's genome.    
22. A gene therapy vector as defined in any one of claims 15‐19 for use in a method of treating a  disease.   
23. The gene therapy vector for use of claim 22, wherein the disease is:    (a) a genetic disease;  (b) a respiratory disease, particularly a genetic respiratory disease;   (c) a  cardiovascular  disease  or  blood  disorder,  particularly  a  genetic  cardiovascular  disease or blood disorder; and/or  (d) selected  from  Surfactant  Protein  B  (SP‐B)  Deficiency;  Surfactant  Protein  C  (SP‐C)  deficiency;  ABCA3  deficiency;  Pulmonary  surfactant  metabolism  dysfunction  2  (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary  Dyskinesia  (PCD);  Alpha  1‐antitrypsin  Deficiency  (A1AD);  Pulmonary  Alveolar  Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory  distress  syndrome  (ARDS);  COVID‐19;  a  pulmonary  fibrotic  disease;  a  pulmonary  allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in  the lungs; and haemophilia.   
24. A cell comprising a nucleic acid cassette as defined  in any one of claims 1‐14 or a gene  therapy vector as defined in any one of claims 15‐19.   
25. A composition comprising a nucleic acid cassette as defined in any one of claims 1‐14 or a  gene  therapy  vector  as  defined  in  any  one  of  claims  15‐19  and  a  pharmaceutically  acceptable carrier, diluent or excipient.   
26. A lentiviral vector pseudotyped with haemagglutinin‐neuraminidase (HN) and fusion (F)  proteins from a respiratory paramyxovirus and comprising a transgene operably linked to  a promoter  for use  in a method of  treating a disease, wherein  the  lentiviral vector  is  administered simultaneously or sequentially with a surfactant.   
27. The lentiviral vector for use of claim 26, wherein the disease is:    (a) a genetic disease;  (b) a respiratory disease, particularly a genetic respiratory disease;   (c) a  cardiovascular  disease  or  blood  disorder,  particularly  a  genetic  cardiovascular  disease or blood disorder; and/or  (d) selected  from  Surfactant  Protein  B  (SP‐B)  Deficiency;  Surfactant  Protein  C  (SP‐C)  deficiency;  ABCA3  deficiency;  Pulmonary  surfactant  metabolism  dysfunction  2  (SMDP2); Pulmonary surfactant metabolism dysfunction 3 (SMDP3); Primary Ciliary  Dyskinesia  (PCD);  Alpha  1‐antitrypsin  Deficiency  (A1AD);  Pulmonary  Alveolar  Proteinosis (PAP); Chronic obstructive pulmonary disease (COPD); Acute respiratory  distress  syndrome  (ARDS);  COVID‐19;  a  pulmonary  fibrotic  disease;  a  pulmonary  allergic condition; a pulmonary bacterial infection; lung cancer; a dysplastic change in  the lungs; and haemophilia.   
28. The lentiviral vector for use of claim 26 or 27, wherein:     (a) the respiratory paramyxovirus is a Sendai virus; and/or  (b) the  lentiviral  vector  is  selected  from  the  group  consisting  of  a  Human  immunodeficiency virus (HIV) vector, a Simian immunodeficiency virus (SIV) vector, a  Feline immunodeficiency virus (FIV) vector, an Equine infectious anaemia virus (EIAV)  vector, and a Visna/maedi virus vector, preferably a SIV vector    29. The lentiviral vector for use of any one of claims 26 to 28, wherein the transgene encodes  a therapeutic protein selected from:    (a) a secreted therapeutic protein selected from: Surfactant Protein B (SP‐B), Surfactant  Protein  C  (SP‐C),  AAT,  Factor  VIII,  Factor  VII,  Factor  IX,  Factor  X,  Factor  XI,  van  Willebrand  Factor,  Granulocyte‐Macrophage  Colony‐Stimulating  Factor  (GM‐CSF),  decorin, an anti‐inflammatory protein (e.g. IL‐10 or TGFβ) or monoclonal antibody, an  anti‐inflammatory decoy, or a monoclonal antibody against an infectious agent; or  (b) ATP‐binding cassette sub‐family A member 3 (ABCA3), TRIM72, CSF2RA, or CSF2RB.    30. The lentiviral vector for use of any one of claims 26 to 29, wherein:    (a) the promoter comprises a SFTPB promoter fragment as defined in any one of claims 1  to 9; and/or   (b) the lentiviral vector comprises a nucleic acid cassette as defined in any one of claims  10 to 14.    31. The lentiviral vector for use of any one of claims 26 to 30, wherein:    (a) the lentiviral vector is administered before the surfactant;  (b) the surfactant is administered before the lentiviral vector; or  Ĩc) the  lentiviral  vector  and  surfactant  are  administered  simultaneously,  optionally  wherein the lentiviral vector and surfactant are mixed prior to administration.       
EP24712568.5A 2023-03-07 2024-03-07 Synthetic promoters Pending EP4676545A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
GBGB2303328.5A GB202303328D0 (en) 2023-03-07 2023-03-07 Synthetic promoters
PCT/GB2024/050608 WO2024184649A1 (en) 2023-03-07 2024-03-07 Synthetic promoters

Publications (1)

Publication Number Publication Date
EP4676545A1 true EP4676545A1 (en) 2026-01-14

Family

ID=85980181

Family Applications (1)

Application Number Title Priority Date Filing Date
EP24712568.5A Pending EP4676545A1 (en) 2023-03-07 2024-03-07 Synthetic promoters

Country Status (6)

Country Link
EP (1) EP4676545A1 (en)
JP (1) JP2026509256A (en)
CN (1) CN121219021A (en)
AU (1) AU2024230905A1 (en)
GB (1) GB202303328D0 (en)
WO (1) WO2024184649A1 (en)

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5223409A (en) 1988-09-02 1993-06-29 Protein Engineering Corp. Directed evolution of novel binding proteins
IL99552A0 (en) 1990-09-28 1992-08-18 Ixsys Inc Compositions containing procaryotic cells,a kit for the preparation of vectors useful for the coexpression of two or more dna sequences and methods for the use thereof
US5976873A (en) * 1994-05-18 1999-11-02 Children's Hospital Medical Center Nucleic acid sequences controlling lung cell-specific gene expression
GB201108879D0 (en) 2011-05-25 2011-07-06 Isis Innovation Vector
GB201118704D0 (en) 2011-10-28 2011-12-14 Univ Oxford Cystic fibrosis treatment
US20230190871A1 (en) * 2020-05-20 2023-06-22 Sana Biotechnology, Inc. Methods and compositions for treatment of viral infections
GB202105277D0 (en) 2021-04-13 2021-05-26 Imperial College Innovations Ltd Signal peptides

Also Published As

Publication number Publication date
AU2024230905A1 (en) 2025-10-23
CN121219021A (en) 2025-12-26
JP2026509256A (en) 2026-03-17
GB202303328D0 (en) 2023-04-19
WO2024184649A1 (en) 2024-09-12

Similar Documents

Publication Publication Date Title
KR102724029B1 (en) Modification of mammalian cells using artificial micro-RNAs and compositions of these products for altering the properties of mammalian cells
AU2018338790B2 (en) Non-human animals comprising a humanized TTR locus and methods of use
ES2640139T3 (en) Heavy chain mice of restricted immunoglobulin
US20230338477A1 (en) Anti-tfr:gaa and anti-cd63:gaa insertion for treatment of pompe disease
KR20190057104A (en) Non-human animal with hexanucleotide repeat extension in C9ORF72 locus
KR20210129108A (en) Compositions and methods for treating glycogen storage disease type 1A
KR102927708B1 (en) Non-human animals comprising a humanized TTR locus with a beta-slip mutation and methods of use thereof
KR102915369B1 (en) CRISPR and AAV Strategies for the Treatment of X-Linked Juvenile Retinoschisis
AU2019403015B2 (en) Nuclease-mediated repeat expansion
KR20230148824A (en) Compositions and methods for delivering nucleic acids
KR20230002788A (en) Artificial expression constructs for selectively modulating gene expression in neocortical layer 5 glutamatergic neurons
US20240197921A1 (en) Signal peptides
US20240325567A1 (en) Cell therapy
EP4676545A1 (en) Synthetic promoters
CN121909216A (en) Anti-TfR: acid sphingomyelinase for the treatment of acid sphingomyelinase deficiency
CN109337928B (en) Methods for improving the efficiency of gene therapy by overexpressing adeno-associated virus receptors
US20250304996A1 (en) Pseudotyped lentiviral vectors
US20250108133A1 (en) Vectors and compositions for gene augmentation of crumbs complex homologue 1 (crb1) mutations
RU2784927C1 (en) Animals other than human, including humanized ttr locus, and application methods
RU2833486C1 (en) Crispr and aav strategies for therapy of x-linked juvenile retinoschisis
CA3208936A1 (en) Retroviral vectors
CN117836420A (en) Recombinant TERT-encoding viral genome and vector
KR20260064747A (en) Anti-TfR:GAA and anti-CD63:GAA insertion for the treatment of Pompe disease
CN114621971A (en) Genetically modified non-human animal, and construction method and application thereof
CN118679250A (en) Anti-TfR:GAA and anti-CD63:GAA insertion for the treatment of Pompe disease

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20251007

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR