EP1616032A2 - Detecting gene expression in live cells using short-lived reporters with enzymatic amplification - Google Patents
Detecting gene expression in live cells using short-lived reporters with enzymatic amplificationInfo
- Publication number
- EP1616032A2 EP1616032A2 EP04749716A EP04749716A EP1616032A2 EP 1616032 A2 EP1616032 A2 EP 1616032A2 EP 04749716 A EP04749716 A EP 04749716A EP 04749716 A EP04749716 A EP 04749716A EP 1616032 A2 EP1616032 A2 EP 1616032A2
- Authority
- EP
- European Patent Office
- Prior art keywords
- cell
- gal
- cells
- reporter protein
- amino acid
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
- 230000014509 gene expression Effects 0.000 title claims abstract description 88
- 230000002255 enzymatic effect Effects 0.000 title claims abstract description 19
- 230000003321 amplification Effects 0.000 title abstract description 8
- 238000003199 nucleic acid amplification method Methods 0.000 title abstract description 8
- 238000000034 method Methods 0.000 claims abstract description 71
- 230000001413 cellular effect Effects 0.000 claims abstract description 39
- 210000004027 cell Anatomy 0.000 claims description 206
- 108090000623 proteins and genes Proteins 0.000 claims description 177
- 102000004169 proteins and genes Human genes 0.000 claims description 97
- 125000003729 nucleotide group Chemical group 0.000 claims description 87
- 239000002773 nucleotide Substances 0.000 claims description 81
- 150000001413 amino acids Chemical group 0.000 claims description 49
- 108010005774 beta-Galactosidase Proteins 0.000 claims description 44
- WQZGKKKJIJFFOK-FPRJBGLDSA-N beta-D-galactose Chemical compound OC[C@H]1O[C@@H](O)[C@H](O)[C@@H](O)[C@H]1O WQZGKKKJIJFFOK-FPRJBGLDSA-N 0.000 claims description 37
- 239000000758 substrate Substances 0.000 claims description 36
- 239000013612 plasmid Substances 0.000 claims description 33
- 108090000765 processed proteins & peptides Proteins 0.000 claims description 26
- 238000002493 microarray Methods 0.000 claims description 25
- 230000004927 fusion Effects 0.000 claims description 24
- 108090000848 Ubiquitin Proteins 0.000 claims description 23
- 102000044159 Ubiquitin Human genes 0.000 claims description 21
- 102000004190 Enzymes Human genes 0.000 claims description 18
- 108090000790 Enzymes Proteins 0.000 claims description 18
- 238000006243 chemical reaction Methods 0.000 claims description 16
- 230000035945 sensitivity Effects 0.000 claims description 15
- 239000007850 fluorescent dye Substances 0.000 claims description 14
- 230000000694 effects Effects 0.000 claims description 12
- ROHFNLRQFUQHCH-YFKPBYRVSA-N L-leucine Chemical compound CC(C)C[C@H](N)C(O)=O ROHFNLRQFUQHCH-YFKPBYRVSA-N 0.000 claims description 11
- ROHFNLRQFUQHCH-UHFFFAOYSA-N Leucine Natural products CC(C)CC(N)C(O)=O ROHFNLRQFUQHCH-UHFFFAOYSA-N 0.000 claims description 11
- UPSFMJHZUCSEHU-JYGUBCOQSA-N n-[(2s,3r,4r,5s,6r)-2-[(2r,3s,4r,5r,6s)-5-acetamido-4-hydroxy-2-(hydroxymethyl)-6-(4-methyl-2-oxochromen-7-yl)oxyoxan-3-yl]oxy-4,5-dihydroxy-6-(hydroxymethyl)oxan-3-yl]acetamide Chemical compound CC(=O)N[C@@H]1[C@@H](O)[C@H](O)[C@@H](CO)O[C@H]1O[C@H]1[C@H](O)[C@@H](NC(C)=O)[C@H](OC=2C=C3OC(=O)C=C(C)C3=CC=2)O[C@@H]1CO UPSFMJHZUCSEHU-JYGUBCOQSA-N 0.000 claims description 10
- 239000004475 Arginine Substances 0.000 claims description 9
- ODKSFYDXXFIFQN-BYPYZUCNSA-P L-argininium(2+) Chemical compound NC(=[NH2+])NCCC[C@H]([NH3+])C(O)=O ODKSFYDXXFIFQN-BYPYZUCNSA-P 0.000 claims description 9
- 108091032917 Transfer-messenger RNA Proteins 0.000 claims description 9
- ODKSFYDXXFIFQN-UHFFFAOYSA-N arginine Natural products OC(=O)C(N)CCCNC(N)=N ODKSFYDXXFIFQN-UHFFFAOYSA-N 0.000 claims description 9
- 210000001236 prokaryotic cell Anatomy 0.000 claims description 9
- 238000001514 detection method Methods 0.000 claims description 8
- 210000003527 eukaryotic cell Anatomy 0.000 claims description 8
- 125000001360 methionine group Chemical group N[C@@H](CCSC)C(=O)* 0.000 claims description 8
- 108010047754 beta-Glucosidase Proteins 0.000 claims description 7
- 239000003795 chemical substances by application Substances 0.000 claims description 7
- 108090000204 Dipeptidase 1 Proteins 0.000 claims description 6
- KDXKERNSBIXSRK-YFKPBYRVSA-N L-lysine Chemical compound NCCCC[C@H](N)C(O)=O KDXKERNSBIXSRK-YFKPBYRVSA-N 0.000 claims description 6
- COLNVLDHVKWLRT-QMMMGPOBSA-N L-phenylalanine Chemical compound OC(=O)[C@@H](N)CC1=CC=CC=C1 COLNVLDHVKWLRT-QMMMGPOBSA-N 0.000 claims description 6
- QIVBCDIJIAJPQS-VIFPVBQESA-N L-tryptophane Chemical compound C1=CC=C2C(C[C@H](N)C(O)=O)=CNC2=C1 QIVBCDIJIAJPQS-VIFPVBQESA-N 0.000 claims description 6
- OUYCCCASQSFEME-QMMMGPOBSA-N L-tyrosine Chemical compound OC(=O)[C@@H](N)CC1=CC=C(O)C=C1 OUYCCCASQSFEME-QMMMGPOBSA-N 0.000 claims description 6
- KDXKERNSBIXSRK-UHFFFAOYSA-N Lysine Natural products NCCCCC(N)C(O)=O KDXKERNSBIXSRK-UHFFFAOYSA-N 0.000 claims description 6
- 239000004472 Lysine Substances 0.000 claims description 6
- QIVBCDIJIAJPQS-UHFFFAOYSA-N Tryptophan Natural products C1=CC=C2C(CC(N)C(O)=O)=CNC2=C1 QIVBCDIJIAJPQS-UHFFFAOYSA-N 0.000 claims description 6
- 102000006995 beta-Glucosidase Human genes 0.000 claims description 6
- 102000006635 beta-lactamase Human genes 0.000 claims description 6
- COLNVLDHVKWLRT-UHFFFAOYSA-N phenylalanine Natural products OC(=O)C(N)CC1=CC=CC=C1 COLNVLDHVKWLRT-UHFFFAOYSA-N 0.000 claims description 6
- OUYCCCASQSFEME-UHFFFAOYSA-N tyrosine Natural products OC(=O)C(N)CC1=CC=C(O)C=C1 OUYCCCASQSFEME-UHFFFAOYSA-N 0.000 claims description 6
- QULZFZMEBOATFS-DISONHOPSA-N 7-[(2s,3r,4s,5r,6r)-3,4,5-trihydroxy-6-(hydroxymethyl)oxan-2-yl]oxyphenoxazin-3-one Chemical compound O[C@@H]1[C@@H](O)[C@@H](O)[C@@H](CO)O[C@H]1OC1=CC=C(N=C2C(=CC(=O)C=C2)O2)C2=C1 QULZFZMEBOATFS-DISONHOPSA-N 0.000 claims description 4
- 238000012544 monitoring process Methods 0.000 claims description 3
- 125000002924 primary amino group Chemical group [H]N([H])* 0.000 claims description 3
- 102000005936 beta-Galactosidase Human genes 0.000 claims 3
- 238000005558 fluorometry Methods 0.000 claims 2
- 238000000870 ultraviolet spectroscopy Methods 0.000 claims 2
- 101710199851 Copy number protein Proteins 0.000 claims 1
- 108700039691 Genetic Promoter Regions Proteins 0.000 claims 1
- 230000003094 perturbing effect Effects 0.000 claims 1
- 230000001052 transient effect Effects 0.000 abstract description 12
- 230000010198 maturation time Effects 0.000 abstract description 10
- 239000000203 mixture Substances 0.000 abstract description 10
- 230000002103 transcriptional effect Effects 0.000 abstract 1
- 241000588724 Escherichia coli Species 0.000 description 54
- 150000007523 nucleic acids Chemical group 0.000 description 51
- 101150066555 lacZ gene Proteins 0.000 description 41
- 108020004707 nucleic acids Proteins 0.000 description 41
- 102000039446 nucleic acids Human genes 0.000 description 41
- 239000013598 vector Substances 0.000 description 35
- 108020004414 DNA Proteins 0.000 description 31
- 108091028043 Nucleic acid sequence Proteins 0.000 description 21
- 108700008625 Reporter Genes Proteins 0.000 description 20
- 239000013604 expression vector Substances 0.000 description 20
- 230000007062 hydrolysis Effects 0.000 description 20
- 238000006460 hydrolysis reaction Methods 0.000 description 20
- 210000000349 chromosome Anatomy 0.000 description 17
- 238000009396 hybridization Methods 0.000 description 17
- 108010043121 Green Fluorescent Proteins Proteins 0.000 description 16
- 102000004144 Green Fluorescent Proteins Human genes 0.000 description 16
- 229940088598 enzyme Drugs 0.000 description 16
- 239000005090 green fluorescent protein Substances 0.000 description 16
- 239000000047 product Substances 0.000 description 16
- BRDJPCFGLMKJRU-UHFFFAOYSA-N DDAO Chemical compound ClC1=C(O)C(Cl)=C2C(C)(C)C3=CC(=O)C=CC3=NC2=C1 BRDJPCFGLMKJRU-UHFFFAOYSA-N 0.000 description 15
- 108091005957 yellow fluorescent proteins Proteins 0.000 description 15
- 108020004999 messenger RNA Proteins 0.000 description 14
- 238000010276 construction Methods 0.000 description 13
- 238000002474 experimental method Methods 0.000 description 13
- 238000003780 insertion Methods 0.000 description 13
- 230000037431 insertion Effects 0.000 description 13
- 239000003550 marker Substances 0.000 description 12
- 239000000523 sample Substances 0.000 description 12
- YMHOBZXQZVXHBM-UHFFFAOYSA-N 2,5-dimethoxy-4-bromophenethylamine Chemical compound COC1=CC(CCN)=C(OC)C=C1Br YMHOBZXQZVXHBM-UHFFFAOYSA-N 0.000 description 11
- 241000545067 Venus Species 0.000 description 11
- 229930182817 methionine Natural products 0.000 description 11
- 230000008569 process Effects 0.000 description 11
- 102000004196 processed proteins & peptides Human genes 0.000 description 11
- 102000007056 Recombinant Fusion Proteins Human genes 0.000 description 10
- 108010008281 Recombinant Fusion Proteins Proteins 0.000 description 10
- 239000012634 fragment Substances 0.000 description 10
- 238000005259 measurement Methods 0.000 description 10
- 230000001105 regulatory effect Effects 0.000 description 10
- 241000894006 Bacteria Species 0.000 description 9
- 241000863430 Shewanella Species 0.000 description 9
- 230000001580 bacterial effect Effects 0.000 description 9
- 239000002299 complementary DNA Substances 0.000 description 9
- 238000002073 fluorescence micrograph Methods 0.000 description 8
- 238000000338 in vitro Methods 0.000 description 8
- 230000010076 replication Effects 0.000 description 8
- 238000013519 translation Methods 0.000 description 8
- FFEARJCKVFRZRR-BYPYZUCNSA-N L-methionine Chemical compound CSCC[C@H](N)C(O)=O FFEARJCKVFRZRR-BYPYZUCNSA-N 0.000 description 7
- 240000004808 Saccharomyces cerevisiae Species 0.000 description 7
- 238000005516 engineering process Methods 0.000 description 7
- 230000006870 function Effects 0.000 description 7
- 102000037865 fusion proteins Human genes 0.000 description 7
- 108020001507 fusion proteins Proteins 0.000 description 7
- 230000001404 mediated effect Effects 0.000 description 7
- 238000010369 molecular cloning Methods 0.000 description 7
- 229920001184 polypeptide Polymers 0.000 description 7
- 238000003259 recombinant expression Methods 0.000 description 7
- 108091032973 (ribonucleotides)n+m Proteins 0.000 description 6
- 108020004705 Codon Proteins 0.000 description 6
- 230000008901 benefit Effects 0.000 description 6
- 230000000295 complement effect Effects 0.000 description 6
- 230000000875 corresponding effect Effects 0.000 description 6
- 238000013518 transcription Methods 0.000 description 6
- 230000035897 transcription Effects 0.000 description 6
- 230000017105 transposition Effects 0.000 description 6
- OPIFSICVWOWJMJ-AEOCFKNESA-N 5-bromo-4-chloro-3-indolyl beta-D-galactoside Chemical compound O[C@@H]1[C@@H](O)[C@@H](O)[C@@H](CO)O[C@H]1OC1=CNC2=CC=C(Br)C(Cl)=C12 OPIFSICVWOWJMJ-AEOCFKNESA-N 0.000 description 5
- 238000004458 analytical method Methods 0.000 description 5
- 238000013459 approach Methods 0.000 description 5
- 238000003491 array Methods 0.000 description 5
- 230000007613 environmental effect Effects 0.000 description 5
- 239000001963 growth medium Substances 0.000 description 5
- 239000002609 medium Substances 0.000 description 5
- 230000037361 pathway Effects 0.000 description 5
- 230000004044 response Effects 0.000 description 5
- 238000001712 DNA sequencing Methods 0.000 description 4
- 108010076504 Protein Sorting Signals Proteins 0.000 description 4
- FAPWRFPIFSIZLT-UHFFFAOYSA-M Sodium chloride Chemical compound [Na+].[Cl-] FAPWRFPIFSIZLT-UHFFFAOYSA-M 0.000 description 4
- 108091081024 Start codon Proteins 0.000 description 4
- 102000018390 Ubiquitin-Specific Proteases Human genes 0.000 description 4
- 108010066496 Ubiquitin-Specific Proteases Proteins 0.000 description 4
- 125000000637 arginyl group Chemical group N[C@@H](CCCNC(N)=N)C(=O)* 0.000 description 4
- 230000006399 behavior Effects 0.000 description 4
- 230000033228 biological regulation Effects 0.000 description 4
- 210000000170 cell membrane Anatomy 0.000 description 4
- 210000002421 cell wall Anatomy 0.000 description 4
- 230000005284 excitation Effects 0.000 description 4
- 239000011521 glass Substances 0.000 description 4
- 230000006801 homologous recombination Effects 0.000 description 4
- 238000002744 homologous recombination Methods 0.000 description 4
- 238000003384 imaging method Methods 0.000 description 4
- 238000004949 mass spectrometry Methods 0.000 description 4
- HSSLDCABUXLXKM-UHFFFAOYSA-N resorufin Chemical compound C1=CC(=O)C=C2OC3=CC(O)=CC=C3N=C21 HSSLDCABUXLXKM-UHFFFAOYSA-N 0.000 description 4
- 239000000126 substance Substances 0.000 description 4
- RWQNBRDOKXIBIV-UHFFFAOYSA-N thymine Chemical compound CC1=CNC(=O)NC1=O RWQNBRDOKXIBIV-UHFFFAOYSA-N 0.000 description 4
- 230000001960 triggered effect Effects 0.000 description 4
- 239000003981 vehicle Substances 0.000 description 4
- 229920000936 Agarose Polymers 0.000 description 3
- 102100026189 Beta-galactosidase Human genes 0.000 description 3
- 102000053602 DNA Human genes 0.000 description 3
- 102000004163 DNA-directed RNA polymerases Human genes 0.000 description 3
- DHMQDGOQFOQNFH-UHFFFAOYSA-N Glycine Natural products NCC(O)=O DHMQDGOQFOQNFH-UHFFFAOYSA-N 0.000 description 3
- 241000124008 Mammalia Species 0.000 description 3
- 125000001429 N-terminal alpha-amino-acid group Chemical group 0.000 description 3
- 125000000729 N-terminal amino-acid group Chemical group 0.000 description 3
- 229930006000 Sucrose Natural products 0.000 description 3
- CZMRCDWAGMRECN-UGDNZRGBSA-N Sucrose Chemical compound O[C@H]1[C@H](O)[C@@H](CO)O[C@@]1(CO)O[C@@H]1[C@H](O)[C@@H](O)[C@H](O)[C@@H](CO)O1 CZMRCDWAGMRECN-UGDNZRGBSA-N 0.000 description 3
- 229960000723 ampicillin Drugs 0.000 description 3
- AVKUERGKIZMTKX-NJBDSQKTSA-N ampicillin Chemical compound C1([C@@H](N)C(=O)N[C@H]2[C@H]3SC([C@@H](N3C2=O)C(O)=O)(C)C)=CC=CC=C1 AVKUERGKIZMTKX-NJBDSQKTSA-N 0.000 description 3
- 238000000137 annealing Methods 0.000 description 3
- 230000031018 biological processes and functions Effects 0.000 description 3
- 230000015556 catabolic process Effects 0.000 description 3
- 238000003776 cleavage reaction Methods 0.000 description 3
- 238000006731 degradation reaction Methods 0.000 description 3
- 238000010494 dissociation reaction Methods 0.000 description 3
- 230000005593 dissociations Effects 0.000 description 3
- 229910052739 hydrogen Inorganic materials 0.000 description 3
- 239000001257 hydrogen Substances 0.000 description 3
- 238000001727 in vivo Methods 0.000 description 3
- 238000002372 labelling Methods 0.000 description 3
- 125000001909 leucine group Chemical group [H]N(*)C(C(*)=O)C([H])([H])C(C([H])([H])[H])C([H])([H])[H] 0.000 description 3
- 230000000670 limiting effect Effects 0.000 description 3
- 238000004519 manufacturing process Methods 0.000 description 3
- 238000013507 mapping Methods 0.000 description 3
- 230000035772 mutation Effects 0.000 description 3
- 230000035699 permeability Effects 0.000 description 3
- 238000006116 polymerization reaction Methods 0.000 description 3
- 108091033319 polynucleotide Proteins 0.000 description 3
- 102000040430 polynucleotide Human genes 0.000 description 3
- 239000002157 polynucleotide Substances 0.000 description 3
- 230000002797 proteolythic effect Effects 0.000 description 3
- 238000000746 purification Methods 0.000 description 3
- 230000007017 scission Effects 0.000 description 3
- 238000012216 screening Methods 0.000 description 3
- 238000004904 shortening Methods 0.000 description 3
- 230000011664 signaling Effects 0.000 description 3
- 239000000243 solution Substances 0.000 description 3
- 239000005720 sucrose Substances 0.000 description 3
- 230000008685 targeting Effects 0.000 description 3
- 210000001519 tissue Anatomy 0.000 description 3
- 238000001890 transfection Methods 0.000 description 3
- 210000005253 yeast cell Anatomy 0.000 description 3
- BUAQBZPEZUPDRF-RHQZKXFESA-N (2r,3r,4r,5r,6s)-2-(hydroxymethyl)-6-[(2-hydroxy-7,8,9,10-tetrahydro-6h-benzo[c]chromen-3-yl)oxy]oxane-3,4,5-triol Chemical compound O[C@@H]1[C@H](O)[C@@H](O)[C@@H](CO)O[C@H]1OC(C(=C1)O)=CC2=C1C(CCCC1)=C1CO2 BUAQBZPEZUPDRF-RHQZKXFESA-N 0.000 description 2
- 108090000994 Catalytic RNA Proteins 0.000 description 2
- 102000053642 Catalytic RNA Human genes 0.000 description 2
- 108020004635 Complementary DNA Proteins 0.000 description 2
- 108090000626 DNA-directed RNA polymerases Proteins 0.000 description 2
- RTZKZFJDLAIYFH-UHFFFAOYSA-N Diethyl ether Chemical compound CCOCC RTZKZFJDLAIYFH-UHFFFAOYSA-N 0.000 description 2
- WQZGKKKJIJFFOK-GASJEMHNSA-N Glucose Natural products OC[C@H]1OC(O)[C@H](O)[C@@H](O)[C@@H]1O WQZGKKKJIJFFOK-GASJEMHNSA-N 0.000 description 2
- 102000004366 Glucosidases Human genes 0.000 description 2
- 108010056771 Glucosidases Proteins 0.000 description 2
- 102000005720 Glutathione transferase Human genes 0.000 description 2
- 108010070675 Glutathione transferase Proteins 0.000 description 2
- 239000004471 Glycine Substances 0.000 description 2
- 102100034343 Integrase Human genes 0.000 description 2
- 108010054278 Lac Repressors Proteins 0.000 description 2
- 108091005804 Peptidases Proteins 0.000 description 2
- 239000004365 Protease Substances 0.000 description 2
- 102100037681 Protein FEV Human genes 0.000 description 2
- 101710198166 Protein FEV Proteins 0.000 description 2
- 108010092799 RNA-directed DNA polymerase Proteins 0.000 description 2
- 241001223867 Shewanella oneidensis Species 0.000 description 2
- ISAKRJDGNUQOIC-UHFFFAOYSA-N Uracil Chemical compound O=C1C=CNC(=O)N1 ISAKRJDGNUQOIC-UHFFFAOYSA-N 0.000 description 2
- 241000700605 Viruses Species 0.000 description 2
- 238000010521 absorption reaction Methods 0.000 description 2
- 238000000862 absorption spectrum Methods 0.000 description 2
- 239000011543 agarose gel Substances 0.000 description 2
- 238000003556 assay Methods 0.000 description 2
- 101150030789 bglB gene Proteins 0.000 description 2
- 210000004899 c-terminal region Anatomy 0.000 description 2
- 229960005091 chloramphenicol Drugs 0.000 description 2
- WIIZWVCIJKGZOK-RKDXNWHRSA-N chloramphenicol Chemical compound ClC(Cl)C(=O)N[C@H](CO)[C@H](O)C1=CC=C([N+]([O-])=O)C=C1 WIIZWVCIJKGZOK-RKDXNWHRSA-N 0.000 description 2
- 238000010367 cloning Methods 0.000 description 2
- OPTASPLRGRRNAP-UHFFFAOYSA-N cytosine Chemical compound NC=1C=CNC(=O)N=1 OPTASPLRGRRNAP-UHFFFAOYSA-N 0.000 description 2
- 238000012217 deletion Methods 0.000 description 2
- 230000037430 deletion Effects 0.000 description 2
- 238000004520 electroporation Methods 0.000 description 2
- 238000000295 emission spectrum Methods 0.000 description 2
- 238000000799 fluorescence microscopy Methods 0.000 description 2
- UYTPUPDQBNUYGX-UHFFFAOYSA-N guanine Chemical compound O=C1NC(N)=NC2=C1N=CN2 UYTPUPDQBNUYGX-UHFFFAOYSA-N 0.000 description 2
- 238000010348 incorporation Methods 0.000 description 2
- 230000006698 induction Effects 0.000 description 2
- 230000001939 inductive effect Effects 0.000 description 2
- 101150001899 lacY gene Proteins 0.000 description 2
- 239000007788 liquid Substances 0.000 description 2
- 239000002207 metabolite Substances 0.000 description 2
- 239000006151 minimal media Substances 0.000 description 2
- 230000004048 modification Effects 0.000 description 2
- 238000012986 modification Methods 0.000 description 2
- 238000002703 mutagenesis Methods 0.000 description 2
- 231100000350 mutagenesis Toxicity 0.000 description 2
- 235000015097 nutrients Nutrition 0.000 description 2
- 230000003647 oxidation Effects 0.000 description 2
- 238000007254 oxidation reaction Methods 0.000 description 2
- 230000036961 partial effect Effects 0.000 description 2
- 239000000816 peptidomimetic Substances 0.000 description 2
- 239000012466 permeate Substances 0.000 description 2
- 230000008488 polyadenylation Effects 0.000 description 2
- 230000001323 posttranslational effect Effects 0.000 description 2
- 238000012545 processing Methods 0.000 description 2
- 230000026447 protein localization Effects 0.000 description 2
- 108091092562 ribozyme Proteins 0.000 description 2
- 101150025220 sacB gene Proteins 0.000 description 2
- 238000011896 sensitive detection Methods 0.000 description 2
- 239000011780 sodium chloride Substances 0.000 description 2
- 239000001509 sodium citrate Substances 0.000 description 2
- NLJMYIDDQXHKNR-UHFFFAOYSA-K sodium citrate Chemical compound O.O.[Na+].[Na+].[Na+].[O-]C(=O)CC(O)(CC([O-])=O)C([O-])=O NLJMYIDDQXHKNR-UHFFFAOYSA-K 0.000 description 2
- 230000002123 temporal effect Effects 0.000 description 2
- 238000012360 testing method Methods 0.000 description 2
- 229940113082 thymine Drugs 0.000 description 2
- 238000000492 total internal reflection fluorescence microscopy Methods 0.000 description 2
- 230000005026 transcription initiation Effects 0.000 description 2
- 230000009466 transformation Effects 0.000 description 2
- 241000701447 unidentified baculovirus Species 0.000 description 2
- 241001515965 unidentified phage Species 0.000 description 2
- 239000013603 viral vector Substances 0.000 description 2
- 230000003612 virological effect Effects 0.000 description 2
- OWEGMIWEEQEYGQ-UHFFFAOYSA-N 100676-05-9 Natural products OC1C(O)C(O)C(CO)OC1OCC1C(O)C(O)C(O)C(OC2C(OC(O)C(O)C2O)CO)O1 OWEGMIWEEQEYGQ-UHFFFAOYSA-N 0.000 description 1
- 229930024421 Adenine Natural products 0.000 description 1
- GFFGJBXGBJISGV-UHFFFAOYSA-N Adenine Chemical compound NC1=NC=NC2=C1N=CN2 GFFGJBXGBJISGV-UHFFFAOYSA-N 0.000 description 1
- 241000194110 Bacillus sp. (in: Bacteria) Species 0.000 description 1
- UXVMQQNJUSDDNG-UHFFFAOYSA-L Calcium chloride Chemical compound [Cl-].[Cl-].[Ca+2] UXVMQQNJUSDDNG-UHFFFAOYSA-L 0.000 description 1
- 108091026890 Coding region Proteins 0.000 description 1
- -1 DNA Chemical class 0.000 description 1
- 238000000018 DNA microarray Methods 0.000 description 1
- 230000006820 DNA synthesis Effects 0.000 description 1
- 108010014303 DNA-directed DNA polymerase Proteins 0.000 description 1
- 102000016928 DNA-directed DNA polymerase Human genes 0.000 description 1
- 241000702421 Dependoparvovirus Species 0.000 description 1
- 229920002307 Dextran Polymers 0.000 description 1
- 108010013369 Enteropeptidase Proteins 0.000 description 1
- 102100029727 Enteropeptidase Human genes 0.000 description 1
- 241000701959 Escherichia virus Lambda Species 0.000 description 1
- 108700024394 Exon Proteins 0.000 description 1
- 108010074860 Factor Xa Proteins 0.000 description 1
- 241000282326 Felis catus Species 0.000 description 1
- 241000238631 Hexapoda Species 0.000 description 1
- 238000012404 In vitro experiment Methods 0.000 description 1
- 108091092195 Intron Proteins 0.000 description 1
- 108090001030 Lipoproteins Proteins 0.000 description 1
- 102000004895 Lipoproteins Human genes 0.000 description 1
- GUBGYTABKSRVRQ-PICCSMPSSA-N Maltose Natural products O[C@@H]1[C@@H](O)[C@H](O)[C@@H](CO)O[C@@H]1O[C@@H]1[C@@H](CO)OC(O)[C@H](O)[C@H]1O GUBGYTABKSRVRQ-PICCSMPSSA-N 0.000 description 1
- 108010021466 Mutant Proteins Proteins 0.000 description 1
- 102000008300 Mutant Proteins Human genes 0.000 description 1
- 238000000636 Northern blotting Methods 0.000 description 1
- 108020004711 Nucleic Acid Probes Proteins 0.000 description 1
- 108010038807 Oligopeptides Proteins 0.000 description 1
- 102000015636 Oligopeptides Human genes 0.000 description 1
- 108700026244 Open Reading Frames Proteins 0.000 description 1
- 102000035195 Peptidases Human genes 0.000 description 1
- 102000009572 RNA Polymerase II Human genes 0.000 description 1
- 108010009460 RNA Polymerase II Proteins 0.000 description 1
- 108020004511 Recombinant DNA Proteins 0.000 description 1
- 102100037486 Reverse transcriptase/ribonuclease H Human genes 0.000 description 1
- 108091028664 Ribonucleotide Proteins 0.000 description 1
- 235000014680 Saccharomyces cerevisiae Nutrition 0.000 description 1
- 101900335766 Saccharomyces cerevisiae Ubiquitin Proteins 0.000 description 1
- 241000235343 Saccharomycetales Species 0.000 description 1
- 238000012300 Sequence Analysis Methods 0.000 description 1
- 238000002105 Southern blotting Methods 0.000 description 1
- 101100309436 Streptococcus mutans serotype c (strain ATCC 700610 / UA159) ftf gene Proteins 0.000 description 1
- 108020005038 Terminator Codon Proteins 0.000 description 1
- 108090000190 Thrombin Proteins 0.000 description 1
- 108020004566 Transfer RNA Proteins 0.000 description 1
- 108020000999 Viral RNA Proteins 0.000 description 1
- 108091005971 Wild-type GFP Proteins 0.000 description 1
- FRYDSOYOHWGSMD-UHFFFAOYSA-N [C].O Chemical class [C].O FRYDSOYOHWGSMD-UHFFFAOYSA-N 0.000 description 1
- 238000009825 accumulation Methods 0.000 description 1
- 229960000643 adenine Drugs 0.000 description 1
- 238000001261 affinity purification Methods 0.000 description 1
- 238000005054 agglomeration Methods 0.000 description 1
- 230000002776 aggregation Effects 0.000 description 1
- 230000004075 alteration Effects 0.000 description 1
- 239000012736 aqueous medium Substances 0.000 description 1
- 108091008324 binding proteins Proteins 0.000 description 1
- 102000023732 binding proteins Human genes 0.000 description 1
- 238000005842 biochemical reaction Methods 0.000 description 1
- 230000003115 biocidal effect Effects 0.000 description 1
- 229920000249 biocompatible polymer Polymers 0.000 description 1
- 230000032770 biofilm formation Effects 0.000 description 1
- 238000005422 blasting Methods 0.000 description 1
- 230000008499 blood brain barrier function Effects 0.000 description 1
- 210000001218 blood-brain barrier Anatomy 0.000 description 1
- 238000010804 cDNA synthesis Methods 0.000 description 1
- 239000001110 calcium chloride Substances 0.000 description 1
- 229910001628 calcium chloride Inorganic materials 0.000 description 1
- 239000001506 calcium phosphate Substances 0.000 description 1
- 229910000389 calcium phosphate Inorganic materials 0.000 description 1
- 235000011010 calcium phosphates Nutrition 0.000 description 1
- 239000003054 catalyst Substances 0.000 description 1
- 238000004113 cell culture Methods 0.000 description 1
- 230000032823 cell division Effects 0.000 description 1
- 230000004670 cellular proteolysis Effects 0.000 description 1
- 230000008859 change Effects 0.000 description 1
- 239000013043 chemical agent Substances 0.000 description 1
- 239000012707 chemical precursor Substances 0.000 description 1
- 239000003153 chemical reaction reagent Substances 0.000 description 1
- 239000003593 chromogenic compound Substances 0.000 description 1
- 230000002759 chromosomal effect Effects 0.000 description 1
- 238000012411 cloning technique Methods 0.000 description 1
- 238000000975 co-precipitation Methods 0.000 description 1
- 238000012875 competitive assay Methods 0.000 description 1
- 150000001875 compounds Chemical class 0.000 description 1
- 238000004624 confocal microscopy Methods 0.000 description 1
- 230000021615 conjugation Effects 0.000 description 1
- 238000011109 contamination Methods 0.000 description 1
- 239000013068 control sample Substances 0.000 description 1
- 230000002079 cooperative effect Effects 0.000 description 1
- 239000006059 cover glass Substances 0.000 description 1
- 239000003431 cross linking reagent Substances 0.000 description 1
- 238000012258 culturing Methods 0.000 description 1
- 229940104302 cytosine Drugs 0.000 description 1
- 230000002950 deficient Effects 0.000 description 1
- 239000005547 deoxyribonucleotide Substances 0.000 description 1
- 125000002637 deoxyribonucleotide group Chemical group 0.000 description 1
- 238000013461 design Methods 0.000 description 1
- 239000008121 dextrose Substances 0.000 description 1
- 238000002405 diagnostic procedure Methods 0.000 description 1
- 238000009826 distribution Methods 0.000 description 1
- 238000001035 drying Methods 0.000 description 1
- 239000002158 endotoxin Substances 0.000 description 1
- 239000003623 enhancer Substances 0.000 description 1
- 230000007071 enzymatic hydrolysis Effects 0.000 description 1
- 238000006047 enzymatic hydrolysis reaction Methods 0.000 description 1
- 238000006911 enzymatic reaction Methods 0.000 description 1
- 102000052116 epidermal growth factor receptor activity proteins Human genes 0.000 description 1
- 108700015053 epidermal growth factor receptor activity proteins Proteins 0.000 description 1
- 150000002148 esters Chemical class 0.000 description 1
- 239000012530 fluid Substances 0.000 description 1
- 230000002538 fungal effect Effects 0.000 description 1
- 238000012215 gene cloning Methods 0.000 description 1
- 238000011223 gene expression profiling Methods 0.000 description 1
- 238000001415 gene therapy Methods 0.000 description 1
- 230000002068 genetic effect Effects 0.000 description 1
- 239000008103 glucose Substances 0.000 description 1
- 229930182478 glucoside Natural products 0.000 description 1
- 150000008131 glucosides Chemical group 0.000 description 1
- 150000004676 glycans Chemical class 0.000 description 1
- 125000003630 glycyl group Chemical group [H]N([H])C([H])([H])C(*)=O 0.000 description 1
- 101150036612 gnl gene Proteins 0.000 description 1
- 239000005556 hormone Substances 0.000 description 1
- 229940088597 hormone Drugs 0.000 description 1
- 101150090192 how gene Proteins 0.000 description 1
- 238000005286 illumination Methods 0.000 description 1
- 230000001771 impaired effect Effects 0.000 description 1
- 238000011534 incubation Methods 0.000 description 1
- 239000000411 inducer Substances 0.000 description 1
- 208000015181 infectious disease Diseases 0.000 description 1
- 230000000977 initiatory effect Effects 0.000 description 1
- 238000011081 inoculation Methods 0.000 description 1
- 230000003993 interaction Effects 0.000 description 1
- 239000000138 intercalating agent Substances 0.000 description 1
- 230000003834 intracellular effect Effects 0.000 description 1
- 238000007852 inverse PCR Methods 0.000 description 1
- 239000010410 layer Substances 0.000 description 1
- 239000003446 ligand Substances 0.000 description 1
- 238000001638 lipofection Methods 0.000 description 1
- 229920006008 lipopolysaccharide Polymers 0.000 description 1
- 239000002502 liposome Substances 0.000 description 1
- 239000006194 liquid suspension Substances 0.000 description 1
- 210000004962 mammalian cell Anatomy 0.000 description 1
- 239000000463 material Substances 0.000 description 1
- 230000007246 mechanism Effects 0.000 description 1
- 239000012528 membrane Substances 0.000 description 1
- 210000004779 membrane envelope Anatomy 0.000 description 1
- 239000002923 metal particle Substances 0.000 description 1
- 230000000813 microbial effect Effects 0.000 description 1
- 238000000386 microscopy Methods 0.000 description 1
- 108091005573 modified proteins Proteins 0.000 description 1
- 102000035118 modified proteins Human genes 0.000 description 1
- 238000001823 molecular biology technique Methods 0.000 description 1
- 239000003068 molecular probe Substances 0.000 description 1
- YOHYSYJDKVYCJI-UHFFFAOYSA-N n-[3-[[6-[3-(trifluoromethyl)anilino]pyrimidin-4-yl]amino]phenyl]cyclopropanecarboxamide Chemical compound FC(F)(F)C1=CC=CC(NC=2N=CN=C(NC=3C=C(NC(=O)C4CC4)C=CC=3)C=2)=C1 YOHYSYJDKVYCJI-UHFFFAOYSA-N 0.000 description 1
- 229920005615 natural polymer Polymers 0.000 description 1
- 239000002853 nucleic acid probe Substances 0.000 description 1
- 230000003287 optical effect Effects 0.000 description 1
- 239000000575 pesticide Substances 0.000 description 1
- 239000008177 pharmaceutical agent Substances 0.000 description 1
- 239000003016 pheromone Substances 0.000 description 1
- 230000004260 plant-type cell wall biogenesis Effects 0.000 description 1
- 229920000642 polymer Polymers 0.000 description 1
- 238000003752 polymerase chain reaction Methods 0.000 description 1
- 229920001282 polysaccharide Polymers 0.000 description 1
- 239000005017 polysaccharide Substances 0.000 description 1
- 230000037452 priming Effects 0.000 description 1
- 230000001737 promoting effect Effects 0.000 description 1
- 230000006337 proteolytic cleavage Effects 0.000 description 1
- 230000002285 radioactive effect Effects 0.000 description 1
- 239000011535 reaction buffer Substances 0.000 description 1
- 238000010223 real-time analysis Methods 0.000 description 1
- 238000011897 real-time detection Methods 0.000 description 1
- 238000010188 recombinant method Methods 0.000 description 1
- 230000006798 recombination Effects 0.000 description 1
- 238000005215 recombination Methods 0.000 description 1
- 230000002829 reductive effect Effects 0.000 description 1
- 230000022532 regulation of transcription, DNA-dependent Effects 0.000 description 1
- 239000002336 ribonucleotide Substances 0.000 description 1
- 125000002652 ribonucleotide group Chemical group 0.000 description 1
- 108020004418 ribosomal RNA Proteins 0.000 description 1
- 238000000926 separation method Methods 0.000 description 1
- 239000002356 single layer Substances 0.000 description 1
- 238000004611 spectroscopical analysis Methods 0.000 description 1
- 238000001228 spectrum Methods 0.000 description 1
- 238000012409 standard PCR amplification Methods 0.000 description 1
- 238000003860 storage Methods 0.000 description 1
- 230000004083 survival effect Effects 0.000 description 1
- 229920001059 synthetic polymer Polymers 0.000 description 1
- 230000009897 systematic effect Effects 0.000 description 1
- 238000011191 terminal modification Methods 0.000 description 1
- 229960004072 thrombin Drugs 0.000 description 1
- 238000000204 total internal reflection microscopy Methods 0.000 description 1
- 239000003053 toxin Substances 0.000 description 1
- 231100000765 toxin Toxicity 0.000 description 1
- 108700012359 toxins Proteins 0.000 description 1
- 238000012546 transfer Methods 0.000 description 1
- 230000001131 transforming effect Effects 0.000 description 1
- QORWJWZARLRLPR-UHFFFAOYSA-H tricalcium bis(phosphate) Chemical compound [Ca+2].[Ca+2].[Ca+2].[O-]P([O-])([O-])=O.[O-]P([O-])([O-])=O QORWJWZARLRLPR-UHFFFAOYSA-H 0.000 description 1
- 230000007306 turnover Effects 0.000 description 1
- 230000014848 ubiquitin-dependent protein catabolic process Effects 0.000 description 1
- 241000701161 unidentified adenovirus Species 0.000 description 1
- 241001430294 unidentified retrovirus Species 0.000 description 1
- 238000011144 upstream manufacturing Methods 0.000 description 1
- 229940035893 uracil Drugs 0.000 description 1
- 238000005406 washing Methods 0.000 description 1
Classifications
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N33/00—Investigating or analysing materials by specific methods not covered by groups G01N1/00 - G01N31/00
- G01N33/48—Biological material, e.g. blood, urine; Haemocytometers
- G01N33/50—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing
- G01N33/5005—Chemical analysis of biological material, e.g. blood, urine; Testing involving biospecific ligand binding methods; Immunological testing involving human or animal cells
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/02—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving viable microorganisms
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B01—PHYSICAL OR CHEMICAL PROCESSES OR APPARATUS IN GENERAL
- B01J—CHEMICAL OR PHYSICAL PROCESSES, e.g. CATALYSIS OR COLLOID CHEMISTRY; THEIR RELEVANT APPARATUS
- B01J2219/00—Chemical, physical or physico-chemical processes in general; Their relevant apparatus
- B01J2219/00274—Sequential or parallel reactions; Apparatus and devices for combinatorial chemistry or for making arrays; Chemical library technology
- B01J2219/00718—Type of compounds synthesised
- B01J2219/0072—Organic compounds
- B01J2219/0074—Biological products
- B01J2219/00743—Cells
Definitions
- the present invention relates to compositions and methods for detecting and analyzing gene expression events occurring in live cells.
- GFP green fluorescent protein
- GFP green fluorescent protein
- Most applications of GFPs have focused on mapping protein localization via fusion constructs.
- current GFPs are not suitable for following fast biological processes on the time-scale of minutes or less. This is primarily due to the fact that GFPs in the cellular environment have a long post-translational maturation time, which is required for the oxidation of the three residues forming the GFP fluorophore.
- a new GFP variant with faster maturation time is needed.
- one GFP molecule only provides one fluorophore, thus it is only suitable for detecting translational product that expressed at high levels.
- the present invention pertains to compositions and methods for detecting and analyzing gene expression events occurring in live cells.
- the present invention pertains to a short-lived reporter with enzymatic amplification.
- the reporters of the present invention have relatively short maturation time and a short cellular lifetime which can be exploited to detect transient events of gene expression in live cells.
- compositions and methods for employing one or more reporters having a short maturation time and a short cellular lifetime to detect transient events of gene expression in individual living cells with high sensitivity and high time resolution are described.
- a reporter gene system employing a reporter, for example, /J-galactosidase ( J-gal).
- the reporter is manipulated in such a manner so as to decrease its cellular lifetime.
- the so-called N-end rule to shorten the cellular lifetime of -gal is utilized.
- the N-end rule states that the cellular lifetime of a protein is related to its N-terminal amino acid residue. This rule applies to all organisms ranging from bacteria to mammals. In E. coli, changing the N-terminal amino acid from the natural methionine to leucine, arginine, lysine, phenylalanine, tryptophan or tyrosine shortens the protein half-life to a few minutes.
- the ubiquitin (ub) fusion technique is used to introduce a lifetime-shortening amino acid (e.g., leucine or arginine) in place of the methionine at the N-terminus of, for example, /J-gal to generate Ub-Leu- J-gal or Ub-Arg- -gal.
- a lifetime-shortening amino acid e.g., leucine or arginine
- the ubiquitin will be cleaved by an ubiquitin-specific protease, thus exposing the leucine or arginine residue and targeting the protein for the proteolytic pathways.
- other means of modifying /3-gal's cellular lifetime are also employed, such as N-terminal and C-terminal signal peptides fusions.
- live-cell microarrays are described.
- multiple libraries of cells are prepared each differing in at least one genotypic property (i.e., the genotype of each cell is different, for example, the reporter gene is inserted at a different position on the chromosome, thereby tagging an operon or a gene).
- a live-cell microarray is comprised of two libraries.
- One library comprises cells each of which has a promoterless lacZ gene encoding for a short-lived -gal with its own ribosome binding site that is inserted into one promoter controlled region in the host cell's genome.
- the second library comprises the same elements except that a gene encoding for a short-lived yellow fluorescent protein YFP (Venus-ssrA) replaces the gene encoding for a short-lived /3-gal in the first library.
- YFP yellow fluorescent protein
- FIG. 1 (a) are chemical structures of 9H-(l,3-dichloro-9,9-dimethylacridin-2- one-7-yl) 3-D-galactopyranoside (DDAO-gal) and its fluorescent product DDAO after hydrolysis by /3-galactosidase, and (b) is a graphical representation of the absorption and emission spectra of DDAO; (c) are chemical structures of resorufin- glucopyranoside (resorufin-glu) and its fluorescent product resorufin after hydrolysis by -glucosidase and (d) is a graphical representation of the absorption and emission spectra of resorufin;
- FIG. 2 (a) shows the location of the gene coding for Ub-Arg- -gal in the lac operon and (b) depicts the nucleic acids and amino acids sequences of Ub-Arg- -gal. Only the sequences of ubiquitin (light-shaded), the arginine residue immediately after ubiquitin, and the linker peptide (unshaded) between ubiquitin and the beginning of -gal (dark-shaded) are shown. Please note that the -gal in this construct lacks its first twenty two amino acids;
- FIG. 3 (a) is a graph measuring the hydrolysis of DDAO-gal in the presence of enzyme products from different gene constructs, and (b) are the amino acid (top) and nucleotide sequence (bottom) for each of the different construct; Please note that only the sequences that differ in these constructs (N-terminus of the lacZ gene) are shown in (b);
- FIG. 4 is a graph showing the DDAO fluorescence generated from the hydrolysis of DDAO-gal by wild type lac77 cells (dark) but not by the lacZ cells (light);
- FIG. 5 depicts the sequence junction of lacZ deletion, wherein the sequence is from the EcorV site of the lacl gene to Nspl site of the lacY gene, and (b) is the amino acid sequence and nucleic acid sequence wherein the numbering of the nucleotides is according to the first base of the lacl gene, the lacZ gene is replaced by lacY gene from the ATG starting codon, the amino acids sequences are shown on top of the DNA sequence panel;
- FIG. 6 is a fluorescence image of E. coli Cells. The signal is from DDAO generated by the basal level expression of unmodified /3-gal;
- FIG. 7 (a) is the fluorescence images observed on single E.coli cells with a gene coding for a short-lived Ub-Arg-j ⁇ -gal incorporated on chromosome. The signal is generated by the basal level expression of ⁇ -gal, For cell 1, only thirteen fluorescence images of cell 1 are shown in fifteen minute intervals for simplicity reasons. For cell 2, the fluorescence images are shown in five minute intervals, (b) is a fluorescence measruement of the production and degradation of /3-gal in one singe E.coli cell under TIR fluorescence microscope;
- FIG. 8 (a) depicts the sequence for the short-lived YFP: Venus-ssrA construct on plasmid pVS5, and (b) is the amino acid sequence and nucleic acid sequence wherein the sequence is from the first base of the yfp gene and to the end of the yfp gene with the addition of 33 bases coding for the ssrA tag;
- FIG. 9 is a graph showing the resorufin fluorescence generated from the hydrolysis of resorufin-glu by E.coli cells expressing -glucosidase ( bgl + , light) but not by the bglB ' cells (dark);
- FIG. 10 is a schematic drawing of the construction of a lacZ library by Tn5 mediated transposition
- FIG. 11 is a schematic drawing of the constructing a lacZ and yfp library
- FIG. 12 is a flow chart showing an automated process for the construction of libraries and the fabrication of the cell array
- FIG. 13(a) is the plasmid map for pBBRlMCS-5.1
- (b) is the nucleotide sequence coding for the short-lived b-gal for the plasmid depicted in (a). Please note that only the sequence at the N-terminus of the ub-leu-lacZ gene is shown;
- FIG. 14 (a) is a fluroescence image of Shewanella oneideinis cells expressing /3-gal from the lacZ7 plasmid pBBR!MCS5.1, and (b) is a graph showing the DDAO fluorescence generated by the hydrolysis of DDAO-gal under various conditions;
- FIG. 15 (a) depicts the nucleotide sequence junction of ub-leu-lacZ gene in Saccharomyce cerevisiae and (b) is the amino acid sequence and nucleic acid sequence for the junction of the ub-leu-lacZ construct on centromeric plasmid transformed into Saccharomyce cerevisiae cell;
- FIG. 16 represents DDAO fluorescence generated from the hydrolysis of DDAO-gal by wild type / ⁇ cZ + cells (dark) but not by the lacZ ' cells (light) in Saccharomyce cerevisiae;
- FIG. 17 is a fluorescence image of S. cerevisiae cells containing unmodified /3-gal.
- FIG. 18 is the fluorescence signal bursts observed on a single S. cerevisiae cell with a short-lived J-gal expressed from a centromeric plasmid.
- the present invention pertains to compositions and methods for detecting and analyzing gene expression events occurring in individual living cells.
- the present invention pertains to short-lived reporters with enzymatic amplification. These reporters of the present invention have relatively short maturation time and a short cellular lifetimes which can be exploited to detect transient events of gene expression in live cells.
- the low time resolutions prevent studies of transient gene expression processes, for example, those involved in cell division; (3) they are not sensitive to low copy number gene products, which often play a prominent role in cellular sensing, signaling and gene regulation; and (4) they can only provide averaged results of large populations of cells rather than behaviors of individual cells: transient and stochastic gene expression events are often masked in the population measurements.
- a method for employing one or moire reporters having a short maturation time and a short cellular lifetime to detect transient events of gene expression in live cells with high sensitivity and a fast time resolution is described.
- a reporting system for monitoring real-time gene expression events in a living cell.
- This reporting system comprises an illuminogenic substrate, wherein said substrate is permeable to said cell.
- the system also comprises at least one reporter protein, wherein said reporter protein facilitates the conversion of said illuminogenic substrate into an illuminescent molecule, and wherein said reporter protein has a short cellular life time.
- the cell can be a prokaryote or eurokaryote.
- the illuminogenic substrate can be any substrate that when acted upon by, for example, hydrolysis, will generate an illuminescent product which emits photons.
- the substrate can be a fluorogenic substrate that when acted upon will generate a fluorescent product that emits fluorescence.
- Chemiluminescence substrates can also be used in the present invention.
- the term illuminogenic is also meant to cover absorption in addition to photon emission, for example, chromogenic substrates.
- the present embodiment is designed to capitalize on the recent advances in sensitive fluorescence microscopy. In the past years, tremendous progress has been made in fluorescence imaging of single-molecules, even in living cells.
- GFP green fluorescent protein
- derivatives thereof as reporter proteins. See, for example, Bongaerts, R.J., et al., Green fluorescent protein as a marker for conditional gene expression in bacterial cells. Methods Enzymol, 2002. 358: p. 43-66; Tsien, R.Y., The green fluorescent protein. Annu Rev Biochem, 1998. 67: p. 509-44; and Chalfie, M., et al., Green fluorescent protein as a marker for gene expression. Science, 1994. 263(5148): p. 802-5, the entire teachings of which are incorporated herein by reference.
- GFPs do not require an exogenous substance or cofactor. Most applications of GFPs have focused on mapping protein localization via fusion constructs. However, GFPs are not suitable for following faster biological processes on the time-scale of minutes or less. This is due to the fact that GFPs in cellular environments have a long post-translational maturation time (Perozzo, M.A., et al., J Biol Chem, 1988. 263(16): p. 7713-6, and Heim, R., D.C. Prasher, and R.Y. Tsien, Proc Natl Acad Sci U S A, 1994. 91(26): p. 12501-4, the entire teachings of which are incorporated herein by reference), which is required for the oxidation of the three residues forming the GFP fluorophore.
- reporter gene As a reporter gene, one GFP molecule only provides for one fluorophore, thus high sensitivity detection is required for low copy numbers. Described herein is a reporter gene system that circumvents these difficulties. To illustrate this new system, ⁇ -galactosidase ("/3-gal”) is used, however, it should be obvious to those skilled in the art that other reporter genes can equally be the subject of the present invention such as jS-glucosidase.
- enzyme-substrate systems that can be employed include, but are not limited to, the following: (a) enzyme: ⁇ -galactosidase, substrates: DDAO- galactopyranoside, Resorufin- galactopyranoside; (b) enzyme: ⁇ -gluocosidase, Substrates: Resorufin-glucopyranoside, DDAO-glucopyranoside; (c) enzyme: ⁇ - lactamase, substrate: CCF2 (see, Zlokarnik et al, Science, 1998, 279(5347), 84-88, and CR2/AM (Gao et al, J.Am.Chem.
- Proteins having between 75% to 85% structural homology (and similar enzymatic activity) with the enzymes described herein are within the scope of the instant invention. Protein having between 85% to 100% structural homology (and similar enzymatic activity) with the enzymes described herein are within the scope of the present invention.
- -gal is a well-studied reporter (encoded by the lacZ gene of E. coli) and has a relatively short maturation time and fast enzymatic hydrolysis rate of fluorogenic substrates.
- DDAO-gal from Molecular Probes
- DDAO's emission maximum is at 660 nm (FIG. lb), having little overlap with autofluorescence of the cell, making it highly suitable for live cell studies. Because one copy of the enzyme ( ⁇ -g&l) generates approximately one thousand fluorescent DDAOs per second, the fluorescent signal is amplified by the enzymatic reaction, making it possible to detect low copy numbers of -gal.
- ⁇ -g&l expression is stochastic. Without inducers, a lac represser binds tightly to a DNA sequence known as the lac operator. When it occasionally falls off the operator sequence of the chromosome, one or more copies of mRNA followed by a few copies of /3-gal are produced through transcription and translation. DDAO-gal can be used to observe this stochastic event of /?-gal expression.
- the so-called N-end rule to shorten the half-life of 3-gal is used, see, Tobias, J.W., etal., Science, 1991. 254(5036): p. 1374-7, the entire teachings of which are incorporated herein by reference.
- the N-end rule states that the cellular half-life of a protein is related to its N-terminal amino acid residue. This rule applies to all organisms ranging from bacteria to mammals. In E. coli, changing the N-terminal amino acid from the natural methionine to leucine, arginine, lysine, phenylalanine, tryptophan or tyrosine shortens the protein half-life to about two minutes.
- the ubiquitin fusion technique is used to introduce a lifetime-shortening amino acid (e.g., leucine or arginine) in place of the methionine at the N-terminus of ⁇ -gal to generate Ub-Leu-/3-gal or Ub-Asg- ⁇ - gal, Seisenberger, G., et al., Science, 2001. 294(5548): p. 1929-32, the entire teaching of which is incorporated herein by reference. After this reporter protein is expressed, the ubquitin will be cleaved by an ubiquitin- specific protease, thus exposing the leucine or arginine residue and targeting the protein for the proteolytic pathways.
- a lifetime-shortening amino acid e.g., leucine or arginine
- Figure 2a depicts the ub-arg-lacZ gene in a chromosomal positioning alignment.
- Figure 2b provides the nucleotide sequence [S ⁇ Q ID NO. 1] and amino sequence [S ⁇ Q ID NO 2].
- a pair of PCR primers (5' GATG GATCCGTCGTTGCTGATTGGCGTTG 3', [S ⁇ Q ID NO. 3] and 5' GATGGATCC CGCAGGCTTCTGCTTCAATC 3', [S ⁇ Q ID NO. 4]) were used to amplify a 2000 bp fragment containing partial lacl, complete lac operon regulation region (the sequence between the end of the lacl gene and the beginning of the lacZ gene) and partial lacZ gene from the E.coli strain kl2 chromosome DNA.
- This fragment was then digested by BamHI, and ligated into a BamHI digested plasmid pBR322 (New England Biolabs) to create plasmid pBR322-IZ using standard cloning protocols Sambrook and Russell, Molecular Cloning, 3 rd Ed, CSHL press.
- Another pair of inverse PCR primers (5' CATAGCTGTT TCCTGTGTGAAATTGTTATCCGC 3', [SEQ ID NO.5] and 5' GGTGCCGGAA AGCTGGCTGGAG 3', [SEQ ID NO. 6]) was used to open this newly constructed pBR322-IZ at the 3' position of the starting codon ATG of the lacZ gene.
- a third pair of PCR primers (5' CAGATTTTCGTCAAGACTTT GACC3', [SEQ ID NO. 7] and 5' GCTTCTGGTGCCGGAAAC 3', [SEQ ID NO. 8]) were used to amplify the ubiquitin gene, the arginine residue (codon AGG) immediately after the C-terminal glycine of ubiquitin and the linker sequence between ubiquitin and lacZ from plasmid pUB23-arg (gift from Professor Daniel Finley, Harvard Medical School).
- This DNA fragment was ligated into the inverse PCR-opened pBR322-IZ and the orientation of the ubiquitin relative to the lacZ gene was verified by DNA sequencing.
- the top row in the sequence identification represents the amino acid sequence and the botton two rows represent nucleotide sequence
- the light-shaded sequence is ubiquitin (Tobias et al., Science, 1991, 254, 1374, the entire teaching of which is incorporated herein by reference)
- the unshaded sequences are the linker sequence between ubiquitin and lacZ, of which and the length and amino acids compositions are altered, and the dark-shaded sequence is the beginning of the lacZ gene without the first twenty two amino acids. It has been demonstrated that in addition to the leucine residue immediately after ubiquitin, the linker sequence has a profound impact on the cellular lifetime of -gal.
- the hydrophobicity of the amino acids composition and the length (or disordered structure) contributes greatly to the overall recognition and delivery of /3-gal to downstream proteases.
- N-terminal signal peptides derived from naturally shortlived proteins are fused to the beginning of -gal to shorten its cellular lifetime.
- the two strains kl2-n3 [S ⁇ Q ID NOS. 15, 16] and kl2-n5 [S ⁇ Q ID NOS. 17, 18] (where the top row in the sequence identification represents the amino acid sequence and the bottom two rows represent the nucleotide sequence) depicted in FIG. 3(a) belong to this group (N-terminal modification).
- the light-shaded sequences are signal peptides taken from the published work of Flynn et al, (Flynn, JM. et al, Molecular Cell, 2003, 11, 671-683, the entire teaching of which is incorporated herein by reference), and the dark-shaded sequence is the beginning of the lacZ gene without the first methionine.
- FIG 3(b) illustrates the different cellular lifetimes of these modified /3-gals expressed from E.coli chromosome, as indicated by the different DDAO-gal hydrolysis rates.
- the measurements were done using a fluorometer, in which DDAO- gal at a final concentration of 100 ⁇ M was added to E. coli cells grown to middle log phase in M9 minimal media.
- the fluorescence of the hydrolyzed product, DDAO was monitored over time at 660 nm with excitation at 638 nm.
- the hydrolysis rate was then calculated by measuring the slope of the fluorescence increase over time.
- the DDAO-gal hydrolysis rate by the wild type kl2 strain is also shown.
- E.coli As a test organism, investigators chose E.coli as a test organism.
- a plasmid encoding for ampicillin resistance gene -lactamase was transformed into E.col strains.
- the presence of the -lactamase allows the usage of the antibiotic ampicillin, which not only keeps the contamination of other bacteria minimal, but also increases the permeability of the E.coli cell wall to the fluorogenic substrate DDAO-gal.
- the mechanism of the increased cell wall permeability is very likely due to the known fact that ampicillin inhibits cell wall synthesis. All the strains described in this invention contain such an ampicillin-encoding plasmid.
- Figure 4 shows the measurements of the DDAO fluorescence signal generated by the hydrolysis of DDAO-gal in the wild type E.coli cells. The measurements were done using a fluorometer under the same conditions as described in FIG 3(b). A fluorescence signal increase can be observed immediately upon the addition of the substrate, demonstrating that DDAO-gal can permeate through cell wall and inner membrane of E. coli.
- FIG. 5 depicts both the amino acid sequence [S ⁇ Q ID NO. 19] and the DNA sequence [S ⁇ Q ID NO. 20] around the region where lacZ is deleted from chromosme.) that is primarily due to autohydrolysis.
- lacZ lacZ deficient strain
- FIG. 5 depicts both the amino acid sequence [S ⁇ Q ID NO. 19] and the DNA sequence [S ⁇ Q ID NO. 20] around the region where lacZ is deleted from chromosme.
- the microscopy experiment was performed using a through-lens total internal reflection (TIR) microscope from Olympus and an intensified CCD camera from Roper Scientific.
- TIR through-lens total internal reflection
- the excitation light was set at 638 nm, wherein autofluorescence of the E. coli cell is negligible. This detection system assures the highest sensitivity available.
- the sample chamber Bioptech
- the sample chamber was maintained at 37°C with M9 minimal medium perfusing through the chamber. E. coli cells were pushed down on the glass coverslip by a droplet of agarose gel.
- DDAO fluorescence signal from individual E. coli cells with wild type -gal (long lifetime about 10 hours) was detected. This was done at the basal level, i.e., the lacZ gene is not induced. DDAO can diffuse out or be expelled by the cell. Once it leaves the cell, DDAO quickly diffuses out from the probe volume. A steady signal was observed. In contrast, as shown in FIG. 6,
- the time trace of the fluorescence bursts exhibits quantized levels corresponding to /3-gal molecules generated and degraded one molecule at a time.
- This demonstrates the signal molecule's sensitivity of this reporting system.
- a short-lived version of a yellow fluorescent protein (YFP) variant, Venus, (Venus- ssrA) is employed. Extensive randomized and directed mutagenesis efforts have produced various GFP and YFP derivatives with faster maturation time than the wild type GFP (30-90 minutes), (Tsien, R.Y., The green fluorescent protein. Annu Rev Biochem, 1998. 67: p.
- One aspect in particular pertains to a short-lived Venus variant by creating a Venus-ssrA construct.
- the ssrA peptide tag sequence (AANDENYAKAAA, [SEQ ID NO. 21]) was encoded at the DNA level as a C- terminal fusion to Venus.
- a bacterial cell uses a ssrA sequence to flag a protein as the result of a prematurely terminated translation (see Kenneth C. Keiler, Patrick R. H. Waller, Robert T. Sauer, Science, 1996, 271, 990-993).
- Tagging Venus with ssrA tag recruits cellular protein degradation machinery and greatly reduces the cellular lifetime of Venus from more than 24 hours to less than 30 minutes. It is straightforward to extend this strategy to other GFP variants for construction of other GFP based short-lived reporter proteins.
- FIG 8(a) illustrates plasmid pVS5 which encodes the Venus-ssrA gene.
- FIG 8 (b) shows the nucleotide [SEQ ID NO. 22] and amino acid sequences [SEQ ID NO. 23] of the Venus-ssrA gene. The first amino acid shown in the figure is the first amino acid of Venus.
- a pair of PCR primers (5' CACCAGC AAGGGCGAGGAGCTGTTC-3' [SEQ ID NO. 24] and 5' TTCTTAGGCGGCTAAGG
- CGTAGTTCTCGTCGTTGGCGGCCTTGTACAGCTCGTCCATGC-3' [SEQ ID NO. 25] ) were used to amplify the Venus gene from a plasmid pCS2/venus (Nagai T, Ibata K, Park ES, Kubota M, Mikoshiba K, Miyawaki A. Nat Biotechnol. 20(1):87- 90) and add the ssrA sequence at the 3' end of the Venus gene.
- the resulting PCR fragment was then ligated into pBAD202/TOPO vector (Invitrogen Inc.) to generate plasmid pVS5.
- Short-lived /3-gal and short-lived YFP are complimentary to each other. Short-lived /3-gal can be used to detect genes that are expressed at low copy numbers because of the enzymatic amplification. Short-lived YFP provides a linear response to high-level gene expression. Real-time analysis of short-lived- YFP-incorporated cells typically work under aerobic conditions, while short-lived S-gal incorporated cells typically work under both aerobic and anaerobic conditions. The combination of the two reporter proteins will cover a broad range of intracellular gene expression levels and applicable organisms
- compositions and methods are described for live-cell microarrays.
- multiple libraries of cells each differing in at least one genotypic property are prepared.
- a live-cell microarray is comprised of two libraries.
- One library comprises cells each of which has a promoterless lacZ gene encoding for a short-lived -gal with its own ribosome binding site that is operatively linked to one promoter controlled region in the host cell's genome.
- the second library comprises the same elements except that a gene encoding for a short-lived YFP (Venus-ssrA) replaces a gene encoding for a shortlived /3-gal.
- the construction of the libraries can be accomplished by random insertion mediated by transposition or by homologous recombination. DNA sequencing around the insertion of the cells in the library will allow a practitioner to identify the position of the insertion with respect to the genome.
- a 75 x 75 element array is sufficient to contain a library with one insertion per gene for a genome has approximately 4000 genes (E.coli has about 4000 genes). (It should be noted that one skilled in the art will appreciate that various other arrays can be employed.)
- the reporter instead of inserting each reporter per gene, the reporter is operatively linked per operon. The size of the array can be smaller if only one insertion is allowed per promoter-controlled region.
- Two sets of live-cell microarrays are made from the two libraries of cells with, for example, liquid handling robots preparing the cells on a substrate such as a glass slide with a micro droplet of agarose containing growth media on top of the cells in order to immobilize the cells for ease of measurement, storage and transportation.
- microarrays Examining the microarrays under a fluorescence microscope, one can study gene expression responses to stimuli and/or environmental changes. For example, parallel movies of all elements of the microarrays can be recorded and vast amounts of data can be compiled and analyzed.
- the microarrays provide first-of-a-kind genome- wide gene expression profiling and massive kinetics data with high sensitivity and time resolution in living cells
- one cell per element of the microarray (e.g., lOO ⁇ m x lOO ⁇ m) can be effectuated.
- lOO ⁇ m x lOO ⁇ m a cell per element of the microarray
- high sensitivity makes it possible to observe the behavior of single bacterial cells in a microbial community. Not only can one detect common trends in expression profiles, a practitioner can also observe how gene expression in one cell affects its neighbors, allowing an investigator to pinpoint cooperative effects among cells.
- the background rejection advantage of confocal or total internal reflection microscopy one has both the high sensitivity to detect low-level expression events and the ability to penetrate multiple layer of biofilm.
- the present invention possesses significant sensitivity for detecting a single copy of reporter proteins in single cells, as exemplified in the Example section (see below). This allows stochastic events of gene expression of low copy number genes to be observed. Stochasticity of gene expression has attracted many experimental and theoretical efforts recently. Combined with the live-cell arrays and short-lived reporter proteins, the highly sensitive measurements of gene expression provide unprecedented information on the working of the genetic network of a genome.
- cell sorting is facilitated by the compositions described herein.
- an illuminogenic substrate such as a fluorescence substrate is introduced to a cell or population of cells, wherein the substrate enters the cells.
- a nucleotide sequence encoding a reporter protein of the instant invention is also introduced to the cells and is operatively linked within the cell's genome.
- the reporter gene i.e., the nucleotide sequence encoding for the reporter protein
- the reporter protein comprises enzymatic activity such that when it is expressed within a host cell it can facilitate the conversion of the illuminogenic substrate to an illuminesence molecule.
- a practitioner can examine various perturbations made upon the cell or cell population and determine if a particular perturbation or set of perturbations trigger the translation of a particular protein. If a particular gene, which is operatively linked to a reporter gene, is expressed upon a perturbation(s) to the cell or any of its components, then an illuminogenic signal will be emitted.
- Cells emitting a particular signal can then be separated from cells not emitting such a signal.
- conventional fluorescence cell sorters are available and can be employed in this embodiment.
- Agents used to perturb a cell can include, but not limited to, pharmaceutical agents, including test agents, pesticides, chemical agents both gaseous and in liquid form, hormones, metabolites, toxins, pheromones, and alike.
- nucleotide is used to include polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Nucleotides can have any three-dimensional structure, and can perform any function, known or unknown. The following are non-limiting examples of nucleotides: a gene or gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant nucleotides, branched nucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers.
- mRNA messenger RNA
- transfer RNA transfer RNA
- ribosomal RNA ribozymes
- cDNA recombinant nucleotides
- branched nucleotides plasmids
- vectors isolated DNA of any sequence, isolated RNA
- a nucleotide can comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the polymer.
- the sequence of nucleotides may be interrupted by non-nucleotide components.
- a nucleotide may be further modified after polymerization, such as by conjugation with a labeling component.
- the term also includes both double- and single-stranded molecules. Unless otherwise specified or required, any embodiment of this invention that is a nucleotide encompasses both the double-stranded form and each of two complementary single-stranded forms known or predicted to make up the double-stranded form.
- a nucleotide is composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); thymine (T); and uracil (U) for thymine when the polynucleotide is RNA.
- nucleotide sequence is the alphabetical representation of a nucleotide molecule. This alphabetical representation can be inputted into databases in a computer having a central processing unit and used for bioinformatics applications such as functional genomics and homology searching.
- a "gene” includes a nucleotide containing at least one open reading frame that is capable of encoding a particular polypeptide or protein after being transcribed and translated. Any of the nucleotide sequences described herein may be used to identify larger fragments or full-length coding sequences of the gene with which they are associated. Methods of isolating larger fragment sequences are known to those of skill in the art, some of which are described herein.
- a “gene product” includes an amino acid, e.g., peptide or polypeptide, generated when a gene is transcribed and then translated.
- a “primer” includes a short nucleotide, generally with a free 3 '.-OH group that binds to a target or “template” present in a sample of interest by hybridizing with the target, and thereafter promoting polymerization of a nucleotide complementary to the target.
- a “polymerase chain reaction” (“PCR”) is a reaction in which replicate copies are made of a target polynucleotide using a "pair of primers” or “set of primers” consisting of "upstream” and a “downstream” primer, and a catalyst of polymerization, such as a DNA polymerase, typically a thermally-stable polymerase enzyme.
- a primer can also be used as a probe in hybridization reactions, such as Southern or Northern blot analyses (see, for example, Sambrook, J., Fritsh, E. F., and Maniatis, T. Molecular Cloning: A Laboratory Manual. 2nd, ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989).
- cDNAs includes complementary DNA, that is mRNA molecules present in a cell or organism made into cDNA with an enzyme such as reverse transcriptase.
- a "cDNA library” includes a collection of mRNA molecules present in a cell or organism, converted into cDNA molecules with the enzyme reverse transcriptase, then inserted into "vectors” (other DNA molecules that can continue to replicate after addition of foreign DNA).
- vectors for libraries include bacteriophage, viruses that infect bacteria, e.g., ⁇ phage. The library can then be probed for the specific cDNA (and thus mRNA) of interest.
- a "delivery vehicle” includes a molecule that is capable of inserting one or more nucleotides into a host cell.
- delivery vehicles are liposomes, biocompatible polymers, including natural polymers and synthetic polymers; lipoproteins; polypeptides; polysaccharides; lipopolysaccharides; artificial viral envelopes; metal particles; and bacteria, viruses and viral vectors, such as baculo virus, adeno virus, and retro virus, bacteriophage, cosmid, plasmid, fungal vector and other recombination vehicles typically used in the art which have been described for replication and/or expression in a variety of eukaryotic and prokaryotic hosts.
- the delivery vehicles may be used for replication of the inserted nucleotide, gene therapy as well as for simply polypeptide and protein expression.
- a "vector” includes a self -replicating nucleic acid molecule that transfers an inserted polynucleotide into and/or between host cells.
- the term is intended to include vectors that function primarily for insertion of a nucleic acid molecule into a cell, replication vectors that function primarily for the replication of nucleic acid and expression vectors that function for transcription and/or translation of the DNA or RNA. Also intended are vectors that provide more than one of the above function.
- a "host cell” is intended to include any individual cell or cell culture that can be or has been a recipient for vectors or for the incorporation of exogenous nucleic acid molecules, nucleotides and/or proteins. It also is intended to include progeny of a single cell. The progeny may not necessarily be completely identical (in morphology or in genomic or total DNA complement) to the original parent cell due to natural, accidental, or deliberate mutation.
- the cells may be prokaryotic, include but are not limited to bacterial cells.
- genetically modified includes a cell containing and/or expressing a foreign gene or nucleic acid sequence that in turn modifies the genotype or phenotype of the cell or its progeny. This term includes any addition, deletion, or disruption to a cell's endogenous nucleotides.
- expression includes the process by which nucleotides are transcribed into mRNA and translated into peptides, polypeptides, or proteins. If the nucleotide is derived from genomic DNA, expression may include splicing of the mRNA, if an appropriate eukaryotic host is selected. Regulatory elements required for expression include promoter sequences to bind RNA polymerase and transcription initiation sequences for ribosome binding.
- a bacterial expression vector includes a promoter such as the lac promoter and for transcription initiation the Shine-Dalgarno sequence and the start codon AUG (Sambrook, J., Fritsh, E. F., and Maniatis, T. Molecular Cloning: A Laboratory Manual.
- a eukaryotic expression vector includes a heterologous or homologous promoter for RNA polymerase II, a downstream polyadenylation signal, the start codon AUG, and a termination codon for detachment of the ribosome.
- RNA polymerase II a heterologous or homologous promoter for RNA polymerase II
- downstream polyadenylation signal a downstream polyadenylation signal
- start codon AUG a downstream polyadenylation signal
- a termination codon for detachment of the ribosome.
- differentially expressed includes the differential production of mRNA transcribed from a gene or a protein product encoded by the gene.
- a differentially expressed gene may be overexpressed or underexpressed as compared to the expression level of a normal or control cell. In one aspect, it includes a differential that is 2.5 times, preferably 5 times or preferably 10 times higher or lower than the expression level detected in a control sample.
- the term "differentially expressed” also includes nucleotide sequences in a cell or tissue which are expressed where silent in a control cell or not expressed where expressed in a control cell.
- peptide includes a compound of two or more subunit amino acids, amino acid analogs, or peptidomimetics.
- the subunits may be linked by peptide bonds. In another embodiment, the subunit may be linked by other bonds, e.g., ester, ether, etc.
- amino acid includes either natural and/or unnatural or synthetic amino acids, including glycine and both the D or L optical isomers, and amino acid analogs and peptidomimetics.
- a peptide of three or more amino acids is commonly referred to as an oligopeptide.
- Peptide chains of greater than three or more amino acids are referred to as a polypeptide or a protein.
- Hybridization includes a reaction in which one or more nucleotides react to form a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues.
- the hydrogen bonding may occur by Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner.
- the complex may comprise two strands forming a duplex structure, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these.
- a hybridization reaction may constitute a step in a more extensive process, such as the initiation of a PCR reaction, or the enzymatic cleavage of a nucleotide by a ribozyme.
- Hybridization reactions can be performed under conditions of different "stringency.”
- the stringency of a hybridization reaction includes the difficulty with which any two nucleic acid molecules will hybridize to one another.
- stringent conditions nucleic acid molecules at least 60%, 65%, 70%, 75% identical to each other remain hybridized to each other, whereas molecules with low percent identity cannot remain hybridized.
- a preferred, non-limiting example of highly stringent hybridization conditions are hybridization in 6 X sodium chloride/sodium citrate (SSC) at about 45°C, followed by one or more washes in 0.2 X SSC, 0.1% SDS at 50°C, preferably at 55°C, more preferably at 60°C, and even more preferably at 65°C.
- a double-stranded nucleotide can be “complementary” or “homologous” to another nucleotide, if hybridization can occur between one of the strands of the first nucleotide and the second.
- “Complementarity” or “homology” is quantifiable in terms of the proportion of bases in opposing strands that are expected to hydrogen bond with each other, according to generally accepted base-pairing rules.
- nucleic acid molecule is intended to include DNA molecules, e.g., cDNA or genomic DNA, and RNA molecules, e.g., mRNA, and analogs of the DNA or RNA generated using nucleotide analogs.
- the nucleic acid molecule can be single-stranded or double-stranded, but preferably is double-stranded DNA.
- isolated nucleic acid molecule includes nucleic acid molecules, which are separated from other nucleic acid molecules that are present in the natural source of the nucleic acid.
- isolated includes nucleic acid molecules that are separated from the chromosome with which the genomic DNA is naturally associated.
- an "isolated" nucleic acid is free of sequences which naturally flank the nucleic acid (i.e., sequences located at the 5' and 3' ends of the nucleic acid) in the genomic DNA of the organism from which the nucleic acid is derived.
- the isolated marker nucleic acid molecule of the invention can contain less than about 5 kb, 4kb, 3kb, 2kb, 1 kb, 0.5 kb or 0.1 kb of nucleotide sequences which naturally flank the nucleic acid molecule in genomic DNA of the cell from which the nucleic acid is derived.
- an "isolated" nucleic acid molecule such as a cDNA molecule, can be substantially free of other cellular material, or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized.
- a nucleic acid molecule of the present invention can be isolated using standard molecular biology techniques and the sequence information provided herein. Using all or portion of the nucleic acid sequence as a hybridization probe, a molecule comprising a nucleotide sequence of the present invention can be isolated using standard hybridization and cloning techniques as described in Sambrook, L, Fritsh, E. F., and Maniatis, T. Molecular Cloning: A Laboratory Manual. 2nd, ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
- a nucleic acid of the invention can be amplified using cDNA, mRNA or alternatively, genomic DNA, as a template and appropriate nucleotide primers according to standard PCR amplification techniques.
- the nucleic acid so amplified can be cloned into an appropriate vector and characterized by DNA sequence analysis.
- nucleotides corresponding to marker nucleotide sequences, or nucleotide sequences encoding a marker of the invention can be prepared by standard synthetic techniques, e.g., using an automated DNA synthesizer.
- a nucleic acid molecule of the invention can comprise only a portion of the nucleic acid sequence of the invention, or a fragment which can be used as a probe or primer.
- the probe/primer typically comprises substantially purified nucleotide.
- Probes based on the nucleotide sequence of a nucleic acid molecule encoding a peptide of the present invention can be used to detect agglomeration proteins.
- the probe comprises a labeling group attached thereto, e.g., the labeling group can be a radioisotope, a fluorescent compound, an enzyme, or an enzyme co-factor.
- the labeling group can be a radioisotope, a fluorescent compound, an enzyme, or an enzyme co-factor.
- Such probes can be used as a part of a diagnostic test kit for identifying cells or tissue which misexpresses, e.g., over- or under-express, a polypeptide of the invention, or which have greater or fewer copies of a gene of the invention.
- hybridizes under stringent conditions is intended to describe conditions for hybridization and washing under which nucleotide sequences at least 60% homologous to each other typically remain hybridized to each other.
- the conditions are such that sequences at least about 70%, more preferably at least about 80%, even more preferably at least about 85% or 90% homologous to each other typically remain hybridized to each other.
- stringent conditions are known to those skilled in the art and can be found in Current Protocols in Molecular- Biology, John Wiley & Sons, N.Y. (1989), 6.3.1-6.3.6.
- a preferred, non-limiting example of stringent hybridization conditions are hybridization in 6 X sodium chloride/sodium citrate (SSC) at about 45°C, followed by one or more washes in 0.2 X SSC, 0.1% SDS at 50°C, preferably at 55°C, more preferably at 60°C, and even more preferably at 65°C.
- SSC sodium chloride/sodium citrate
- an isolated nucleic acid molecule of the invention that hybridizes under stringent conditions to the sequence of SEQ ID NO. 1-10.
- a "naturally-occurring" nucleic acid molecule includes an RNA or DNA molecule having a nucleotide sequence that occurs in nature, e.g., encodes a natural protein.
- the nucleotides of the invention can include other appended groups such as peptides, e.g., for targeting host cell receptors in vivo, or agents facilitating transport across the cell membrane (see, e.g., Letsinger et al. (1989) Proc. Natl. Acad. Sci. USA 86:6553-6556; Lemairre et al. (1987) Proc. Natl Acad. Sci. USA 84:648-652; PCT Publication No. W088/09810) or the blood-brain barrier (see, e.g., PCT Publication No. W0 89/10134).
- peptides e.g., for targeting host cell receptors in vivo
- agents facilitating transport across the cell membrane see, e.g., Letsinger et al. (1989) Proc. Natl. Acad. Sci. USA 86:6553-6556; Lemairre et al. (1987) Proc. Natl
- nucleotides can be modified with hybridization-triggered cleavage agents (see, Krol et al. (1988) Bio- Techniques 6:958-976) or intercalating agents (see, Zon (1988) Pharm. Res. 5:539- 549).
- the nucleotide may be conjugated to another molecule, e.g., a peptide, hybridization triggered cross-linking agent, transport agent, or hybridization- triggered cleavage agent.
- the nucleotide may be detectably labeled, either such that the label is detected by the addition of another reagent, e.g., a substrate for an enzymatic label, or is detectable immediately upon hybridization of the nucleotide, e.g., a radioactive label or a fluorescent label, e.g., a molecular beacon as described in U.S. Patent 5,876,930.
- Another aspect of the invention pertains to vectors, preferably expression vectors, containing a nucleic acid encoding a marker protein of the invention (or a portion thereof).
- the term "vector” includes a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked.
- vectors which includes a circular double stranded DNA loop into which additional DNA segments can be ligated.
- viral vector Another type of vector is a viral vector, wherein additional DNA segments can be ligated into the viral genome.
- Certain vectors are capable of autonomous replication in a host cell into which they are introduced, e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors.
- Other vectors e.g., non-episomal mammalian vectors, are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome.
- certain vectors are capable of directing the expression of genes to which they are operatively linked.
- expression vectors are referred to herein as "expression vectors.”
- expression vectors of utility in recombinant DNA techniques are often in the form of plasmids.
- plasmid and vector can be used interchangeably as the plasmid is the most commonly used form of vector.
- the recombinant expression vectors of the invention comprise a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, which means that the recombinant expression vectors include one or more regulatory sequences, selected on the basis of the host cells to be used for expression, which is operatively linked to the nucleic acid sequence to be expressed.
- "operatively linked" is intended to mean that the nucleotide sequence of interest is linked to the regulatory sequence(s) in a manner which allows for expression of the nucleotide sequence, e.g., in an in vitro transcription/translation system or in a host cell when the vector is introduced into the host cell.
- regulatory sequence is intended to include promoters, enhancers and other expression control elements, e.g., polyadenylation signals. Such regulatory sequences are described, for example, in Goeddel; Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990). Regulatory sequences include those which direct constitutive expression of a nucleotide sequence in many types of host cells and those which direct expression of the nucleotide sequence only in certain host cells, e.g., tissue-specific regulatory sequences. It will be appreciated by those skilled in the art that the design of the expression vector can depend on such factors as the choice of the host cell to be transformed, the level of expression of protein desired, and the like.
- the expression vectors of the invention can be introduced into host cells to thereby produce proteins or peptides, including fusion proteins or peptides, encoded by nucleic acids as described herein, e.g., marker proteins, mutant forms of marker proteins, fusion proteins, and the like.
- the recombinant expression vectors of the invention can be designed for expression of marker proteins in prokaryotic or eukaryotic cells.
- proteins can be expressed in bacterial cells such as E. coli, insect cells (using baculovirus expression vectors) yeast cells or mammalian cells. Suitable host cells are discussed further in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990).
- the recombinant expression vector can be transcribed and translated in vitro, for example using T7 promoter regulatory sequences and T7 polymerase.
- Fusion vectors add a number of amino acids to a protein encoded therein, usually to the amino terminus of the recombinant protein.
- Such fusion vectors typically serve three purposes: 1) to increase expression of recombinant protein; 2) to increase the solubility of the recombinant protein; and 3) to aid in the purification of the recombinant protein by acting as a ligand in affinity purification.
- a proteolytic cleavage site is introduced at the junction of the fusion moiety and the recombinant protein to enable separation of the recombinant protein from the fusion moiety subsequent to purification of the fusion protein.
- enzymes, and their cognate recognition sequences include Factor Xa, thrombin and enterokinase.
- Typical fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith, D.B. and Johnson, K.S.
- fusion proteins can be utilized in marker activity assays, e.g., direct assays or competitive assays described in detail below, or to generate antibodies specific for marker proteins for example.
- Suitable inducible non-fusion E. coli expression vectors include pTrc (Amann et al, (1988) Gene 69:301-315) and pET 1 Id (Studier et al, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, California (1990) 60-89).
- Target gene expression from the pTrc vector relies on host RNA polymerase transcription from a hybrid trp-lac fusion promoter.
- Target gene expression from the pET 1 Id vector relies on transcription from a T7 gnlO-lac fusion promoter mediated by a coexpressed viral RNA polymerase (T7 gnl). This viral polymerase is supplied by host strains BL21(DE3) or HMS174(DE3) from a resident prophage harboring a T7 gnl gene under the transcriptional control of the lacUV 5 promoter.
- One strategy to maximize recombinant protein expression in E. coli is to express the protein in a host bacteria with an impaired capacity to proteolytically cleave the recombinant protein (Gottesman, S., Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, California (1990) 119- 128).
- Another strategy is to alter the nucleic acid sequence of the nucleic acid to be inserted into an expression vector so that the individual codons for each amino acid are those preferentially utilized in E. coli (Wada et al, (1992) Nucleic Acids Res. 20:2111-2118). Such alteration of nucleic acid sequences of the invention can be carried out by standard DNA synthesis techniques.
- Another aspect of the invention pertains to host cells into which a nucleic acid ' molecule of the invention is introduced within a recombinant expression vector or a nucleic acid molecule of the invention containing sequences which allow it to homologously recombine into a specific site of the host cell's genome.
- host cell and "recombinant host cell” are used interchangeably herein. It is understood that such terms refer not only to the particular subject cell but also to the progeny or potential progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term as used herein.
- a host cell can be any prokaryotic or eukaryotic cell.
- the host cell is a prokaryotic cell.
- the invention can be expressed in bacterial cells such as E. coli.
- Other suitable host cells are known to those skilled in the art.
- Vector DNA can be introduced into host cells via conventional transformation or transfection techniques.
- a host cell of the invention such as a host cell in culture, can be used to produce, i.e., express, a recombinant protein.
- the invention further provides methods for producing a protein using the host cells of the invention.
- the method comprises culturing the host cell of invention (into which a recombinant expression vector encoding a protein, or proteins, has been introduced) in a suitable medium such that a protein of the invention is produced.
- the method further comprises isolating a protein from the medium or the host cell.
- Example 1 Detection of transient gene expression in single living E.coli cells with sensitivity for one protein molecule
- N-end rule states that the cellular lifetime of a protein is related to its N-terminal amino acid residue.
- changing N-terminal amino acid from the natural methionine to leucine, arginine, lysine, phenylalanine, tryptophan or tyrosine shortens the protein's half -life to a few minutes.
- the ubiquitin fusion technique was used in order to introduce a lifetime- shortening amino acid (e.g., leucine or arginine) replacing the methionine at the N-terminus of ⁇ -gal to generate Ub-Leu- ⁇ -gal or Ub-Arg- ⁇ -gal , see, Bachmair, A., D. Finley, and A. Varshavsky, Science, 1986. 234(4773): p. 179- 86, the entire teaching of which is incorporated herein by reference.
- the ubiquitin is cleaved by ubiquitin- specific protease, thus the argine or the leucine residue is exposed to the proteolytic pathways in E.coli.
- ⁇ -gal is but one system to demonstrate the proof of principle of short-lived reporter proteins.
- This same strategy described herein can be used with other reporter genes in order to track transient behavior.
- These reporter genes can be used to make fusion proteins for multiplexing observation of gene expression processes.
- One will be able to study gene regulatory circuits by examining the effects of one gene on another.
- Such work will offer detailed information about the interactions and regulation among gene products.
- another reporter ⁇ -glucosidase with a molecular mass of 82 kDa, encoded by the gene bglB from Bacillus sp. GL1 (Arch. Ciochem. Biophys., vol 360, No. 1, pp 1-9, 1998) is employed.
- This enzyme hydrolyzes the non-reducing terminal glucoside from either carbon hydrates or artificial substrates such as resorufin-glucopyranoside (see FIG. 1 (c) and (d) for the substrate structure and product spectrum).
- resorufin-glucopyranoside see FIG. 1 (c) and (d) for the substrate structure and product spectrum.
- the strain that express the ⁇ -glucosidase gene showed very high hydrolysis activity on resorufin-glucopyranoside, while the wild type E. Coli (does not contain the gene encoding for ⁇ -glucosidase) showed negligible glucosidase activity (FIG. 9).
- ⁇ -lactamase which hydrolyzes fluoregenic substrates such as CCF2 and CR2/AM (see, Zlokarnik et al, Science, 1998, 279(5347), 84-88, Gao et al, J.Am.Chem. Soc, 2003, 125, 11146-11147, the entire teachings of which are incorporated herein by reference.), can also be genetically modified and employed in the reporting system.
- a DNA cassette including a promoter-less ub-x-lacZ gene (the x between ub and lacZ represents any amino acid that shortens the cellular lifetime of the resulting ⁇ -gal) will be cloned into a transposon construction vector pMOD-2 (Epicentre Technologies), flanked by two Tn5 -recognizable 19 bp ME sequence.
- This ub-x-lacZ gene contains its own ribosome binding site (RBS), in front of which a stop codon will be placed to avoid a translation read-through from a previous gene. See, FIG. 10.
- FIG 10 is a schematic drawing of the construction of the lacZ library by Tn5 mediated transposition.
- ME represents Tn5 recognizable mosaic ends sequence (triangles); RBS are the ribosome binding sites (rectangles); and the box joined by a hitched box indicates the ub-x-lacZ gene and the oval with a turn arrow on top indicates a promoter on the chromosome.
- the selection for desired colonies containing the reporter genes will be based on blue/white colony screening on X-gal plates.
- the expression of the promoter-less Ub-X-lacZ gene from a functional promoter on the chromosome will result in blue colonies due to the conversion of X-gal into blue insoluble precipitant by ⁇ -gal. Since the conversion of X-gal by ⁇ -gal is highly efficient and can accumulate, even colonies transiently expressing ⁇ -gal can be identified if sufficient growing time is allowed. Investigators have observed that the E. coli colonies contain one single copy or less of short-lived ⁇ -gal on average produce easily visible blue color after 16 hours incubation.
- Figure 11 depicts an alternative method for constructing the lacZ and YFP libraries simultaneously.
- the last selection step generates both lacZ and yfp libraries based on blue and white colonies screening. (The notations are the same as FIG. 10.)
- a methylated DNA cassette will be randomly inserted into E.coli genome by Tn5 mediated in vitro transposition as described above.
- This DNA cassette will contain a copy of ub-x-lacZ (contains a stop codon and its own ribosome binding site in front), and also a copy of Venus-ssrA with its 3' end flanked by approximately 500 bp sequence, which is homologous to the 3' end of the lacZ gene.
- ub-x-lacZ contains a stop codon and its own ribosome binding site in front
- Venus-ssrA with its 3' end flanked by approximately 500 bp sequence
- the first round selection for the incorporation of this DNA cassette into the chromosome will be based on the ⁇ -galactosidase activity on the X-gal plate or chloramphenicol resistance.
- the colonies from the first round of selection will be pooled and plated on sucrose plates supplied with X-gal. Blue colonies that survived on the sucrose plates indicate the presence of the ub-x-lacZ gene on the chromosome, thus forming the lacZ library, while the white colonies indicate the presence of Venus-ssrA, forming the YFP library. Both libraries will then be replicated on chloramphenicol plates to ensure that the survival on sucrose plates is not due to the mutation of the sacB gene (see, Link, A.J., D. Phillips, and G.M. Church).
- Abundant single-stranded D ⁇ A will be first generated by one primer, which specifically targets one end of the ub-x-lacZ gene and goes outward relative to the transposon D ⁇ A.
- these ssD ⁇ As will be amplified by random priming at low annealing temperature using the same primer to produce double- stranded D ⁇ A (dsD ⁇ A) with different lengths.
- dsD ⁇ A double- stranded D ⁇ A
- dsD ⁇ As double- stranded D ⁇ A
- dsD ⁇ As double- stranded D ⁇ A
- a new primer which targets specifically a sequence lying in the middle of the ME sequence and the first primer binding sequence on the transposon, will be used to sequence the amplified dsD ⁇ A. The sequence will then be compared to E.coli genome to identify the position of insertion.
- investigators will select from the initial library based on the sequence data according to the following criteria: 1) least disruption of a gene; 2) least polar effect to the downstream genes caused by the insertion; and 3) at least one insertion for one promoter or operon.
- the final library is estimated to contain at least 3000 strains including both unique and multiple reporter gene insertions for each predicted promoter.
- Nanolitres of aqueous media containing E.coli cells will be pipetted onto a #1 glass coverslip using a robotic micro-arrayer; the Omnigrid (GeneMachine) arrayer is capable of dispensing a minimum of 300 picolitre of fluid.
- Sub-microlitres of low-melting-temperature agarose containing growth media will be immediately applied on top of the cell solutions to prevent drying of cells.
- the weight of the agarose will compress the media droplets and create a monolayer of cells on the surface of the cover glass.
- the resulting spot will be approximately hundreds of micrometers in diameter, which is about the size of the view-field on a microscope. Macroscopic version of this technique has been consistently demonstrated, and cells stay viable and divide for many generations on the slide.
- the array Once the array is printed, it will be capped by a gasket and Microaqueduct slide manufactured by Bioptech Inc, as shown in FIG. 12.
- the microaqueduct slide will allow laminar flow through the chamber and keep the temperature constant via an add-on thermoelectric heater unit.
- a solution of DDAO-gal and growth media can be perfused through the chamber to keep the cells supplied with nutrients and fluorogenic substrates, while allowing fluorescent product DDAO and cellular metabolites to flow away.
- this set-up whereby media is allowed to flow over microbes that are fixed in place, is very similar to those routinely used for analysis of biofilm formation, and is therefore, more representative of how bacteria exist in the environment.
- the microarray encased in a flow chamber offers a versatile and durable platform with a controlled environment and constant supply of nutrients.
- This chamber will be mounted on TIRFM microscope (Nikon Te-2000E) with a built-in motorized XY stage. Equipped with rotary encoders and feedback stepper motors, the XY stage can visit each micro-colony on the microarray with a repeatability of one micron.
- the objective lens can auto-focus before acquiring an image at each spot.
- Shutters and filter wheels controlled by commercial software can precisely time illumination with laser and Xe lamp, to acquire fluorescent and phase-contrast images. These images can be stacked into movies for each point in the microarray. Fluorescent time trajectories will be extracted for each individual cell and proper statistics analysis will be performed.
- High-throughput real-time data provides quantitative information on system- wide gene expression kinetics. This first-of-a-kind systems biology dataset provides an opportunity for mathematical modeling.
- Figure 14a shows the fluorescence image of the individual transformed cells supplied with DDAO-gal without induction. In contrast, under the same condition, no fluorescence signal was observed in the wild-type strain, see, FIG. 14b. This experiment proves that DDAO-gal can permeate through the Shewanella oneideinsis cell membrane and the fluorescence signal is specifically due to the presence of ⁇ -gal.
- the N-end rule has been demonstrated to be universal in organisms examined such as E. coli, yeast and mammals, (see, Varshavsky, A., The N-end rule: functions, mystery, uses. Proc Natl Acad Sci U S A, 1996. 93(22): p. 12142-9, the entire teaching of which is incorporated herein by reference).
- Shewanella is closely related to E. coli, therefore, it is reasonable to assume that the same rule also applies in Shewanella.
- FIG. 14b when a short-lived ⁇ -gal (Ub-Leu- ⁇ -gal) is expressed together with the ubiquitin-specific protease, the hydrolysis rate decrease dramatically, indicating shortened cellular lifetime of ⁇ -gal.
- Example 3 ⁇ -gal applied to Saccharomyce cerevisiae
- Saccharomyce cerevisiae (budding yeast) to probe stochastic gene expression events. Saccharyomyce cerevisiae has extensive ubiquitin-dependent protein degradation pathways, thereby enabling a cellular lifetime of modified ⁇ -gal less than a few minutes.
- the ub-leu-lacZ reporter gene was generated using standard cloning protocols (Sambrook and Russell, Molecular Cloning, 3 rd Ed, CSHL press, the entire teaching of which is incorporated herein by reference) with a pair of PCR primers (5' CTTGGTA CCATGCAGATTTTCGTCAAGACTTTG 3' [SEQ ID NO. 27], and 5' GAGCGGC CGCTTTTGACACCAGACC 3' [SEQ ID NO. 28]) to amplify a 4000bp fragment containing ub-leu-lacZ from pUB23 plasmid generated by Varshavsky, et al. This DNA fragment was ligated into the pYC2/CT plasmid (Invitrogen, Inc.) and the resulting construct was verified by DNA sequencing.
- Figure 15a depicts the nucleotide sequence junction of ub-leu-lacZ, and (b) is the amino acid sequence [SEQ ID NO. 29] and nucleic acid sequence [SEQ ID NO. 30] for the junction of the ub-leu-lacZ construct on centromeric plasmid: the sequence is from the Gall promoter site to the Bsu26I site of the lacZ gene, and numbering of the nucleotides is according to the first base of the Gall promoter; the ubiquitin gene is joined by an modified Z cZ gene with its first methionine residue replaced by a leucine residue. The amino acids sequences are shown on top of the DNA sequence panel.
- Figure 16 shows DDAO fluorescence generated from the hydrolysis of DDAO-gal by lacZ7 (dark) but not by the lacZ (light) yeast cells measured in a fluorometer.
- Final concentration of DDAO-gal was 50 ⁇ M and S. cerevisiae cells was grown to middle log phase in synthetic dextrose medium.
- the significantly different hydrolysis rates between the two strains demonstrated that (i) fluorescence substrate DDAO-gal is permeable to S. cerevisiae cell wall and plasma membrane; (ii) DDAO- gal is hydrolyzed by ⁇ -gal with remarkable specificity and high turnover rate.
- Figure 17 is a fluorescence image of S. cerevisiae cells expressing wild type ⁇ -gal. This experiment was done using fluorescence microscope with excitation at 568nm. The other setup is identical to those used in the E.coli and Shewanella experiments as described above. The presence of glucose in the growth media represses the Gall promoter, resulting in a low basal level expression of ⁇ -gal. Experimental conditions were chosen to minimize the background and autofluorescence of yeast cells.
- Figure 18 shows the fluorescence burst observed on a single S. cerevisiae cell with a short-lived ⁇ -gal expressed from the centromeric plasmid.
- the burst in the time trace indicates a single lacZ gene expression event, resulting from the stochastic dissociation of the repressor from its binding site.
- the rise of the burst indicates the generation of ⁇ -gal and the decay indicates the degradation of ⁇ -gal.
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Engineering & Computer Science (AREA)
- Immunology (AREA)
- Organic Chemistry (AREA)
- Molecular Biology (AREA)
- Zoology (AREA)
- Hematology (AREA)
- Microbiology (AREA)
- Wood Science & Technology (AREA)
- Urology & Nephrology (AREA)
- Biotechnology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Biomedical Technology (AREA)
- Physics & Mathematics (AREA)
- Analytical Chemistry (AREA)
- Biochemistry (AREA)
- General Health & Medical Sciences (AREA)
- Medicinal Chemistry (AREA)
- Pathology (AREA)
- Biophysics (AREA)
- General Physics & Mathematics (AREA)
- Food Science & Technology (AREA)
- Cell Biology (AREA)
- Tropical Medicine & Parasitology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
Abstract
Disclosed herein are compositions and methods for detecting and analyzing both transcriptional and translational events occurring in live cells. In particular, short-lived reporters with enzymatic amplification are described. These reporters have relatively short maturation time and a short cellular lifetime which can be exploited to detect transient events of gene expression in live cells.
Description
IN THE UNITED STATES PATENT AND TRADEMARK OFFICE A NON-PROVISIONAL PATENT APPLICATION
FOR
DETECTING GENE EXPRESSION IN LIYE CELLS USING SHORT-LIYEB REPORTERS WITH ENZYMATIC AMPLIFICATION
RELATED APPLICATIONS
The present application claims priority to and the benefit of US provisional application serial number 60/459,897, filed April 2, 2003.
FIELD OF THE INVENTION
The present invention relates to compositions and methods for detecting and analyzing gene expression events occurring in live cells.
BACKGROUND
One of the major challenges in the post-genomic era is to understand how genes are expressed and regulated. Gene expression can be tracked at the mRNA and protein level. Despite considerable progress in transcription and translational profiling with micorarray and mass spectrometry, methods that continuously monitor gene expression dynamics in live cells are in high demand. In addition, current microarray and mass spectrometry technologies cannot detect low copy number gene products, which often play a prominent role in sensing, signaling and gene regulation.
In recent years, tremendous progress has been made in the area of single- molecule' detection in biological systems. It is fair to say that the single-molecule approach has changed the way many biological problems are addressed and interpreted. New insights derived from this approach are continuing to emerge. Although most of the single-molecule work has been carried out in vitro, single molecule experiments in living cells are beginning to appear. Indeed, gene expression in a single cell is a single molecule problem. In addition, the low copy numbers of
mRNA and proteins exhibit stochastic fluctuations similar to those seen in single- molecule experiments.
The use of reporter proteins have been employed to detect events of gene expression in cells. Typically, green fluorescent protein (GFP) and its derivatives are used as reporter proteins. The main advantage of GFPs is that they do not require exogenous substrate or cofactor. Most applications of GFPs have focused on mapping protein localization via fusion constructs. However, current GFPs are not suitable for following fast biological processes on the time-scale of minutes or less. This is primarily due to the fact that GFPs in the cellular environment have a long post-translational maturation time, which is required for the oxidation of the three residues forming the GFP fluorophore. Thus, a new GFP variant with faster maturation time is needed. However, even with such a variant, one GFP molecule only provides one fluorophore, thus it is only suitable for detecting translational product that expressed at high levels.
Currently, a need exists for a new reporting system that allows real-time detection of low copy number translational products in individual live cells. Moreover, there is a concomitant need for such a reporter system that employs compositions, which shorten the cellular lifetime of the reporter protein, thus allowing for following real-time biological processes while obtaining background-free measurements with high sensitivity.
SUMMARY OF THE INVENTION
The present invention pertains to compositions and methods for detecting and analyzing gene expression events occurring in live cells. In one aspect, the present invention pertains to a short-lived reporter with enzymatic amplification. The reporters of the present invention have relatively short maturation time and a short cellular lifetime which can be exploited to detect transient events of gene expression in live cells.
In one embodiment of the present invention, compositions and methods for employing one or more reporters having a short maturation time and a short cellular
lifetime to detect transient events of gene expression in individual living cells with high sensitivity and high time resolution are described. Also, described herein is a reporter gene system employing a reporter, for example, /J-galactosidase ( J-gal). In one aspect of this embodiment, the reporter is manipulated in such a manner so as to decrease its cellular lifetime.
In this aspect, the so-called N-end rule to shorten the cellular lifetime of -gal is utilized. The N-end rule states that the cellular lifetime of a protein is related to its N-terminal amino acid residue. This rule applies to all organisms ranging from bacteria to mammals. In E. coli, changing the N-terminal amino acid from the natural methionine to leucine, arginine, lysine, phenylalanine, tryptophan or tyrosine shortens the protein half-life to a few minutes. Since all newly translated proteins have methionine at the N- terminus (the translation start codon encodes for methionine), the ubiquitin (ub) fusion technique is used to introduce a lifetime-shortening amino acid (e.g., leucine or arginine) in place of the methionine at the N-terminus of, for example, /J-gal to generate Ub-Leu- J-gal or Ub-Arg- -gal. After this reporter protein is expressed, the ubiquitin will be cleaved by an ubiquitin-specific protease, thus exposing the leucine or arginine residue and targeting the protein for the proteolytic pathways. In addition to the N-end rule, other means of modifying /3-gal's cellular lifetime are also employed, such as N-terminal and C-terminal signal peptides fusions.
In another embodiment, live-cell microarrays are described. In this embodiment, multiple libraries of cells are prepared each differing in at least one genotypic property (i.e., the genotype of each cell is different, for example, the reporter gene is inserted at a different position on the chromosome, thereby tagging an operon or a gene). In one aspect, a live-cell microarray is comprised of two libraries. One library comprises cells each of which has a promoterless lacZ gene encoding for a short-lived -gal with its own ribosome binding site that is inserted into one promoter controlled region in the host cell's genome. The second library comprises the same elements except that a gene encoding for a short-lived yellow fluorescent protein YFP (Venus-ssrA) replaces the gene encoding for a short-lived /3-gal in the first library.
BRIEF DESCRIPTION OF THE DRAWINGS
FIG. 1 (a) are chemical structures of 9H-(l,3-dichloro-9,9-dimethylacridin-2- one-7-yl) 3-D-galactopyranoside (DDAO-gal) and its fluorescent product DDAO after hydrolysis by /3-galactosidase, and (b) is a graphical representation of the absorption and emission spectra of DDAO; (c) are chemical structures of resorufin- glucopyranoside (resorufin-glu) and its fluorescent product resorufin after hydrolysis by -glucosidase and (d) is a graphical representation of the absorption and emission spectra of resorufin;
FIG. 2 (a) shows the location of the gene coding for Ub-Arg- -gal in the lac operon and (b) depicts the nucleic acids and amino acids sequences of Ub-Arg- -gal. Only the sequences of ubiquitin (light-shaded), the arginine residue immediately after ubiquitin, and the linker peptide (unshaded) between ubiquitin and the beginning of -gal (dark-shaded) are shown. Please note that the -gal in this construct lacks its first twenty two amino acids;
FIG. 3 (a) is a graph measuring the hydrolysis of DDAO-gal in the presence of enzyme products from different gene constructs, and (b) are the amino acid (top) and nucleotide sequence (bottom) for each of the different construct; Please note that only the sequences that differ in these constructs (N-terminus of the lacZ gene) are shown in (b);
FIG. 4 is a graph showing the DDAO fluorescence generated from the hydrolysis of DDAO-gal by wild type lac77 cells (dark) but not by the lacZ cells (light);
FIG. 5 (a) depicts the sequence junction of lacZ deletion, wherein the sequence is from the EcorV site of the lacl gene to Nspl site of the lacY gene, and (b) is the amino acid sequence and nucleic acid sequence wherein the numbering of the nucleotides is according to the first base of the lacl gene, the lacZ gene is replaced by
lacY gene from the ATG starting codon, the amino acids sequences are shown on top of the DNA sequence panel;
FIG. 6 is a fluorescence image of E. coli Cells. The signal is from DDAO generated by the basal level expression of unmodified /3-gal;
FIG. 7 (a) is the fluorescence images observed on single E.coli cells with a gene coding for a short-lived Ub-Arg-jβ-gal incorporated on chromosome. The signal is generated by the basal level expression of β-gal, For cell 1, only thirteen fluorescence images of cell 1 are shown in fifteen minute intervals for simplicity reasons. For cell 2, the fluorescence images are shown in five minute intervals, (b) is a fluorescence measruement of the production and degradation of /3-gal in one singe E.coli cell under TIR fluorescence microscope;
FIG. 8 (a) depicts the sequence for the short-lived YFP: Venus-ssrA construct on plasmid pVS5, and (b) is the amino acid sequence and nucleic acid sequence wherein the sequence is from the first base of the yfp gene and to the end of the yfp gene with the addition of 33 bases coding for the ssrA tag;
FIG. 9 is a graph showing the resorufin fluorescence generated from the hydrolysis of resorufin-glu by E.coli cells expressing -glucosidase ( bgl +, light) but not by the bglB' cells (dark);
FIG. 10 is a schematic drawing of the construction of a lacZ library by Tn5 mediated transposition;
FIG. 11 is a schematic drawing of the constructing a lacZ and yfp library;
FIG. 12 is a flow chart showing an automated process for the construction of libraries and the fabrication of the cell array;
FIG. 13(a) is the plasmid map for pBBRlMCS-5.1, and (b) is the nucleotide sequence coding for the short-lived b-gal for the plasmid depicted in (a). Please note that only the sequence at the N-terminus of the ub-leu-lacZ gene is shown;
FIG. 14 (a) is a fluroescence image of Shewanella oneideinis cells expressing /3-gal from the lacZ7 plasmid pBBR!MCS5.1, and (b) is a graph showing the DDAO fluorescence generated by the hydrolysis of DDAO-gal under various conditions;
FIG. 15 (a) depicts the nucleotide sequence junction of ub-leu-lacZ gene in Saccharomyce cerevisiae and (b) is the amino acid sequence and nucleic acid sequence for the junction of the ub-leu-lacZ construct on centromeric plasmid transformed into Saccharomyce cerevisiae cell;
FIG. 16 represents DDAO fluorescence generated from the hydrolysis of DDAO-gal by wild type /αcZ+ cells (dark) but not by the lacZ' cells (light) in Saccharomyce cerevisiae;
FIG. 17 is a fluorescence image of S. cerevisiae cells containing unmodified /3-gal; and
FIG. 18 is the fluorescence signal bursts observed on a single S. cerevisiae cell with a short-lived J-gal expressed from a centromeric plasmid.
DETAILED DESCRIPTION
The present invention pertains to compositions and methods for detecting and analyzing gene expression events occurring in individual living cells. In particular, the present invention pertains to short-lived reporters with enzymatic amplification. These reporters of the present invention have relatively short maturation time and a short cellular lifetimes which can be exploited to detect transient events of gene expression in live cells.
Tremendous progress has been made to track gene expression at the mRNA level by DNA arrays and at the protein level by mass spectrometry. Although current
DNA microarray and mass spectrometry technologies have started to address compelling biological problems at a genome-wide scale, they suffer from a few disadvantages: (1) they cannot continuously monitor temporal evolution of expression - multiple samples have to be taken in order to evaluate the response to a stimulus or an environmental change; (2) they cannot follow fast gene expression processes on the time scale of minutes. The low time resolutions prevent studies of transient gene expression processes, for example, those involved in cell division; (3) they are not sensitive to low copy number gene products, which often play a prominent role in cellular sensing, signaling and gene regulation; and (4) they can only provide averaged results of large populations of cells rather than behaviors of individual cells: transient and stochastic gene expression events are often masked in the population measurements.
In one embodiment of the present invention, a method for employing one or moire reporters having a short maturation time and a short cellular lifetime to detect transient events of gene expression in live cells with high sensitivity and a fast time resolution is described.
In one aspect, a reporting system for monitoring real-time gene expression events in a living cell is disclosed. This reporting system comprises an illuminogenic substrate, wherein said substrate is permeable to said cell. The system also comprises at least one reporter protein, wherein said reporter protein facilitates the conversion of said illuminogenic substrate into an illuminescent molecule, and wherein said reporter protein has a short cellular life time.
In this aspect, the cell can be a prokaryote or eurokaryote. The illuminogenic substrate can be any substrate that when acted upon by, for example, hydrolysis, will generate an illuminescent product which emits photons. For example, the substrate can be a fluorogenic substrate that when acted upon will generate a fluorescent product that emits fluorescence. Chemiluminescence substrates can also be used in the present invention. The term illuminogenic is also meant to cover absorption in addition to photon emission, for example, chromogenic substrates.
The present embodiment is designed to capitalize on the recent advances in sensitive fluorescence microscopy. In the past years, tremendous progress has been made in fluorescence imaging of single-molecules, even in living cells. See, for example, Sako, Y. and T. Uyemura, Total Internal Reflection Fluorescence Microscopy for Single-molecule Imaging in Living Cells. Cell Struct Funct, 2002. 27(5): p. 357-65; Sako, Y., S. Minoghchi, and T. Yanagida, Single-molecule imaging of EGFR signalling on the surface of living cells. Nat Cell Biol, 2000. 2(3): p. 168- 72; Seisenberger, G., et al., Real-time single-molecule imaging of the infection pathway of an adeno-associated virus. Science, 2001. 294(5548): p. 1929-32; and the entire teachings of which are incorporated herein by reference. State-of-the-art microscopes are more than capable of imaging single or multiple numbers of gene products of a single gene, if not single fluorophores in a live cell.
A popular approach for real-time observation of gene expression in live cells is the use of green fluorescent protein (GFP) and derivatives thereof as reporter proteins. See, for example, Bongaerts, R.J., et al., Green fluorescent protein as a marker for conditional gene expression in bacterial cells. Methods Enzymol, 2002. 358: p. 43-66; Tsien, R.Y., The green fluorescent protein. Annu Rev Biochem, 1998. 67: p. 509-44; and Chalfie, M., et al., Green fluorescent protein as a marker for gene expression. Science, 1994. 263(5148): p. 802-5, the entire teachings of which are incorporated herein by reference. The main advantage of GFPs is that they do not require an exogenous substance or cofactor. Most applications of GFPs have focused on mapping protein localization via fusion constructs. However, GFPs are not suitable for following faster biological processes on the time-scale of minutes or less. This is due to the fact that GFPs in cellular environments have a long post-translational maturation time (Perozzo, M.A., et al., J Biol Chem, 1988. 263(16): p. 7713-6, and Heim, R., D.C. Prasher, and R.Y. Tsien, Proc Natl Acad Sci U S A, 1994. 91(26): p. 12501-4, the entire teachings of which are incorporated herein by reference), which is required for the oxidation of the three residues forming the GFP fluorophore.
As a reporter gene, one GFP molecule only provides for one fluorophore, thus high sensitivity detection is required for low copy numbers. Described herein is a
reporter gene system that circumvents these difficulties. To illustrate this new system, β-galactosidase ("/3-gal") is used, however, it should be obvious to those skilled in the art that other reporter genes can equally be the subject of the present invention such as jS-glucosidase.
Other enzyme-substrate systems that can be employed include, but are not limited to, the following: (a) enzyme: β -galactosidase, substrates: DDAO- galactopyranoside, Resorufin- galactopyranoside; (b) enzyme: β -gluocosidase, Substrates: Resorufin-glucopyranoside, DDAO-glucopyranoside; (c) enzyme: β- lactamase, substrate: CCF2 (see, Zlokarnik et al, Science, 1998, 279(5347), 84-88, and CR2/AM (Gao et al, J.Am.Chem. Soc, 2003, 125, 11146-11147, the entire teachings of which are incorporated herein by reference.) It should be understood that other enzyme activities similar to those just listed are also encompassed within the present invention. Additionally, modified proteins having identical or similar enzymatic activities are also encompassed within the present invention. For example, proteins that have between 45% to 65% structural homology (and similar enzymatic activity) with the enzymes described herein are within the scope of the invention. (Unless otherwise stated, the terms protein and peptide can be used interchangeably herein.) Proteins having between 65% to 75% structural homology (and similar enzymatic activity) with the enzymes mentioned above are within the scope of the invention. Proteins having between 75% to 85% structural homology (and similar enzymatic activity) with the enzymes described herein are within the scope of the instant invention. Protein having between 85% to 100% structural homology (and similar enzymatic activity) with the enzymes described herein are within the scope of the present invention.
-gal is a well-studied reporter (encoded by the lacZ gene of E. coli) and has a relatively short maturation time and fast enzymatic hydrolysis rate of fluorogenic substrates. DDAO-gal (from Molecular Probes) is a good fluorogenic substrate for assaying /3-gal activity in vivo (FIG. la). DDAO's emission maximum is at 660 nm (FIG. lb), having little overlap with autofluorescence of the cell, making it highly suitable for live cell studies. Because one copy of the enzyme (β-g&l) generates approximately one thousand fluorescent DDAOs per second, the fluorescent signal is
amplified by the enzymatic reaction, making it possible to detect low copy numbers of -gal. Without induction, there are about 10 -gal per E.coli cell (Sambrook, J. and D. Russell, Molecular Cloning. 3rd ed. Vol. 3. 2001, Cold Spring Harbor, New York: Cold Spring Harbor Laboratory Press. 15.57, the entire teaching of which is incorporated herein by reference), providing a good model system for detecting genes that expressed at low copy numbers. (It should be noted that other substrates can also be employed such as resorufin-glu, whose hydrolyzed product resorufin has am maximum absorption at 571 nm, and emission at 585 nm.)
On the chromosome of E.coli, β-g&l expression is stochastic. Without inducers, a lac represser binds tightly to a DNA sequence known as the lac operator. When it occasionally falls off the operator sequence of the chromosome, one or more copies of mRNA followed by a few copies of /3-gal are produced through transcription and translation. DDAO-gal can be used to observe this stochastic event of /?-gal expression.
However, in E. coli, the lifetime of -gal is longer than 10 hours, Tobias, J.W., etal, Science, 1991. 254(5036): p. 1374-7, the entire teaching of which is incorporated herein by reference. This presents a general obstacle to follow dynamic processes. A long-lived reporter protein leaves a constant background that prevents the detection of small and transient variations. A solution to this problem, which is the subject of this invention, is to shorten the cellular lifetime of reporter proteins. The reporter proteins are degraded shortly after they are expressed, generating a background free condition for sensitive detection. This provides a general approach for visualizing individual gene expression events in realtime, as these events can be observed as discrete fluorescence bursts.
In one aspect of the invention, the so-called N-end rule to shorten the half-life of 3-gal is used, see, Tobias, J.W., etal., Science, 1991. 254(5036): p. 1374-7, the entire teachings of which are incorporated herein by reference. The N-end rule states that the cellular half-life of a protein is related to its N-terminal amino acid residue. This rule applies to all organisms ranging from bacteria to mammals. In E. coli, changing the N-terminal amino acid from the natural methionine to leucine, arginine,
lysine, phenylalanine, tryptophan or tyrosine shortens the protein half-life to about two minutes. Since all newly translated proteins have methionine at the N-terminus (the translation start codon encodes for methionine), the ubiquitin fusion technique is used to introduce a lifetime-shortening amino acid (e.g., leucine or arginine) in place of the methionine at the N-terminus of β-gal to generate Ub-Leu-/3-gal or Ub-Asg-β- gal, Seisenberger, G., et al., Science, 2001. 294(5548): p. 1929-32, the entire teaching of which is incorporated herein by reference. After this reporter protein is expressed, the ubquitin will be cleaved by an ubiquitin- specific protease, thus exposing the leucine or arginine residue and targeting the protein for the proteolytic pathways.
Using an ub-arg-lacZ reporter gene (coding for Ub-Arg- -gal) on the chromosome of E.coli, it has been demonstrated that an in vivo half-life of about two minutes in E. coli can be obtained. Figure 2a depicts the ub-arg-lacZ gene in a chromosomal positioning alignment. Figure 2b provides the nucleotide sequence [SΕQ ID NO. 1] and amino sequence [SΕQ ID NO 2].
To generate this ub-arg-lacZ reporter gene, a pair of PCR primers (5' GATG GATCCGTCGTTGCTGATTGGCGTTG 3', [SΕQ ID NO. 3] and 5' GATGGATCC CGCAGGCTTCTGCTTCAATC 3', [SΕQ ID NO. 4]) were used to amplify a 2000 bp fragment containing partial lacl, complete lac operon regulation region (the sequence between the end of the lacl gene and the beginning of the lacZ gene) and partial lacZ gene from the E.coli strain kl2 chromosome DNA. This fragment was then digested by BamHI, and ligated into a BamHI digested plasmid pBR322 (New England Biolabs) to create plasmid pBR322-IZ using standard cloning protocols Sambrook and Russell, Molecular Cloning, 3rd Ed, CSHL press. Another pair of inverse PCR primers (5' CATAGCTGTT TCCTGTGTGAAATTGTTATCCGC 3', [SEQ ID NO.5] and 5' GGTGCCGGAA AGCTGGCTGGAG 3', [SEQ ID NO. 6]) was used to open this newly constructed pBR322-IZ at the 3' position of the starting codon ATG of the lacZ gene. A third pair of PCR primers (5' CAGATTTTCGTCAAGACTTT GACC3', [SEQ ID NO. 7] and 5' GCTTCTGGTGCCGGAAAC 3', [SEQ ID NO. 8]) were used to amplify the ubiquitin gene, the arginine residue (codon AGG) immediately after the C-terminal glycine of ubiquitin and the linker sequence between ubiquitin and lacZ from plasmid
pUB23-arg (gift from Professor Daniel Finley, Harvard Medical School). This DNA fragment was ligated into the inverse PCR-opened pBR322-IZ and the orientation of the ubiquitin relative to the lacZ gene was verified by DNA sequencing. Next, the replacement of the wild type lacZ gene on the E. coli chromosome was achieved by homologous recombination using a gene replacement vector pKO3, see, Link, et ah, J Bacteriol, 1997. 179(20): p. 6228-37, the entire teaching of which is incorporated herein by reference. The final resulting construct on the chromosome is depicted in FIG. 2.
In addition to this Ub-Arg - S-gal construct, a repertoire of short-lived jS-gals with different cellular lifetimes were constructed. In one group (N-end rule), the linker sequence lying between ubiquitin and lacZ, referred to as "eK" sequence, was varied. See FIG. 3(a) kl2-el [SΕQ ID NOS. 9, 10], kl2-e2 [SΕQ ID NOS. 11, 12] and kl2-e3a [SΕQ ID NOS. 13, 14] (where the top row in the sequence identification represents the amino acid sequence and the botton two rows represent nucleotide sequence), where the light-shaded sequence is ubiquitin (Tobias et al., Science, 1991, 254, 1374, the entire teaching of which is incorporated herein by reference), the unshaded sequences are the linker sequence between ubiquitin and lacZ, of which and the length and amino acids compositions are altered, and the dark-shaded sequence is the beginning of the lacZ gene without the first twenty two amino acids. It has been demonstrated that in addition to the leucine residue immediately after ubiquitin, the linker sequence has a profound impact on the cellular lifetime of -gal. The hydrophobicity of the amino acids composition and the length (or disordered structure) contributes greatly to the overall recognition and delivery of /3-gal to downstream proteases.
In another group, N-terminal signal peptides derived from naturally shortlived proteins are fused to the beginning of -gal to shorten its cellular lifetime. The two strains kl2-n3 [SΕQ ID NOS. 15, 16] and kl2-n5 [SΕQ ID NOS. 17, 18] (where the top row in the sequence identification represents the amino acid sequence and the bottom two rows represent the nucleotide sequence) depicted in FIG. 3(a) belong to this group (N-terminal modification). In the sequence panel of kl2-n3 and k2-n5, the light-shaded sequences are signal peptides taken from the published work of Flynn et
al, (Flynn, JM. et al, Molecular Cell, 2003, 11, 671-683, the entire teaching of which is incorporated herein by reference), and the dark-shaded sequence is the beginning of the lacZ gene without the first methionine.
FIG 3(b) illustrates the different cellular lifetimes of these modified /3-gals expressed from E.coli chromosome, as indicated by the different DDAO-gal hydrolysis rates. The measurements were done using a fluorometer, in which DDAO- gal at a final concentration of 100 μM was added to E. coli cells grown to middle log phase in M9 minimal media. The fluorescence of the hydrolyzed product, DDAO, was monitored over time at 660 nm with excitation at 638 nm. The hydrolysis rate was then calculated by measuring the slope of the fluorescence increase over time. As a reference, the DDAO-gal hydrolysis rate by the wild type kl2 strain is also shown.
To illustrate the use of jS-gal as a reporter gene, investigators chose E.coli as a test organism. A plasmid encoding for ampicillin resistance gene -lactamase was transformed into E.col strains. The presence of the -lactamase allows the usage of the antibiotic ampicillin, which not only keeps the contamination of other bacteria minimal, but also increases the permeability of the E.coli cell wall to the fluorogenic substrate DDAO-gal. The mechanism of the increased cell wall permeability is very likely due to the known fact that ampicillin inhibits cell wall synthesis. All the strains described in this invention contain such an ampicillin-encoding plasmid. Figure 4 shows the measurements of the DDAO fluorescence signal generated by the hydrolysis of DDAO-gal in the wild type E.coli cells. The measurements were done using a fluorometer under the same conditions as described in FIG 3(b). A fluorescence signal increase can be observed immediately upon the addition of the substrate, demonstrating that DDAO-gal can permeate through cell wall and inner membrane of E. coli.
In contrast, as a control experiment, a negligible rate of DDAO-gal hydrolysis was observed in a lacZ deficient strain (lacZ) of E. coli (FIG. 5 depicts both the amino acid sequence [SΕQ ID NO. 19] and the DNA sequence [SΕQ ID NO. 20] around the region where lacZ is deleted from chromosme.) that is primarily due to
autohydrolysis. This experiment demonstrates that the hydrolysis of DDAO-gal is specific to the presence of /3-gal.
The microscopy experiment was performed using a through-lens total internal reflection (TIR) microscope from Olympus and an intensified CCD camera from Roper Scientific. The total internal reflection excitation allows detection of only a thin layer
(< 400 nm) above the cover slip and effectively suppresses the fluorescene background of the medium. The excitation light was set at 638 nm, wherein autofluorescence of the E. coli cell is negligible. This detection system assures the highest sensitivity available. The sample chamber (Bioptech) was maintained at 37°C with M9 minimal medium perfusing through the chamber. E. coli cells were pushed down on the glass coverslip by a droplet of agarose gel.
As shown in FIG. 6, a strong DDAO fluorescence signal from individual E. coli cells with wild type -gal (long lifetime about 10 hours) was detected. This was done at the basal level, i.e., the lacZ gene is not induced. DDAO can diffuse out or be expelled by the cell. Once it leaves the cell, DDAO quickly diffuses out from the probe volume. A steady signal was observed. In contrast, as shown in FIG. 7, when the gene encoding for a short-lived Ub-Arg- -gal (see FIG 2 for sequence) replaced the wild type lacZ gene encoding for the long-lived S-gal on the chromosome, single fluorescence bursts corresponding to stochastic expression of the lacZ gene were observed in real time in single E.coli cell (see FIG. 7(a) for the fluorescence images of E.coli cells). Each burst is triggered by the dissociation of the lac repressor from the lac operator on the E.coli chromosome. The fluorescence off time corresponds to the time required for the repressor to dissociate from the operator sequence, while the fluorescence on time corresponds to the time required for the degradation of /3-gal. Moreover, as shown in FIG. 7(b), the time trace of the fluorescence bursts exhibits quantized levels corresponding to /3-gal molecules generated and degraded one molecule at a time. This demonstrates the signal molecule's sensitivity of this reporting system.
In this embodiment, in order to detect genes with higher expression levels, a short-lived version of a yellow fluorescent protein (YFP) variant, Venus, (Venus- ssrA) is employed. Extensive randomized and directed mutagenesis efforts have produced various GFP and YFP derivatives with faster maturation time than the wild type GFP (30-90 minutes), (Tsien, R.Y., The green fluorescent protein. Annu Rev Biochem, 1998. 67: p. 509-44, the entire teaching of which is incorporated herein by reference), thereby enabling the use of GFPs and YFPs as reporters for transient dynamic changes. One of the more promising YFP variants is "Venus," which matures in ~ 3 minutes in vitro, see, Nagai, T., et al., Nat Biotechnol, 2002. 20(1): p. 87-90, the entire teaching of which is incorporated herein by reference. Like other YFPs, the matured Venus is stable in the cell with a lifetime of ~24 hours, see, Li, X., et al, J Biol Chem, 1998. 273(52): p. 34970-5, the entire teaching of which is incorporated herein by reference.
One aspect in particular pertains to a short-lived Venus variant by creating a Venus-ssrA construct. In this construct, the ssrA peptide tag sequence (AANDENYAKAAA, [SEQ ID NO. 21]) was encoded at the DNA level as a C- terminal fusion to Venus. Normally, a bacterial cell uses a ssrA sequence to flag a protein as the result of a prematurely terminated translation (see Kenneth C. Keiler, Patrick R. H. Waller, Robert T. Sauer, Science, 1996, 271, 990-993). Tagging Venus with ssrA tag recruits cellular protein degradation machinery and greatly reduces the cellular lifetime of Venus from more than 24 hours to less than 30 minutes. It is straightforward to extend this strategy to other GFP variants for construction of other GFP based short-lived reporter proteins.
FIG 8(a) illustrates plasmid pVS5 which encodes the Venus-ssrA gene. FIG 8 (b) shows the nucleotide [SEQ ID NO. 22] and amino acid sequences [SEQ ID NO. 23] of the Venus-ssrA gene. The first amino acid shown in the figure is the first amino acid of Venus. To generate the venus-ssrA reporter gene, a pair of PCR primers (5' CACCAGC AAGGGCGAGGAGCTGTTC-3' [SEQ ID NO. 24] and 5' TTCTTAGGCGGCTAAGG
CGTAGTTCTCGTCGTTGGCGGCCTTGTACAGCTCGTCCATGC-3' [SEQ ID NO. 25] ) were used to amplify the Venus gene from a plasmid pCS2/venus (Nagai T,
Ibata K, Park ES, Kubota M, Mikoshiba K, Miyawaki A. Nat Biotechnol. 20(1):87- 90) and add the ssrA sequence at the 3' end of the Venus gene. The resulting PCR fragment was then ligated into pBAD202/TOPO vector (Invitrogen Inc.) to generate plasmid pVS5.
Again, this general strategy allows for highly sensitive detection of dynamic processes in living cells free from the complication of large fluorescence background associated with protein accumulation.
Short-lived /3-gal and short-lived YFP are complimentary to each other. Short-lived /3-gal can be used to detect genes that are expressed at low copy numbers because of the enzymatic amplification. Short-lived YFP provides a linear response to high-level gene expression. Real-time analysis of short-lived- YFP-incorporated cells typically work under aerobic conditions, while short-lived S-gal incorporated cells typically work under both aerobic and anaerobic conditions. The combination of the two reporter proteins will cover a broad range of intracellular gene expression levels and applicable organisms
In another embodiment, compositions and methods are described for live-cell microarrays. In this embodiment, multiple libraries of cells each differing in at least one genotypic property are prepared. In one aspect, a live-cell microarray is comprised of two libraries. One library comprises cells each of which has a promoterless lacZ gene encoding for a short-lived -gal with its own ribosome binding site that is operatively linked to one promoter controlled region in the host cell's genome. The second library comprises the same elements except that a gene encoding for a short-lived YFP (Venus-ssrA) replaces a gene encoding for a shortlived /3-gal.
The construction of the libraries can be accomplished by random insertion mediated by transposition or by homologous recombination. DNA sequencing around the insertion of the cells in the library will allow a practitioner to identify the position of the insertion with respect to the genome. In one particular aspect, a 75 x 75 element array is sufficient to contain a library with one insertion per gene for a
genome has approximately 4000 genes (E.coli has about 4000 genes). (It should be noted that one skilled in the art will appreciate that various other arrays can be employed.) In another particular aspect, instead of inserting each reporter per gene, the reporter is operatively linked per operon. The size of the array can be smaller if only one insertion is allowed per promoter-controlled region.
Two sets of live-cell microarrays are made from the two libraries of cells with, for example, liquid handling robots preparing the cells on a substrate such as a glass slide with a micro droplet of agarose containing growth media on top of the cells in order to immobilize the cells for ease of measurement, storage and transportation.
Examining the microarrays under a fluorescence microscope, one can study gene expression responses to stimuli and/or environmental changes. For example, parallel movies of all elements of the microarrays can be recorded and vast amounts of data can be compiled and analyzed. The microarrays provide first-of-a-kind genome- wide gene expression profiling and massive kinetics data with high sensitivity and time resolution in living cells
The advantages of employing live-cell microarrays can be summarized as follows:
(1) real-time and parallel observations; (2) high throughput system- wide profiling; (3) quantitative analyses of gene expression levels; (4) background free measurements due to short cellular lifetime of the reporter proteins; (5) high sensitivity for low copy number genes due to enzymatic amplification; (6) single cell sensitivity enabling observation of stochastic events; (7) high time resolution (minute) allowing observations of transient behaviors; (8) broad dynamics range afforded by the combination of two reporter genes; (9) ease and low cost in studying the microarrays with commercially available fluorescence microscopes in non-specialized laboratories; and (10) low cost in microarray replication for distribution to the scientific community.
In one aspect, one cell per element of the microarray (e.g., lOOμm x lOOμm) can be effectuated. In order to obtain reliable statistics, however, one may wish to
place a larger number of cells (10-100) per array element. In addition, high sensitivity makes it possible to observe the behavior of single bacterial cells in a microbial community. Not only can one detect common trends in expression profiles, a practitioner can also observe how gene expression in one cell affects its neighbors, allowing an investigator to pinpoint cooperative effects among cells. Finally, with the background rejection advantage of confocal or total internal reflection microscopy, one has both the high sensitivity to detect low-level expression events and the ability to penetrate multiple layer of biofilm.
It is important to stress that the present invention possesses significant sensitivity for detecting a single copy of reporter proteins in single cells, as exemplified in the Example section (see below). This allows stochastic events of gene expression of low copy number genes to be observed. Stochasticity of gene expression has attracted many experimental and theoretical efforts recently. Combined with the live-cell arrays and short-lived reporter proteins, the highly sensitive measurements of gene expression provide unprecedented information on the working of the genetic network of a genome.
Systematic analyses of the gene expression patterns and their temporal evolution are expected to provide detailed information and generate new insights into function and control of gene expression processes.
In one embodiment of the present invention, cell sorting is facilitated by the compositions described herein. In this embodiment, an illuminogenic substrate, such a fluorescence substrate is introduced to a cell or population of cells, wherein the substrate enters the cells. A nucleotide sequence encoding a reporter protein of the instant invention is also introduced to the cells and is operatively linked within the cell's genome. For instance, the reporter gene (i.e., the nucleotide sequence encoding for the reporter protein) can be operatively linked to a predetermined host gene.
As described above, the reporter protein comprises enzymatic activity such that when it is expressed within a host cell it can facilitate the conversion of the illuminogenic substrate to an illuminesence molecule. With this system in place, a
practitioner can examine various perturbations made upon the cell or cell population and determine if a particular perturbation or set of perturbations trigger the translation of a particular protein. If a particular gene, which is operatively linked to a reporter gene, is expressed upon a perturbation(s) to the cell or any of its components, then an illuminogenic signal will be emitted.
Cells emitting a particular signal can then be separated from cells not emitting such a signal. For example, conventional fluorescence cell sorters are available and can be employed in this embodiment.
Agents used to perturb a cell can include, but not limited to, pharmaceutical agents, including test agents, pesticides, chemical agents both gaseous and in liquid form, hormones, metabolites, toxins, pheromones, and alike.
To facilitate the understanding of the present invention, a number of terms and phrases are defined below:
As used herein, the term "nucleotide" is used to include polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Nucleotides can have any three-dimensional structure, and can perform any function, known or unknown. The following are non-limiting examples of nucleotides: a gene or gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, recombinant nucleotides, branched nucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, nucleic acid probes, and primers. A nucleotide can comprise modified nucleotides, such as methylated nucleotides and nucleotide analogs. If present, modifications to the nucleotide structure can be imparted before or after assembly of the polymer. The sequence of nucleotides may be interrupted by non-nucleotide components. A nucleotide may be further modified after polymerization, such as by conjugation with a labeling component. The term also includes both double- and single-stranded molecules. Unless otherwise specified or required, any embodiment of this invention that is a nucleotide encompasses both the double-stranded form and
each of two complementary single-stranded forms known or predicted to make up the double-stranded form.
A nucleotide is composed of a specific sequence of four nucleotide bases: adenine (A); cytosine (C); guanine (G); thymine (T); and uracil (U) for thymine when the polynucleotide is RNA. Thus, the term "nucleotide sequence" is the alphabetical representation of a nucleotide molecule. This alphabetical representation can be inputted into databases in a computer having a central processing unit and used for bioinformatics applications such as functional genomics and homology searching.
A "gene" includes a nucleotide containing at least one open reading frame that is capable of encoding a particular polypeptide or protein after being transcribed and translated. Any of the nucleotide sequences described herein may be used to identify larger fragments or full-length coding sequences of the gene with which they are associated. Methods of isolating larger fragment sequences are known to those of skill in the art, some of which are described herein.
A "gene product" includes an amino acid, e.g., peptide or polypeptide, generated when a gene is transcribed and then translated.
A "primer" includes a short nucleotide, generally with a free 3 '.-OH group that binds to a target or "template" present in a sample of interest by hybridizing with the target, and thereafter promoting polymerization of a nucleotide complementary to the target. A "polymerase chain reaction" ("PCR") is a reaction in which replicate copies are made of a target polynucleotide using a "pair of primers" or "set of primers" consisting of "upstream" and a "downstream" primer, and a catalyst of polymerization, such as a DNA polymerase, typically a thermally-stable polymerase enzyme. Methods for PCR are well known in the art, and are taught, for example, in MacPherson et al, LRL Press at Oxford University Press (1991). All processes of producing replicate copies of a nucleotide, such as PCR or gene cloning, are collectively referred to herein as "replication". A primer can also be used as a probe in hybridization reactions, such as Southern or Northern blot analyses (see, for example, Sambrook, J., Fritsh, E. F., and Maniatis, T. Molecular Cloning: A
Laboratory Manual. 2nd, ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989).
The term "cDNAs" includes complementary DNA, that is mRNA molecules present in a cell or organism made into cDNA with an enzyme such as reverse transcriptase. A "cDNA library" includes a collection of mRNA molecules present in a cell or organism, converted into cDNA molecules with the enzyme reverse transcriptase, then inserted into "vectors" (other DNA molecules that can continue to replicate after addition of foreign DNA). Exemplary vectors for libraries include bacteriophage, viruses that infect bacteria, e.g., λ phage. The library can then be probed for the specific cDNA (and thus mRNA) of interest.
A "delivery vehicle" includes a molecule that is capable of inserting one or more nucleotides into a host cell. Examples of delivery vehicles are liposomes, biocompatible polymers, including natural polymers and synthetic polymers; lipoproteins; polypeptides; polysaccharides; lipopolysaccharides; artificial viral envelopes; metal particles; and bacteria, viruses and viral vectors, such as baculo virus, adeno virus, and retro virus, bacteriophage, cosmid, plasmid, fungal vector and other recombination vehicles typically used in the art which have been described for replication and/or expression in a variety of eukaryotic and prokaryotic hosts. The delivery vehicles may be used for replication of the inserted nucleotide, gene therapy as well as for simply polypeptide and protein expression.
A "vector" includes a self -replicating nucleic acid molecule that transfers an inserted polynucleotide into and/or between host cells. The term is intended to include vectors that function primarily for insertion of a nucleic acid molecule into a cell, replication vectors that function primarily for the replication of nucleic acid and expression vectors that function for transcription and/or translation of the DNA or RNA. Also intended are vectors that provide more than one of the above function.
A "host cell" is intended to include any individual cell or cell culture that can be or has been a recipient for vectors or for the incorporation of exogenous nucleic acid molecules, nucleotides and/or proteins. It also is intended to include progeny of
a single cell. The progeny may not necessarily be completely identical (in morphology or in genomic or total DNA complement) to the original parent cell due to natural, accidental, or deliberate mutation. The cells may be prokaryotic, include but are not limited to bacterial cells.
The term "genetically modified" includes a cell containing and/or expressing a foreign gene or nucleic acid sequence that in turn modifies the genotype or phenotype of the cell or its progeny. This term includes any addition, deletion, or disruption to a cell's endogenous nucleotides.
As used herein, "expression" includes the process by which nucleotides are transcribed into mRNA and translated into peptides, polypeptides, or proteins. If the nucleotide is derived from genomic DNA, expression may include splicing of the mRNA, if an appropriate eukaryotic host is selected. Regulatory elements required for expression include promoter sequences to bind RNA polymerase and transcription initiation sequences for ribosome binding. For example, a bacterial expression vector includes a promoter such as the lac promoter and for transcription initiation the Shine-Dalgarno sequence and the start codon AUG (Sambrook, J., Fritsh, E. F., and Maniatis, T. Molecular Cloning: A Laboratory Manual. 2nd, ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989). Similarly, a eukaryotic expression vector includes a heterologous or homologous promoter for RNA polymerase II, a downstream polyadenylation signal, the start codon AUG, and a termination codon for detachment of the ribosome. Such vectors can be obtained commercially or assembled by the sequences described in methods well known in the art, for example, the methods described below for constructing vectors in general.
"Differentially expressed", as applied to a gene, includes the differential production of mRNA transcribed from a gene or a protein product encoded by the gene. A differentially expressed gene may be overexpressed or underexpressed as compared to the expression level of a normal or control cell. In one aspect, it includes a differential that is 2.5 times, preferably 5 times or preferably 10 times higher or lower than the expression level detected in a control sample. The term "differentially
expressed" also includes nucleotide sequences in a cell or tissue which are expressed where silent in a control cell or not expressed where expressed in a control cell.
The term "peptide" includes a compound of two or more subunit amino acids, amino acid analogs, or peptidomimetics. The subunits may be linked by peptide bonds. In another embodiment, the subunit may be linked by other bonds, e.g., ester, ether, etc. As used herein the term "amino acid" includes either natural and/or unnatural or synthetic amino acids, including glycine and both the D or L optical isomers, and amino acid analogs and peptidomimetics. A peptide of three or more amino acids is commonly referred to as an oligopeptide. Peptide chains of greater than three or more amino acids are referred to as a polypeptide or a protein.
"Hybridization" includes a reaction in which one or more nucleotides react to form a complex that is stabilized via hydrogen bonding between the bases of the nucleotide residues. The hydrogen bonding may occur by Watson-Crick base pairing, Hoogstein binding, or in any other sequence-specific manner. The complex may comprise two strands forming a duplex structure, three or more strands forming a multi-stranded complex, a single self-hybridizing strand, or any combination of these. A hybridization reaction may constitute a step in a more extensive process, such as the initiation of a PCR reaction, or the enzymatic cleavage of a nucleotide by a ribozyme.
Hybridization reactions can be performed under conditions of different "stringency." The stringency of a hybridization reaction includes the difficulty with which any two nucleic acid molecules will hybridize to one another. Under stringent conditions, nucleic acid molecules at least 60%, 65%, 70%, 75% identical to each other remain hybridized to each other, whereas molecules with low percent identity cannot remain hybridized. A preferred, non-limiting example of highly stringent hybridization conditions are hybridization in 6 X sodium chloride/sodium citrate (SSC) at about 45°C, followed by one or more washes in 0.2 X SSC, 0.1% SDS at 50°C, preferably at 55°C, more preferably at 60°C, and even more preferably at 65°C.
When hybridization occurs in an antiparallel configuration between two single-stranded nucleotides, the reaction is called "annealing" and those nucleotides are described as "complementary". A double-stranded nucleotide can be "complementary" or "homologous" to another nucleotide, if hybridization can occur between one of the strands of the first nucleotide and the second. "Complementarity" or "homology" (the degree that one nucleotide is complementary with another) is quantifiable in terms of the proportion of bases in opposing strands that are expected to hydrogen bond with each other, according to generally accepted base-pairing rules.
As used herein, the term "nucleic acid molecule" is intended to include DNA molecules, e.g., cDNA or genomic DNA, and RNA molecules, e.g., mRNA, and analogs of the DNA or RNA generated using nucleotide analogs. The nucleic acid molecule can be single-stranded or double-stranded, but preferably is double-stranded DNA.
The term "isolated nucleic acid molecule" includes nucleic acid molecules, which are separated from other nucleic acid molecules that are present in the natural source of the nucleic acid. For example, with regards to genomic DNA, the term "isolated" includes nucleic acid molecules that are separated from the chromosome with which the genomic DNA is naturally associated. Preferably, an "isolated" nucleic acid is free of sequences which naturally flank the nucleic acid (i.e., sequences located at the 5' and 3' ends of the nucleic acid) in the genomic DNA of the organism from which the nucleic acid is derived. For example, in various embodiments, the isolated marker nucleic acid molecule of the invention, or nucleic acid molecule encoding a peptide marker of the invention, can contain less than about 5 kb, 4kb, 3kb, 2kb, 1 kb, 0.5 kb or 0.1 kb of nucleotide sequences which naturally flank the nucleic acid molecule in genomic DNA of the cell from which the nucleic acid is derived. Moreover, an "isolated" nucleic acid molecule, such as a cDNA molecule, can be substantially free of other cellular material, or culture medium when produced by recombinant techniques, or substantially free of chemical precursors or other chemicals when chemically synthesized.
A nucleic acid molecule of the present invention can be isolated using standard molecular biology techniques and the sequence information provided herein. Using all or portion of the nucleic acid sequence as a hybridization probe, a molecule comprising a nucleotide sequence of the present invention can be isolated using standard hybridization and cloning techniques as described in Sambrook, L, Fritsh, E. F., and Maniatis, T. Molecular Cloning: A Laboratory Manual. 2nd, ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989.
A nucleic acid of the invention can be amplified using cDNA, mRNA or alternatively, genomic DNA, as a template and appropriate nucleotide primers according to standard PCR amplification techniques. The nucleic acid so amplified can be cloned into an appropriate vector and characterized by DNA sequence analysis. Furthermore, nucleotides corresponding to marker nucleotide sequences, or nucleotide sequences encoding a marker of the invention can be prepared by standard synthetic techniques, e.g., using an automated DNA synthesizer.
A nucleic acid molecule of the invention, moreover, can comprise only a portion of the nucleic acid sequence of the invention, or a fragment which can be used as a probe or primer. The probe/primer typically comprises substantially purified nucleotide.
Probes based on the nucleotide sequence of a nucleic acid molecule encoding a peptide of the present invention can be used to detect agglomeration proteins. In other embodiments, the probe comprises a labeling group attached thereto, e.g., the labeling group can be a radioisotope, a fluorescent compound, an enzyme, or an enzyme co-factor. Such probes can be used as a part of a diagnostic test kit for identifying cells or tissue which misexpresses, e.g., over- or under-express, a polypeptide of the invention, or which have greater or fewer copies of a gene of the invention.
As used herein, the term "hybridizes under stringent conditions" is intended to describe conditions for hybridization and washing under which nucleotide sequences at least 60% homologous to each other typically remain hybridized to each other. Preferably, the conditions are such that sequences at least about 70%, more preferably at least about 80%, even more preferably at least about 85% or 90% homologous to each other typically remain hybridized to each other. Such stringent conditions are known to those skilled in the art and can be found in Current Protocols in Molecular- Biology, John Wiley & Sons, N.Y. (1989), 6.3.1-6.3.6. A preferred, non-limiting example of stringent hybridization conditions are hybridization in 6 X sodium chloride/sodium citrate (SSC) at about 45°C, followed by one or more washes in 0.2 X SSC, 0.1% SDS at 50°C, preferably at 55°C, more preferably at 60°C, and even more preferably at 65°C. Preferably, an isolated nucleic acid molecule of the invention that hybridizes under stringent conditions to the sequence of SEQ ID NO. 1-10. As used herein, a "naturally-occurring" nucleic acid molecule includes an RNA or DNA molecule having a nucleotide sequence that occurs in nature, e.g., encodes a natural protein.
In other embodiments, the nucleotides of the invention can include other appended groups such as peptides, e.g., for targeting host cell receptors in vivo, or agents facilitating transport across the cell membrane (see, e.g., Letsinger et al. (1989) Proc. Natl. Acad. Sci. USA 86:6553-6556; Lemairre et al. (1987) Proc. Natl Acad. Sci. USA 84:648-652; PCT Publication No. W088/09810) or the blood-brain barrier (see, e.g., PCT Publication No. W0 89/10134). In addition, nucleotides can be modified with hybridization-triggered cleavage agents (see, Krol et al. (1988) Bio- Techniques 6:958-976) or intercalating agents (see, Zon (1988) Pharm. Res. 5:539- 549). To this end, the nucleotide may be conjugated to another molecule, e.g., a peptide, hybridization triggered cross-linking agent, transport agent, or hybridization- triggered cleavage agent. Finally, the nucleotide may be detectably labeled, either such that the label is detected by the addition of another reagent, e.g., a substrate for an enzymatic label, or is detectable immediately upon hybridization of the nucleotide, e.g., a radioactive label or a fluorescent label, e.g., a molecular beacon as described in U.S. Patent 5,876,930.
Another aspect of the invention pertains to vectors, preferably expression vectors, containing a nucleic acid encoding a marker protein of the invention (or a portion thereof). As used herein, the term "vector" includes a nucleic acid molecule capable of transporting another nucleic acid to which it has been linked. One type of vector is a "plasmid", which includes a circular double stranded DNA loop into which additional DNA segments can be ligated. Another type of vector is a viral vector, wherein additional DNA segments can be ligated into the viral genome. Certain vectors are capable of autonomous replication in a host cell into which they are introduced, e.g., bacterial vectors having a bacterial origin of replication and episomal mammalian vectors. Other vectors, e.g., non-episomal mammalian vectors, are integrated into the genome of a host cell upon introduction into the host cell, and thereby are replicated along with the host genome. Moreover, certain vectors are capable of directing the expression of genes to which they are operatively linked. Such vectors are referred to herein as "expression vectors." In general, expression vectors of utility in recombinant DNA techniques are often in the form of plasmids. In the present specification, "plasmid" and "vector" can be used interchangeably as the plasmid is the most commonly used form of vector.
The recombinant expression vectors of the invention comprise a nucleic acid of the invention in a form suitable for expression of the nucleic acid in a host cell, which means that the recombinant expression vectors include one or more regulatory sequences, selected on the basis of the host cells to be used for expression, which is operatively linked to the nucleic acid sequence to be expressed. Within a recombinant expression vector, "operatively linked" is intended to mean that the nucleotide sequence of interest is linked to the regulatory sequence(s) in a manner which allows for expression of the nucleotide sequence, e.g., in an in vitro transcription/translation system or in a host cell when the vector is introduced into the host cell. The term "regulatory sequence" is intended to include promoters, enhancers and other expression control elements, e.g., polyadenylation signals. Such regulatory sequences are described, for example, in Goeddel; Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990). Regulatory sequences include those which direct constitutive expression of a nucleotide sequence in many types of host cells and those which direct expression of the nucleotide
sequence only in certain host cells, e.g., tissue-specific regulatory sequences. It will be appreciated by those skilled in the art that the design of the expression vector can depend on such factors as the choice of the host cell to be transformed, the level of expression of protein desired, and the like. The expression vectors of the invention can be introduced into host cells to thereby produce proteins or peptides, including fusion proteins or peptides, encoded by nucleic acids as described herein, e.g., marker proteins, mutant forms of marker proteins, fusion proteins, and the like.
The recombinant expression vectors of the invention can be designed for expression of marker proteins in prokaryotic or eukaryotic cells. For example, proteins can be expressed in bacterial cells such as E. coli, insect cells (using baculovirus expression vectors) yeast cells or mammalian cells. Suitable host cells are discussed further in Goeddel, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, CA (1990). Alternatively, the recombinant expression vector can be transcribed and translated in vitro, for example using T7 promoter regulatory sequences and T7 polymerase.
Expression of proteins in prokaryotes is most often carried out in E. coli with vectors containing constitutive or inducible promoters directing the expression of either fusion or non-fusion proteins. Fusion vectors add a number of amino acids to a protein encoded therein, usually to the amino terminus of the recombinant protein. Such fusion vectors typically serve three purposes: 1) to increase expression of recombinant protein; 2) to increase the solubility of the recombinant protein; and 3) to aid in the purification of the recombinant protein by acting as a ligand in affinity purification. Often, in fusion expression vectors, a proteolytic cleavage site is introduced at the junction of the fusion moiety and the recombinant protein to enable separation of the recombinant protein from the fusion moiety subsequent to purification of the fusion protein. Such enzymes, and their cognate recognition sequences, include Factor Xa, thrombin and enterokinase. Typical fusion expression vectors include pGEX (Pharmacia Biotech Inc; Smith, D.B. and Johnson, K.S. (1988) Gene 67:31-40), pMAL (New England Biolabs, Beverly, MA) and pRIT5 (Pharmacia, Piscataway, NJ) which fuse glutathione S-transferase (GST), maltose E binding protein, or protein A, respectively, to the target recombinant protein.
Purified fusion proteins can be utilized in marker activity assays, e.g., direct assays or competitive assays described in detail below, or to generate antibodies specific for marker proteins for example.
Examples of suitable inducible non-fusion E. coli expression vectors include pTrc (Amann et al, (1988) Gene 69:301-315) and pET 1 Id (Studier et al, Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, California (1990) 60-89). Target gene expression from the pTrc vector relies on host RNA polymerase transcription from a hybrid trp-lac fusion promoter. Target gene expression from the pET 1 Id vector relies on transcription from a T7 gnlO-lac fusion promoter mediated by a coexpressed viral RNA polymerase (T7 gnl). This viral polymerase is supplied by host strains BL21(DE3) or HMS174(DE3) from a resident prophage harboring a T7 gnl gene under the transcriptional control of the lacUV 5 promoter.
One strategy to maximize recombinant protein expression in E. coli is to express the protein in a host bacteria with an impaired capacity to proteolytically cleave the recombinant protein (Gottesman, S., Gene Expression Technology: Methods in Enzymology 185, Academic Press, San Diego, California (1990) 119- 128). Another strategy is to alter the nucleic acid sequence of the nucleic acid to be inserted into an expression vector so that the individual codons for each amino acid are those preferentially utilized in E. coli (Wada et al, (1992) Nucleic Acids Res. 20:2111-2118). Such alteration of nucleic acid sequences of the invention can be carried out by standard DNA synthesis techniques.
Another aspect of the invention pertains to host cells into which a nucleic acid ' molecule of the invention is introduced within a recombinant expression vector or a nucleic acid molecule of the invention containing sequences which allow it to homologously recombine into a specific site of the host cell's genome. The terms "host cell" and "recombinant host cell" are used interchangeably herein. It is understood that such terms refer not only to the particular subject cell but also to the progeny or potential progeny of such a cell. Because certain modifications may occur
in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term as used herein.
A host cell can be any prokaryotic or eukaryotic cell. Preferably, the host cell is a prokaryotic cell. For example, the invention can be expressed in bacterial cells such as E. coli. Other suitable host cells are known to those skilled in the art.
Vector DNA can be introduced into host cells via conventional transformation or transfection techniques. As used herein, the terms "transformation" and "transfection" are intended to refer to a variety of art-recognized techniques for introducing foreign nucleic acid, e.g., DNA, into a host cell, including calcium phosphate or calcium chloride co-precipitation, DEAE-dextran-mediated transfection, lipofection, or electroporation. Suitable methods for transforming or transfecting host cells can be found in Sambrook, et al. (Molecular Cloning: A Laboratory Manual. 2nd, ed., Cold Spring Harbor Laboratory, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY, 1989), and other laboratory manuals.
A host cell of the invention, such as a host cell in culture, can be used to produce, i.e., express, a recombinant protein. Accordingly, the invention further provides methods for producing a protein using the host cells of the invention. In one embodiment, the method comprises culturing the host cell of invention (into which a recombinant expression vector encoding a protein, or proteins, has been introduced) in a suitable medium such that a protein of the invention is produced. In another embodiment, the method further comprises isolating a protein from the medium or the host cell.
Of course, one skilled in the art will appreciate further features and advantages of the invention based on the above-described embodiments. Accordingly, the invention is not to be limited by what has been particularly shown and described, except as indicated by the appended claims.
EXAMPLES
Example 1 : Detection of transient gene expression in single living E.coli cells with sensitivity for one protein molecule
(i) Construction of a short-lived β-gal
As discussed supra, in order to observe individual events involved in the expression of the lacZ gene, one must construct an E. coli strain that expresses shortlived β-gal. To achieve this goal, the inventors employed the so-called N-end rule (Tobias, J.W., et al, Science, 1991. 254(5036): p. 1374-7, and Varshavsky, A., Proc Natl Acad Sci USA, 1996.93(22): p. 12142-9, the entire teaching of which is incorporated herein by reference) and N-terminal signal peptides (Flynn, JM. et al, Molecular Cell, 2003, 11, 671-683, the entire teaching of which is incorporated herein by reference) to shorten the cellular lifetime of β-gal. The N-end rule states that the cellular lifetime of a protein is related to its N-terminal amino acid residue. In E. coli, changing N-terminal amino acid from the natural methionine to leucine, arginine, lysine, phenylalanine, tryptophan or tyrosine shortens the protein's half -life to a few minutes. In this experiment, the ubiquitin fusion technique was used in order to introduce a lifetime- shortening amino acid (e.g., leucine or arginine) replacing the methionine at the N-terminus of β-gal to generate Ub-Leu-β-gal or Ub-Arg-β-gal , see, Bachmair, A., D. Finley, and A. Varshavsky, Science, 1986. 234(4773): p. 179- 86, the entire teaching of which is incorporated herein by reference. After this fusion protein is expressed, the ubiquitin is cleaved by ubiquitin- specific protease, thus the argine or the leucine residue is exposed to the proteolytic pathways in E.coli.
An ub-srg-lacZ fusion gene was constructed on a plasmid. However, expressing the fusion gene off of the constructed plasmid would introduce many complications due to the variable copy number of plasmids from cell to cell. Therefore, in order to observe stochastic expression of lacZ at the basal level, this fusion gene was integrated into the E. coli genome by homologous recombination, see, Link, et al, J Bacteriol, 1997. 179(20): p. 6228-37, the entire teaching of which is incorporated herein by reference. The same method was used to construct a lacZ
strain, in which the entire coding sequence of β-gal is deleted. Hydrolysis of DDAO- gal in this lac strain is essentially abolished as compared to that of wild type (see, FIG. 4). This demonstrated that the hydrolysis of DDAO-Gal by β-gal is specific and the fluorescence signal observed is related to the expression of β-gal only.
After replacing the endogenous β-gal gene with Ub-Arg-β-gal, the plasmid pRB293, containing the UBPl gene, the Saccharomyces cerevisiae ubiquitin-specific processing protease (see, Tobias, J.W. and A. Varshavsky, J Biol Chem, 1991. 266(18): p.
12021-8, the entire teaching of which is incorporated herein by reference) was transformed into the cell. The altered cellular lifetime of Ub-Arg-β-gal using spectroscopic methods was then examined.
(ii) Live cell observation.
After obtaining the Ub-Arg-β-gal construct, live cell experiments were conducted employing a total internal reflection fluorescence (TIRF) microscope. The experiments were conducted using various concentrations from 1 μM to 50 μM of DDAO-gal in M9 minimal media, which is perfused through the sample chamber above the E. coli cells pushed down on the glass coverslip by a droplet of agarose gel. Detectable DDAO fluorescence bursts associated with the stochastic events of gene expression even at the reduced number of β-gal were observed (see, FIG. 7). Most . importantly, the time traces of the fluorescence bursts exhibit quantized levels corresponding to β-gal molecules generated and degraded one at a time. This demonstrated that this method has the ultimate sensitivity for even one protein molecule. From these individual lacZ expression events, important parameters were extracted, such as the dissociation constant &d between the lac repressor and the operator and the expression efficiency in the cellular environment. These measurements are important because the thermodynamics and kinetics of biochemical reactions, in principle, can be distinctly different in the cellular environment than in vitro. The previous understanding of these parameters were either obtained from in vitro experiments or deduced indirectly from ensemble cellular measurements.
(iii) Expanding reporter repertoire.
β-gal is but one system to demonstrate the proof of principle of short-lived reporter proteins. This same strategy described herein can be used with other reporter genes in order to track transient behavior. These reporter genes can be used to make fusion proteins for multiplexing observation of gene expression processes. One will be able to study gene regulatory circuits by examining the effects of one gene on another. Such work will offer detailed information about the interactions and regulation among gene products. For example, another reporter, β-glucosidase with a molecular mass of 82 kDa, encoded by the gene bglB from Bacillus sp. GL1 (Arch. Ciochem. Biophys., vol 360, No. 1, pp 1-9, 1998) is employed. This enzyme hydrolyzes the non-reducing terminal glucoside from either carbon hydrates or artificial substrates such as resorufin-glucopyranoside (see FIG. 1 (c) and (d) for the substrate structure and product spectrum). We have expressed recombinant β-D- glucosidase in E. Coli. The strain that express the β-glucosidase gene showed very high hydrolysis activity on resorufin-glucopyranoside, while the wild type E. Coli (does not contain the gene encoding for β-glucosidase) showed negligible glucosidase activity (FIG. 9). As another example, β-lactamase, which hydrolyzes fluoregenic substrates such as CCF2 and CR2/AM (see, Zlokarnik et al, Science, 1998, 279(5347), 84-88, Gao et al, J.Am.Chem. Soc, 2003, 125, 11146-11147, the entire teachings of which are incorporated herein by reference.), can also be genetically modified and employed in the reporting system.
Example 2: Construction of live cell array of E.coli
The construction of two libraries comprised of short-lived β -gal and shortlived YFP and the corresponding live cell arrays are illustrated in E.coli as described in detail in the following. However, it should be obvious to those skilled in the art that other short-lived reporter genes such as β-glucosidase or β -lactamase , and other organisms such as Saccharomyce cerevisiae and Shewanella oneideinsis can equally be the subjects of the present invention.
(i) Construction of a short-lived β-gal and YFP reporter libraries of E.coli
Investigators will proceed with an in vitro Tn5-based transposon system, (see, Gory shin, I.Y., et al., Insertional transposon mutagenesis by electroporation of released Tn5 transposition complexes. Nat Biotechnol, 2000. 18(1): p. 97-100, the entire teaching of which is incorporated herein by reference)
To generate the lacZ library using the in vitro Tn5 mediated transposition, a DNA cassette including a promoter-less ub-x-lacZ gene (the x between ub and lacZ represents any amino acid that shortens the cellular lifetime of the resulting β-gal) will be cloned into a transposon construction vector pMOD-2 (Epicentre Technologies), flanked by two Tn5 -recognizable 19 bp ME sequence. This ub-x-lacZ gene contains its own ribosome binding site (RBS), in front of which a stop codon will be placed to avoid a translation read-through from a previous gene. See, FIG. 10.
Figure 10 is a schematic drawing of the construction of the lacZ library by Tn5 mediated transposition. ME represents Tn5 recognizable mosaic ends sequence (triangles); RBS are the ribosome binding sites (rectangles); and the box joined by a hitched box indicates the ub-x-lacZ gene and the oval with a turn arrow on top indicates a promoter on the chromosome.
The selection for desired colonies containing the reporter genes will be based on blue/white colony screening on X-gal plates. The expression of the promoter-less Ub-X-lacZ gene from a functional promoter on the chromosome will result in blue colonies due to the conversion of X-gal into blue insoluble precipitant by β-gal. Since the conversion of X-gal by β-gal is highly efficient and can accumulate, even colonies transiently expressing β-gal can be identified if sufficient growing time is allowed. Investigators have observed that the E. coli colonies contain one single copy or less of short-lived β-gal on average produce easily visible blue color after 16 hours incubation. Thus, investigators expect nearly all promoters, even the tightly controlled promoters at its basal level activity can be identified using this blue/white screening.
Figure 11 depicts an alternative method for constructing the lacZ and YFP libraries simultaneously. In this method, the last selection step generates both lacZ and yfp libraries based on blue and white colonies screening. (The notations are the same as FIG. 10.)
The approach for constructing the two libraries simultaneously is planned as follows: First, a methylated DNA cassette will be randomly inserted into E.coli genome by Tn5 mediated in vitro transposition as described above. This DNA cassette will contain a copy of ub-x-lacZ (contains a stop codon and its own ribosome binding site in front), and also a copy of Venus-ssrA with its 3' end flanked by approximately 500 bp sequence, which is homologous to the 3' end of the lacZ gene. Between the ub-x-lacZ and Venus-ssrA lies the cat and sacB genes. The first round selection for the incorporation of this DNA cassette into the chromosome will be based on the β-galactosidase activity on the X-gal plate or chloramphenicol resistance. The colonies from the first round of selection will be pooled and plated on sucrose plates supplied with X-gal. Blue colonies that survived on the sucrose plates indicate the presence of the ub-x-lacZ gene on the chromosome, thus forming the lacZ library, while the white colonies indicate the presence of Venus-ssrA, forming the YFP library. Both libraries will then be replicated on chloramphenicol plates to ensure that the survival on sucrose plates is not due to the mutation of the sacB gene (see, Link, A.J., D. Phillips, and G.M. Church).
The investigators are also aware that the construction of the two libraries will probably still leave some promoters not covered. In such cases, insertion of the reporter genes after the specific promoters will be done separately using homologous recombination individually, assuming the numbers of the promoters not covered by the methods described above is not significant.
All the above work can be automated as explained in the following: the blue colony picking, inoculation into 96-well plates, and the subsequent master plates making with arranged colonies will be performed by the Q-bot (Genetix). See, Fig. 12. Colonies from the master plates will be picked directly into 96-well PCR tubes containing reaction buffers prepared by the Genesis liquid-handling robot (Tecan).
The PCR reactions will then be cleaned using Qiaquick 96 PCR purification kit (Qiagen Inc.) on a Beckman BioMek FX robot. Automated DNA sequencing will be performed by a commercial company. The process of reading the DNA sequence files, blasting and mapping onto the Shewanella genome can also be automated by a home-made program.
To identify the position of the reporter gene on the chromosome, regions before and after the Ub-x-lacZ gene in each strain will be sequenced. A modified colony PCR using a protocol called random amplification of transposon ends (RATE), (see, Karlyshev, AN., M.J. Pallen, and B.W. Wren, Single-primer PCR procedure for rapid identification of transposon insertion sites. Biotechniques, 2000. 28(6): p. 1078, 1080, 1082, the entire teaching of which is incorporated herein by reference), will be employed to amplify the regions around the ub-x-lacZ gene. Abundant single-stranded DΝA (ssDΝA) will be first generated by one primer, which specifically targets one end of the ub-x-lacZ gene and goes outward relative to the transposon DΝA. Second, these ssDΝAs will be amplified by random priming at low annealing temperature using the same primer to produce double- stranded DΝA (dsDΝA) with different lengths. Finally, these dsDΝAs will be used as templates and amplified using the same primer at stringent annealing temperature. A new primer, which targets specifically a sequence lying in the middle of the ME sequence and the first primer binding sequence on the transposon, will be used to sequence the amplified dsDΝA. The sequence will then be compared to E.coli genome to identify the position of insertion.
At this step, investigators will pick at least 2 x 104 blue colonies to establish the initial library. At this size, one would expect on average one insertion per 250 bp on the genome (the genome size of E.coli is approximately 4.6 x 106). The goal is to tag every possible promoter (few thousands in total, judging by the predicted 4000 genes in E.coli) with a reporter. Therefore, the activity of each promoter or operon of the entire genome in response to different stimuli can be studied in parallel on one or two live cell arrays. To achieve this goal, investigators will select from the initial library based on the sequence data according to the following criteria: 1) least
disruption of a gene; 2) least polar effect to the downstream genes caused by the insertion; and 3) at least one insertion for one promoter or operon.
The final library is estimated to contain at least 3000 strains including both unique and multiple reporter gene insertions for each predicted promoter.
(b) Construction of live-cell microarrays
Once the construction of the reporter library is completed, investigators will print the reporter-labeled cell strains into microarrays. Nanolitres of aqueous media containing E.coli cells will be pipetted onto a #1 glass coverslip using a robotic micro-arrayer; the Omnigrid (GeneMachine) arrayer is capable of dispensing a minimum of 300 picolitre of fluid. Sub-microlitres of low-melting-temperature agarose containing growth media will be immediately applied on top of the cell solutions to prevent drying of cells. The weight of the agarose will compress the media droplets and create a monolayer of cells on the surface of the cover glass. The resulting spot will be approximately hundreds of micrometers in diameter, which is about the size of the view-field on a microscope. Macroscopic version of this technique has been consistently demonstrated, and cells stay viable and divide for many generations on the slide.
Allowing the spacing between adjacent spots to be 1mm, it is possible to print an array of micro-colonies corresponding to the library of 1000 strains in a 60mm x 60mm area. That is well within the travel range of an automated x-y stage.
Once the array is printed, it will be capped by a gasket and Microaqueduct slide manufactured by Bioptech Inc, as shown in FIG. 12. The microaqueduct slide will allow laminar flow through the chamber and keep the temperature constant via an add-on thermoelectric heater unit. A solution of DDAO-gal and growth media can be perfused through the chamber to keep the cells supplied with nutrients and fluorogenic substrates, while allowing fluorescent product DDAO and cellular metabolites to flow away. Unlike liquid suspensions, this set-up, whereby media is allowed to flow over microbes that are fixed in place, is very similar to those
routinely used for analysis of biofilm formation, and is therefore, more representative of how bacteria exist in the environment.
The microarray encased in a flow chamber offers a versatile and durable platform with a controlled environment and constant supply of nutrients. This chamber will be mounted on TIRFM microscope (Nikon Te-2000E) with a built-in motorized XY stage. Equipped with rotary encoders and feedback stepper motors, the XY stage can visit each micro-colony on the microarray with a repeatability of one micron. With an external Z-drive, the objective lens can auto-focus before acquiring an image at each spot. Shutters and filter wheels controlled by commercial software (such as Metamorph, Universal Imagining Inc.) can precisely time illumination with laser and Xe lamp, to acquire fluorescent and phase-contrast images. These images can be stacked into movies for each point in the microarray. Fluorescent time trajectories will be extracted for each individual cell and proper statistics analysis will be performed.
Investigators will conduct experiments to monitor gene expression responses to various environmental factors with a complete library. Since the environmental changes and factors are uniform for all microcolonies on the array, observation of fluorescence changes at each spot will reflect changes in expression of each tagged promoter induced by the stimuli. Given the capability of the motorized stage and the control software, it can easily scan 100 spots per minute, collecting time-point trajectories of approximately 100 cells at each spot.
High-throughput real-time data provides quantitative information on system- wide gene expression kinetics. This first-of-a-kind systems biology dataset provides an opportunity for mathematical modeling.
Example 3: Real-time gene expression of live Shewanella oneideinsis
(i) Demonstration of DDAO-gal permeability
Investigators have demonstrated that the fluorogenic substrate of β-gal, DDAO-gal, is permeable to the Shewanella oneidensis cell membrane. To do so, they have transformed wild-type (lacZ) Shewanella oneidensis cells with pBBRlMCS5.1 (see, FIG. 13), a plasmid containing the lac operon along with the lacl repressor gene. Figure 10a depicts a plasmid map of the pBBRlMCS-5.1 plasmid. Figure 13b is the nucleotide sequence [SEQ ID NO. 26] encoding the ubiquitin and part of the beginning of β-gal on the plasmid pBBRlMCS5.1.
Figure 14a shows the fluorescence image of the individual transformed cells supplied with DDAO-gal without induction. In contrast, under the same condition, no fluorescence signal was observed in the wild-type strain, see, FIG. 14b. This experiment proves that DDAO-gal can permeate through the Shewanella oneideinsis cell membrane and the fluorescence signal is specifically due to the presence of β-gal.
(ii) Demonstration of a short-lived X-β-gal in Shewanella
The N-end rule has been demonstrated to be universal in organisms examined such as E. coli, yeast and mammals, (see, Varshavsky, A., The N-end rule: functions, mysteries, uses. Proc Natl Acad Sci U S A, 1996. 93(22): p. 12142-9, the entire teaching of which is incorporated herein by reference). Shewanella is closely related to E. coli, therefore, it is reasonable to assume that the same rule also applies in Shewanella. As shown in FIG. 14b, when a short-lived β-gal (Ub-Leu-β-gal) is expressed together with the ubiquitin-specific protease, the hydrolysis rate decrease dramatically, indicating shortened cellular lifetime of β-gal.
Example 3: β-gal applied to Saccharomyce cerevisiae
A short-lived version of β-gal (ub-leu-lacZ) was used in the eukaryotic model organism Saccharomyce cerevisiae (budding yeast) to probe stochastic gene expression events. Saccharyomyce cerevisiae has extensive ubiquitin-dependent protein degradation pathways, thereby enabling a cellular lifetime of modified β-gal less than a few minutes.
The ub-leu-lacZ reporter gene was generated using standard cloning protocols (Sambrook and Russell, Molecular Cloning, 3rd Ed, CSHL press, the entire teaching
of which is incorporated herein by reference) with a pair of PCR primers (5' CTTGGTA CCATGCAGATTTTCGTCAAGACTTTG 3' [SEQ ID NO. 27], and 5' GAGCGGC CGCTTTTGACACCAGACC 3' [SEQ ID NO. 28]) to amplify a 4000bp fragment containing ub-leu-lacZ from pUB23 plasmid generated by Varshavsky, et al. This DNA fragment was ligated into the pYC2/CT plasmid (Invitrogen, Inc.) and the resulting construct was verified by DNA sequencing.
Figure 15a depicts the nucleotide sequence junction of ub-leu-lacZ, and (b) is the amino acid sequence [SEQ ID NO. 29] and nucleic acid sequence [SEQ ID NO. 30] for the junction of the ub-leu-lacZ construct on centromeric plasmid: the sequence is from the Gall promoter site to the Bsu26I site of the lacZ gene, and numbering of the nucleotides is according to the first base of the Gall promoter; the ubiquitin gene is joined by an modified Z cZ gene with its first methionine residue replaced by a leucine residue. The amino acids sequences are shown on top of the DNA sequence panel.
Figure 16 shows DDAO fluorescence generated from the hydrolysis of DDAO-gal by lacZ7 (dark) but not by the lacZ (light) yeast cells measured in a fluorometer. Final concentration of DDAO-gal was 50 μM and S. cerevisiae cells was grown to middle log phase in synthetic dextrose medium. The significantly different hydrolysis rates between the two strains demonstrated that (i) fluorescence substrate DDAO-gal is permeable to S. cerevisiae cell wall and plasma membrane; (ii) DDAO- gal is hydrolyzed by β-gal with remarkable specificity and high turnover rate.
Figure 17 is a fluorescence image of S. cerevisiae cells expressing wild type β-gal. This experiment was done using fluorescence microscope with excitation at 568nm. The other setup is identical to those used in the E.coli and Shewanella experiments as described above. The presence of glucose in the growth media represses the Gall promoter, resulting in a low basal level expression of β-gal. Experimental conditions were chosen to minimize the background and autofluorescence of yeast cells.
Figure 18 shows the fluorescence burst observed on a single S. cerevisiae cell with a short-lived β-gal expressed from the centromeric plasmid. The burst in the time trace indicates a single lacZ gene expression event, resulting from the stochastic dissociation of the repressor from its binding site. The rise of the burst indicates the generation of β-gal and the decay indicates the degradation of β-gal. This is the first experiment demonstrating that the short-lived β-gal reporter system is capable of detecting low copy number translational product and following the gene expression events in realtime in live eukaryotic cells.
Claims
1. A reporting system for monitoring real-time gene expression events in a cell, comprising: an illuminogenic substrate, wherein said substrate is contained within said cell; and at least one reporter protein, wherein said reporter protein facilitates the conversion of said illuminogenic substrate into an illuminescent molecule, and wherein said reporter protein has a short cellular life time.
2. The system of claim 1, wherein said cell can be either a prokaryotic cell or an euckaryotic cell.
3. The system of claim 1 , wherein said illuminogenic substrate is a fluorogenic substrate.
4. The system of claim 3, wherein said fluorogenic substrate is selected from the group consisting of DDAO-galactopyranoside, resorufin-galactopyranoside, resorufin-glucopyranoside, DDAO-glucopyranoside, CCF2, CR2 and alike.
5. The system of claim 1, wherein said system has high sensitivity for low copy number proteins.
6. The system of claim 1, wherein said reporter protein comprises enzymatic activity.
7. The system of claim 6, wherein said enzymatic activity is selected from the group consisting of β-galactosidase, β-glucosidase, β-lactamase, and alike.
8. The system of claim 1 , wherein said short cellular life time for said reporter protein ranges from about 2 minutes to about half an hour.
9. The system of claim 1 , wherein said reporter protein is constructed of an amino acid sequence using the N-end rule in order to shorten the cellular life time of said reporter protein.
10. The system of claim 9, wherein a methionine residue positioned at the amino terminus of said reporter protein is replaced with an amino acid selected from the group consisting of leucine, arginine, lysine, phenylalanine, tryptophan, and tyrosine.
11. The system of claim 10, wherein an ubiquitin nucleotide fusion construct is used to introduce a predetermined nucleotide sequence that encodes for a predetermined amino acid sequence that is expressed resulting in said reporter protein.
12. The system of claim 11 , wherein an ubiquitin fusion expression peptide product is selected from the group consisting of Ub-leu-β-gal, Ub-arg-β-gal, and alike.
13. The system of claim 12, wherein said ubiquitin peptide fusion construct is Ub- arg-ZαcZ.
14. The system of claim 13, wherein a nucleotide sequence for said Ub-arg-ZαcZ is SEQ ID NO. 1, and wherein an amino acid sequence for said Ub-arg-ZαcZ is SEQ ID NO. 2.
15. The system of claim 12, wherein said ubiquitin fusion peptide product is Ub- arg-β-gal.
16. The system of claim 15, wherein said Ub-arg-β-gal has a linker sequence between ubiquitin and β-gal.
17. The system of claim 16, wherein said linker sequence comprises an amino acid sequence selected from group consisting of SEQ ID NO. 9, SEQ ID NO. 11, and SEQ ID NO. 13.
18. The system of claim 17, wherein said linker sequence is encoded by a nucleotide nucleotide sequence selected from the group consisting of SEQ ID NO. 10, SEQ ID NO. 12, and SEQ ID NO. 14.
19. The system of claim 1 , wherein said reporter protein comprise a N-terminal signal amino acid sequence operatively linked to a β-gal amino acid sequence in order to shorten the cellular life time of said reporter protein, wherein said N-terminal signal amino acid sequence is selected from the group consisting of SEQ ID NO. 15 and SEQ ID NO. 17.
20. The system of claim 19, wherein said N-terminal signal amino acid sequence is encoded by a nucleotide sequence selected from the group consisting of SEQ ID NO. 16 and SEQ ID NO. 18.
21. The system of claim 1 , wherein a nucleotide sequence encoding for said reporter protein is operatively linked within the genome of said cell.
22. The system of claim 21, wherein said nucleotide sequence encoding for said reporter protein is operatively linked to a host cell nucleotide sequence that encodes a host specific protein.
23. A live-cell microarray, comprising multiple libraries of cells each of which differs from the rest in at least one genotypic property.
24. The microarray of claim 23, wherein said cells are selected from the group consisting of prokaryotic cells and eukaryotic cells.
25. The microarray of claim 23 comprising two libraries of cells.
26. The microarray of claim 25, wherein a first library of cells has cells that have a promoterless ZαcZ gene that encodes for a short lived β-gal having its own ribosome binding site, wherein said promoterless ZαcZ gene is inserted within a host's promoter region, and wherein a second library of cells has the same elements as said first library of cells except that a gene encoding for a short lived YFP replaces said gene encoding for said short lived β-gal.
27. The microarray of claim 26, wherein said YFP is a Venus-ssrA nucleotide construct.
28. The microarray of claim 27, wherein said Venus-ssrA nucleotide construct encodes a ssrA peptide tag, wherein said ssrA peptide tag amino acid sequence is SEQ ID NO. 21.
29. The microarray of claim 28, wherein said Venus-ssrA nucleotide construct is operatively linked to a nucleotide sequence within the pVS5 plasmid, and wherein said Venus-ssrA nucleotide sequence is SEQ ID NO. 22, wherein said ssrA nucleotide encodes for an amino acid sequence, and wherein said amino acid sequence is SEQ ID NO. 23.
30. A method of monitoring gene expression in real-time in a living cell, comprising the steps of:
(a) obtaining a cell population in which at least one cell comprises at least one nucleotide sequence encoding a reporter protein operatively linked to said cell's genome, wherein said reporter protein when expressed has enzymatic activity and has a short cellular life time;
(b) introducing an illuminogenic substrate to said cell under conditions suitable for the conversion of said illuminogenic substrate to an illuminescent molecule, wherein said substrate enters said cell; and
(c) detecting an illuminescent signal.
31. The method of claim 30, wherein the cell population can comprise either prokaryotic cells, eukaryotic cells, or a combination of both.
32. The method of claim 31 , wherein the cell population is selected from the group consisting of prokaryotic cells and eukaryotic cells.
33. The method of claim 30, wherein said nucleotide sequence encoding for a reporter protein is operatively linked to a predetermined host cell's nucleotide sequence that encodes a specific protein.
34. The method of claim 30, wherein said reporter protein comprises an amino acid sequence selected from the group consisting of Ub-leu-β-gal, Ub-arg-β-gal, and alike.
35. The method of claim 30, wherein said illuminogenic substrate is a fluorogenic substrate.
36. The method of claim 35, wherein said fluorogenic substrate is selected from the group consisting of DDAO-galactopyranoside, resorufin-galactopyranoside, resorufin-glucopyranoside, DDAO-glucopyranoside, CCF2, CR2 and alike.
37. The method of claim 30, wherein said enzyme activity is selected from the group consisting of β-galactosidase, β-glucosidase, β-lactamase, and alike.
38. The method of claim 30, wherein said reporter protein comprises a substituted amino acid in place of the amino-terminus methionine residue.
39. The method of claim 38, wherein said substituted amino acid is selected from the group consisting of leucine, arginine, lysine, phenylalanine, tryptophan, and tyrosine.
40. The method of claim 30, wherein said detection is accomplished by any means of detecting a signal including visible and UV spectrometry, fluorometry, and alike.
41. A method for cell sorting, comprising the steps of: (a) obtaining a cell population in which at least one cell comprises at least one nucleotide sequence encoding a reporter protein operatively linked to said cell's genome, wherein said reporter protein when expressed has enzymatic activity and has a short cellular life time;
(b) introducing an illuminogenic substrate to said cell under conditions suitable for the conversion of said illuminogenic substrate to an illuminescent molecule, wherein said substrate enters said cell;
(c) contacting said cell population with a perturbing agent; and
(d) detecting an illuminescent signal.
41. The method of claim 40, wherein the cell population can comprise either prokaryotic cells, eukaryotic cells, or a combination of both.
42. The method of claim 41 , wherein the cell population is selected from the group consisting of prokaryotic cells and eukaryotic cells.
43. The method of claim 40, wherein said nucleotide sequence encoding for a reporter protein is operatively linked to a predetermined host cell's nucleotide sequence that encodes a specific protein.
44. The method of claim 40, wherein said reporter protein comprises an amino acid sequence selected from the group consisting of Ub-leu-β-gal, Ub-arg-β-gal and alike.
45. The method of claim 40, wherein said illuminogenic substrate is a fluorogenic substrate.
46. The method of claim 45, wherein said fluorogenic substrate is selected from the group consisting of DDAO-galactopyranoside, resorufin-galactopyranoside, resorufin-glucopyranoside, DDAO-glucopyranoside, CCF2, CR2 and alike.
47. The method of claim 40, wherein said enzyme activity is selected from the group consisting of β-galactosidase, β-glucosidase, β-lactamase, and alike.
48. The method of claim 40, wherein said reporter protein comprises a substituted amino acid in place of the amino-terminus methionine residue.
49. The method of claim 48, wherein said substituted amino acid is selected from the group consisting of leucine, arginine, lysine, phenylalanine, tryptophan, and tyrosine.
50. The method of claim 40, wherein said detection is accomplished by any means of detecting a signal including visible and UV spectrometry, fluorometry, and alike.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US45989703P | 2003-04-02 | 2003-04-02 | |
| PCT/US2004/010341 WO2004090104A2 (en) | 2003-04-02 | 2004-04-02 | Detecting gene expression in live cells using short-lived reporters with enzymatic amplification |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP1616032A2 true EP1616032A2 (en) | 2006-01-18 |
Family
ID=33159706
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP04749716A Withdrawn EP1616032A2 (en) | 2003-04-02 | 2004-04-02 | Detecting gene expression in live cells using short-lived reporters with enzymatic amplification |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP1616032A2 (en) |
| WO (1) | WO2004090104A2 (en) |
-
2004
- 2004-04-02 WO PCT/US2004/010341 patent/WO2004090104A2/en not_active Ceased
- 2004-04-02 EP EP04749716A patent/EP1616032A2/en not_active Withdrawn
Non-Patent Citations (1)
| Title |
|---|
| See references of WO2004090104A3 * |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2004090104A2 (en) | 2004-10-21 |
| WO2004090104A3 (en) | 2005-03-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US20210340527A1 (en) | Encoding of dna vector identity via iterative hybridization detection of a barcode transcript | |
| US20190249169A1 (en) | Methods for screening proteins using dna encoded chemical libraries as templates for enzyme catalysis | |
| US20100212040A1 (en) | Isolation of living cells and preparation of cell lines based on detection and quantification of preselected cellular ribonucleic acid sequences | |
| US20050070005A1 (en) | High throughput or capillary-based screening for a bioactivity or biomolecule | |
| US11795581B2 (en) | Platform for discovery and analysis of therapeutic agents | |
| JP2005501217A (en) | High-throughput or capillary-based screening for bioactivity or biomolecules | |
| US20170253938A1 (en) | Dividing of reporter proteins by dna sequences and its application in site specific recombination | |
| WO2014189768A1 (en) | Devices and methods for display of encoded peptides, polypeptides, and proteins on dna | |
| WO2002068698A2 (en) | Use of nucleic acid libraries to create toxicological profiles | |
| US20040175765A1 (en) | Cell-screening assay and composition | |
| EP1616032A2 (en) | Detecting gene expression in live cells using short-lived reporters with enzymatic amplification | |
| US20040265835A1 (en) | Method of sorting vesicle-entrapped, coupled nucleic acid-protein displays | |
| EP2824456A1 (en) | Screening for inhibitors of ribosome biogenesis | |
| AU2011203213A1 (en) | Selection and Isolation of Living Cells Using mRNA-Binding Probes | |
| JP2006061023A (en) | Screening method and screening apparatus using microchamber array | |
| AU2008202162B2 (en) | Selection and Isolation of Living Cells Using mRNA-Binding Probes | |
| KR101132052B1 (en) | DNA Selection and Detection Method Using Proteins Binding Specific DNA Base Sequences | |
| KR101104817B1 (en) | Esterase (ESTL120P), which can be used as a reporter as a fusion partner, and a method for producing a cloning vector using the indicator | |
| HK1079238B (en) | Selection and isolation of living cells using rna-binding probes | |
| HK1245340B (en) | Platform for discovery and analysis of therapeutic agents | |
| HK1146637B (en) | Selection and isolation of living cells using mrna-binding probes | |
| HK1146637A1 (en) | Selection and isolation of living cells using mrna-binding probes |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20051102 |
|
| AK | Designated contracting states |
Kind code of ref document: A2 Designated state(s): AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LI LU MC NL PL PT RO SE SI SK TR |
|
| AX | Request for extension of the european patent |
Extension state: AL HR LT LV MK |
|
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION HAS BEEN WITHDRAWN |
|
| 18W | Application withdrawn |
Effective date: 20060713 |