EP2788490A1 - Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification - Google Patents
Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplificationInfo
- Publication number
- EP2788490A1 EP2788490A1 EP12854779.1A EP12854779A EP2788490A1 EP 2788490 A1 EP2788490 A1 EP 2788490A1 EP 12854779 A EP12854779 A EP 12854779A EP 2788490 A1 EP2788490 A1 EP 2788490A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- hcv
- sequences
- virus
- seq
- sequence
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Withdrawn
Links
- 241000700605 Viruses Species 0.000 title claims description 90
- 238000013412 genome amplification Methods 0.000 title description 5
- 238000010367 cloning Methods 0.000 title description 3
- 238000000034 method Methods 0.000 claims abstract description 49
- 230000014599 transmission of virus Effects 0.000 claims abstract description 8
- 102000040430 polynucleotide Human genes 0.000 claims description 37
- 108091033319 polynucleotide Proteins 0.000 claims description 37
- 239000002157 polynucleotide Substances 0.000 claims description 37
- 108090000623 proteins and genes Proteins 0.000 claims description 36
- 229960005486 vaccine Drugs 0.000 claims description 21
- 239000000203 mixture Substances 0.000 claims description 19
- 108020000999 Viral RNA Proteins 0.000 claims description 17
- 230000002163 immunogen Effects 0.000 claims description 16
- 239000000523 sample Substances 0.000 claims description 13
- 238000012163 sequencing technique Methods 0.000 claims description 13
- 230000028993 immune response Effects 0.000 claims description 12
- 239000013610 patient sample Substances 0.000 claims description 9
- 238000002864 sequence alignment Methods 0.000 claims description 8
- 238000003559 RNA-seq method Methods 0.000 claims description 2
- 241000711549 Hepacivirus C Species 0.000 abstract description 183
- 208000015181 infectious disease Diseases 0.000 abstract description 60
- 108091028043 Nucleic acid sequence Proteins 0.000 abstract description 7
- 108090000765 processed proteins & peptides Proteins 0.000 description 70
- 102000004196 processed proteins & peptides Human genes 0.000 description 68
- 229920001184 polypeptide Polymers 0.000 description 66
- 230000003612 virological effect Effects 0.000 description 38
- 230000005540 biological transmission Effects 0.000 description 37
- 230000001154 acute effect Effects 0.000 description 35
- 239000012634 fragment Substances 0.000 description 33
- 230000035772 mutation Effects 0.000 description 30
- 238000003752 polymerase chain reaction Methods 0.000 description 29
- 150000007523 nucleic acids Chemical class 0.000 description 27
- 239000013615 primer Substances 0.000 description 27
- 239000002773 nucleotide Substances 0.000 description 25
- 125000003729 nucleotide group Chemical group 0.000 description 25
- 230000002068 genetic effect Effects 0.000 description 24
- 238000007476 Maximum Likelihood Methods 0.000 description 21
- 108091032973 (ribonucleotides)n+m Proteins 0.000 description 20
- 108020004414 DNA Proteins 0.000 description 19
- 238000004458 analytical method Methods 0.000 description 18
- 239000000047 product Substances 0.000 description 18
- 102000039446 nucleic acids Human genes 0.000 description 17
- 108020004707 nucleic acids Proteins 0.000 description 17
- 102000004169 proteins and genes Human genes 0.000 description 17
- 238000006243 chemical reaction Methods 0.000 description 16
- 238000005304 joining Methods 0.000 description 16
- 108091093088 Amplicon Proteins 0.000 description 15
- 108091034135 Vault RNA Proteins 0.000 description 15
- 210000004027 cell Anatomy 0.000 description 14
- 238000010839 reverse transcription Methods 0.000 description 13
- 230000000890 antigenic effect Effects 0.000 description 12
- 238000010804 cDNA synthesis Methods 0.000 description 12
- 238000006467 substitution reaction Methods 0.000 description 12
- 241000251539 Vertebrata <Metazoa> Species 0.000 description 11
- 150000001413 amino acids Chemical class 0.000 description 11
- 241000713772 Human immunodeficiency virus 1 Species 0.000 description 10
- 108060004795 Methyltransferase Proteins 0.000 description 10
- 230000014509 gene expression Effects 0.000 description 10
- 108020004999 messenger RNA Proteins 0.000 description 10
- 230000004048 modification Effects 0.000 description 10
- 238000012986 modification Methods 0.000 description 10
- 230000006798 recombination Effects 0.000 description 10
- 238000005215 recombination Methods 0.000 description 10
- 238000005070 sampling Methods 0.000 description 10
- 230000000692 anti-sense effect Effects 0.000 description 9
- 210000001151 cytotoxic T lymphocyte Anatomy 0.000 description 9
- 239000003814 drug Substances 0.000 description 9
- 101710144111 Non-structural protein 3 Proteins 0.000 description 8
- 230000001684 chronic effect Effects 0.000 description 8
- BASFCYQUMIYNBI-UHFFFAOYSA-N platinum Chemical compound [Pt] BASFCYQUMIYNBI-UHFFFAOYSA-N 0.000 description 8
- 238000012408 PCR amplification Methods 0.000 description 7
- 239000000427 antigen Substances 0.000 description 7
- 108091007433 antigens Proteins 0.000 description 7
- 102000036639 antigens Human genes 0.000 description 7
- 239000000872 buffer Substances 0.000 description 7
- 239000002299 complementary DNA Substances 0.000 description 7
- 238000012217 deletion Methods 0.000 description 7
- 230000037430 deletion Effects 0.000 description 7
- 238000000338 in vitro Methods 0.000 description 7
- 238000001727 in vivo Methods 0.000 description 7
- 239000013612 plasmid Substances 0.000 description 7
- 238000013518 transcription Methods 0.000 description 7
- 230000035897 transcription Effects 0.000 description 7
- 108091035707 Consensus sequence Proteins 0.000 description 6
- 102100031780 Endonuclease Human genes 0.000 description 6
- 108010092799 RNA-directed DNA polymerase Proteins 0.000 description 6
- 108010006785 Taq Polymerase Proteins 0.000 description 6
- 208000036142 Viral infection Diseases 0.000 description 6
- 230000003321 amplification Effects 0.000 description 6
- 238000009826 distribution Methods 0.000 description 6
- 238000003199 nucleic acid amplification method Methods 0.000 description 6
- 230000022532 regulation of transcription, DNA-dependent Effects 0.000 description 6
- 230000010076 replication Effects 0.000 description 6
- 108020003589 5' Untranslated Regions Proteins 0.000 description 5
- 108091026890 Coding region Proteins 0.000 description 5
- 125000003275 alpha amino acid group Chemical group 0.000 description 5
- 102000054765 polymorphisms of proteins Human genes 0.000 description 5
- 238000012360 testing method Methods 0.000 description 5
- 238000011282 treatment Methods 0.000 description 5
- 239000001226 triphosphate Substances 0.000 description 5
- 239000013598 vector Substances 0.000 description 5
- 230000029812 viral genome replication Effects 0.000 description 5
- 230000009385 viral infection Effects 0.000 description 5
- 101800001020 Non-structural protein 4A Proteins 0.000 description 4
- 108091034117 Oligonucleotide Proteins 0.000 description 4
- 101800001554 RNA-directed RNA polymerase Proteins 0.000 description 4
- 101800001838 Serine protease/helicase NS3 Proteins 0.000 description 4
- 241000713311 Simian immunodeficiency virus Species 0.000 description 4
- 238000000137 annealing Methods 0.000 description 4
- 238000004364 calculation method Methods 0.000 description 4
- 238000011161 development Methods 0.000 description 4
- 229940079593 drug Drugs 0.000 description 4
- 239000000499 gel Substances 0.000 description 4
- 230000001965 increasing effect Effects 0.000 description 4
- 238000003780 insertion Methods 0.000 description 4
- 230000037431 insertion Effects 0.000 description 4
- 230000001404 mediated effect Effects 0.000 description 4
- 230000036961 partial effect Effects 0.000 description 4
- 229910052697 platinum Inorganic materials 0.000 description 4
- 230000001105 regulatory effect Effects 0.000 description 4
- 102200111182 rs35520672 Human genes 0.000 description 4
- 230000028327 secretion Effects 0.000 description 4
- 235000011178 triphosphate Nutrition 0.000 description 4
- UNXRWKVEANCORM-UHFFFAOYSA-N triphosphoric acid Chemical compound OP(O)(=O)OP(O)(=O)OP(O)(O)=O UNXRWKVEANCORM-UHFFFAOYSA-N 0.000 description 4
- 238000012800 visualization Methods 0.000 description 4
- 206010059866 Drug resistance Diseases 0.000 description 3
- 241000282412 Homo Species 0.000 description 3
- 108010050904 Interferons Proteins 0.000 description 3
- 102000014150 Interferons Human genes 0.000 description 3
- 101800001014 Non-structural protein 5A Proteins 0.000 description 3
- 108010076504 Protein Sorting Signals Proteins 0.000 description 3
- 108010026552 Proteome Proteins 0.000 description 3
- 101710086015 RNA ligase Proteins 0.000 description 3
- 230000035508 accumulation Effects 0.000 description 3
- 238000009825 accumulation Methods 0.000 description 3
- 239000000654 additive Substances 0.000 description 3
- 238000013459 approach Methods 0.000 description 3
- 229940028617 conventional vaccine Drugs 0.000 description 3
- 238000005516 engineering process Methods 0.000 description 3
- 239000003623 enhancer Substances 0.000 description 3
- 210000003494 hepatocyte Anatomy 0.000 description 3
- 239000000137 peptide hydrolase inhibitor Substances 0.000 description 3
- 230000002688 persistence Effects 0.000 description 3
- 238000012545 processing Methods 0.000 description 3
- 230000004044 response Effects 0.000 description 3
- 239000003161 ribonuclease inhibitor Substances 0.000 description 3
- 230000007704 transition Effects 0.000 description 3
- 238000013519 translation Methods 0.000 description 3
- 239000013603 viral vector Substances 0.000 description 3
- 229920000936 Agarose Polymers 0.000 description 2
- 108020004705 Codon Proteins 0.000 description 2
- 241001465754 Metazoa Species 0.000 description 2
- 101800001019 Non-structural protein 4B Proteins 0.000 description 2
- 101800000440 Non-structural protein NS3A Proteins 0.000 description 2
- 108091093037 Peptide nucleic acid Proteins 0.000 description 2
- 241000701370 Plasmavirus Species 0.000 description 2
- 229940124158 Protease/peptidase inhibitor Drugs 0.000 description 2
- 238000002123 RNA extraction Methods 0.000 description 2
- IWUCXVSUMQZMFG-AFCXAGJDSA-N Ribavirin Chemical compound N1=C(C(=O)N)N=CN1[C@H]1[C@H](O)[C@H](O)[C@@H](CO)O1 IWUCXVSUMQZMFG-AFCXAGJDSA-N 0.000 description 2
- 108091008874 T cell receptors Proteins 0.000 description 2
- 102000016266 T-Cell Antigen Receptors Human genes 0.000 description 2
- 108020005038 Terminator Codon Proteins 0.000 description 2
- 206010058874 Viraemia Diseases 0.000 description 2
- JLCPHMBAVCMARE-UHFFFAOYSA-N [3-[[3-[[3-[[3-[[3-[[3-[[3-[[3-[[3-[[3-[[3-[[5-(2-amino-6-oxo-1H-purin-9-yl)-3-[[3-[[3-[[3-[[3-[[3-[[5-(2-amino-6-oxo-1H-purin-9-yl)-3-[[5-(2-amino-6-oxo-1H-purin-9-yl)-3-hydroxyoxolan-2-yl]methoxy-hydroxyphosphoryl]oxyoxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(5-methyl-2,4-dioxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxyoxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(5-methyl-2,4-dioxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(5-methyl-2,4-dioxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(5-methyl-2,4-dioxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(6-aminopurin-9-yl)oxolan-2-yl]methoxy-hydroxyphosphoryl]oxy-5-(4-amino-2-oxopyrimidin-1-yl)oxolan-2-yl]methyl [5-(6-aminopurin-9-yl)-2-(hydroxymethyl)oxolan-3-yl] hydrogen phosphate Polymers Cc1cn(C2CC(OP(O)(=O)OCC3OC(CC3OP(O)(=O)OCC3OC(CC3O)n3cnc4c3nc(N)[nH]c4=O)n3cnc4c3nc(N)[nH]c4=O)C(COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3COP(O)(=O)OC3CC(OC3CO)n3cnc4c(N)ncnc34)n3ccc(N)nc3=O)n3cnc4c(N)ncnc34)n3ccc(N)nc3=O)n3ccc(N)nc3=O)n3ccc(N)nc3=O)n3cnc4c(N)ncnc34)n3cnc4c(N)ncnc34)n3cc(C)c(=O)[nH]c3=O)n3cc(C)c(=O)[nH]c3=O)n3ccc(N)nc3=O)n3cc(C)c(=O)[nH]c3=O)n3cnc4c3nc(N)[nH]c4=O)n3cnc4c(N)ncnc34)n3cnc4c(N)ncnc34)n3cnc4c(N)ncnc34)n3cnc4c(N)ncnc34)O2)c(=O)[nH]c1=O JLCPHMBAVCMARE-UHFFFAOYSA-N 0.000 description 2
- 239000002253 acid Substances 0.000 description 2
- 238000007792 addition Methods 0.000 description 2
- 239000011543 agarose gel Substances 0.000 description 2
- 239000007864 aqueous solution Substances 0.000 description 2
- 230000002238 attenuated effect Effects 0.000 description 2
- 230000009286 beneficial effect Effects 0.000 description 2
- 230000008901 benefit Effects 0.000 description 2
- 230000027455 binding Effects 0.000 description 2
- 210000004369 blood Anatomy 0.000 description 2
- 239000008280 blood Substances 0.000 description 2
- 238000004422 calculation algorithm Methods 0.000 description 2
- 239000013065 commercial product Substances 0.000 description 2
- 238000012937 correction Methods 0.000 description 2
- 238000010790 dilution Methods 0.000 description 2
- 239000012895 dilution Substances 0.000 description 2
- LOKCTEFSRHRXRJ-UHFFFAOYSA-I dipotassium trisodium dihydrogen phosphate hydrogen phosphate dichloride Chemical compound P(=O)(O)(O)[O-].[K+].P(=O)(O)([O-])[O-].[Na+].[Na+].[Cl-].[K+].[Cl-].[Na+] LOKCTEFSRHRXRJ-UHFFFAOYSA-I 0.000 description 2
- 238000002474 experimental method Methods 0.000 description 2
- 230000005847 immunogenicity Effects 0.000 description 2
- 230000002458 infectious effect Effects 0.000 description 2
- 238000002347 injection Methods 0.000 description 2
- 239000007924 injection Substances 0.000 description 2
- 238000007689 inspection Methods 0.000 description 2
- 229940079322 interferon Drugs 0.000 description 2
- 229910052943 magnesium sulfate Inorganic materials 0.000 description 2
- 239000000463 material Substances 0.000 description 2
- 238000013178 mathematical model Methods 0.000 description 2
- 238000010369 molecular cloning Methods 0.000 description 2
- 230000037361 pathway Effects 0.000 description 2
- 239000002953 phosphate buffered saline Substances 0.000 description 2
- 238000013081 phylogenetic analysis Methods 0.000 description 2
- 238000002360 preparation method Methods 0.000 description 2
- 230000008569 process Effects 0.000 description 2
- 230000000644 propagated effect Effects 0.000 description 2
- 238000012175 pyrosequencing Methods 0.000 description 2
- 239000011535 reaction buffer Substances 0.000 description 2
- 108091008146 restriction endonucleases Proteins 0.000 description 2
- 238000003757 reverse transcription PCR Methods 0.000 description 2
- 230000002441 reversible effect Effects 0.000 description 2
- 229960000329 ribavirin Drugs 0.000 description 2
- HZCAHMRRMINHDJ-DBRKOABJSA-N ribavirin Natural products O[C@@H]1[C@H](O)[C@@H](CO)O[C@H]1N1N=CN=C1 HZCAHMRRMINHDJ-DBRKOABJSA-N 0.000 description 2
- 230000003248 secreting effect Effects 0.000 description 2
- 230000035945 sensitivity Effects 0.000 description 2
- 238000000926 separation method Methods 0.000 description 2
- 239000001488 sodium phosphate Substances 0.000 description 2
- 229910000162 sodium phosphate Inorganic materials 0.000 description 2
- 239000000243 solution Substances 0.000 description 2
- 238000003786 synthesis reaction Methods 0.000 description 2
- RYFMWSXOAZQYPI-UHFFFAOYSA-K trisodium phosphate Chemical compound [Na+].[Na+].[Na+].[O-]P([O-])([O-])=O RYFMWSXOAZQYPI-UHFFFAOYSA-K 0.000 description 2
- 238000002255 vaccination Methods 0.000 description 2
- 210000002845 virion Anatomy 0.000 description 2
- 102000040650 (ribonucleotides)n+m Human genes 0.000 description 1
- 208000030507 AIDS Diseases 0.000 description 1
- 102000007469 Actins Human genes 0.000 description 1
- 108010085238 Actins Proteins 0.000 description 1
- 108700028369 Alleles Proteins 0.000 description 1
- 241000271566 Aves Species 0.000 description 1
- 101710132601 Capsid protein Proteins 0.000 description 1
- 208000035473 Communicable disease Diseases 0.000 description 1
- 208000037041 Community-Acquired Infections Diseases 0.000 description 1
- 108090000695 Cytokines Proteins 0.000 description 1
- 102000004127 Cytokines Human genes 0.000 description 1
- 241000701022 Cytomegalovirus Species 0.000 description 1
- 239000003155 DNA primer Substances 0.000 description 1
- 238000001712 DNA sequencing Methods 0.000 description 1
- 108010014303 DNA-directed DNA polymerase Proteins 0.000 description 1
- 102000016928 DNA-directed DNA polymerase Human genes 0.000 description 1
- 108090000626 DNA-directed RNA polymerases Proteins 0.000 description 1
- 102000004163 DNA-directed RNA polymerases Human genes 0.000 description 1
- 208000001490 Dengue Diseases 0.000 description 1
- 206010012310 Dengue fever Diseases 0.000 description 1
- 108010016626 Dipeptides Proteins 0.000 description 1
- 102000004190 Enzymes Human genes 0.000 description 1
- 108090000790 Enzymes Proteins 0.000 description 1
- 241000710781 Flaviviridae Species 0.000 description 1
- 241000710831 Flavivirus Species 0.000 description 1
- 108700039691 Genetic Promoter Regions Proteins 0.000 description 1
- 208000031886 HIV Infections Diseases 0.000 description 1
- 102000002812 Heat-Shock Proteins Human genes 0.000 description 1
- 108010004889 Heat-Shock Proteins Proteins 0.000 description 1
- 102100021519 Hemoglobin subunit beta Human genes 0.000 description 1
- 108091005904 Hemoglobin subunit beta Proteins 0.000 description 1
- 241000711557 Hepacivirus Species 0.000 description 1
- 108091027305 Heteroduplex Proteins 0.000 description 1
- 101001002466 Homo sapiens Interferon lambda-3 Proteins 0.000 description 1
- 102100020992 Interferon lambda-3 Human genes 0.000 description 1
- 102000015696 Interleukins Human genes 0.000 description 1
- 108010063738 Interleukins Proteins 0.000 description 1
- 108020004684 Internal Ribosome Entry Sites Proteins 0.000 description 1
- 108091092195 Intron Proteins 0.000 description 1
- 241000282560 Macaca mulatta Species 0.000 description 1
- 241000829100 Macaca mulatta polyomavirus 1 Species 0.000 description 1
- 241001529936 Murinae Species 0.000 description 1
- 101000933447 Mus musculus Beta-glucuronidase Proteins 0.000 description 1
- 244000061176 Nicotiana tabacum Species 0.000 description 1
- 235000002637 Nicotiana tabacum Nutrition 0.000 description 1
- 241000283973 Oryctolagus cuniculus Species 0.000 description 1
- 241000282577 Pan troglodytes Species 0.000 description 1
- 208000037581 Persistent Infection Diseases 0.000 description 1
- 241000709664 Picornaviridae Species 0.000 description 1
- 102000009609 Pyrophosphatases Human genes 0.000 description 1
- 108010009413 Pyrophosphatases Proteins 0.000 description 1
- 102000007056 Recombinant Fusion Proteins Human genes 0.000 description 1
- 108010008281 Recombinant Fusion Proteins Proteins 0.000 description 1
- 208000035999 Recurrence Diseases 0.000 description 1
- 241000714474 Rous sarcoma virus Species 0.000 description 1
- 238000012300 Sequence Analysis Methods 0.000 description 1
- FAPWRFPIFSIZLT-UHFFFAOYSA-M Sodium chloride Chemical compound [Na+].[Cl-] FAPWRFPIFSIZLT-UHFFFAOYSA-M 0.000 description 1
- 108091027544 Subgenomic mRNA Proteins 0.000 description 1
- 206010042566 Superinfection Diseases 0.000 description 1
- 230000024932 T cell mediated immunity Effects 0.000 description 1
- 210000001744 T-lymphocyte Anatomy 0.000 description 1
- 101710136739 Teichoic acid poly(glycerol phosphate) polymerase Proteins 0.000 description 1
- 102000003978 Tissue Plasminogen Activator Human genes 0.000 description 1
- 108090000373 Tissue Plasminogen Activator Proteins 0.000 description 1
- 239000007983 Tris buffer Substances 0.000 description 1
- 241000710886 West Nile virus Species 0.000 description 1
- 241000710772 Yellow fever virus Species 0.000 description 1
- 230000021736 acetylation Effects 0.000 description 1
- 238000006640 acetylation reaction Methods 0.000 description 1
- 230000033289 adaptive immune response Effects 0.000 description 1
- 230000000996 additive effect Effects 0.000 description 1
- 239000002671 adjuvant Substances 0.000 description 1
- 238000011166 aliquoting Methods 0.000 description 1
- 230000009435 amidation Effects 0.000 description 1
- 238000007112 amidation reaction Methods 0.000 description 1
- 230000000840 anti-viral effect Effects 0.000 description 1
- 238000011225 antiretroviral therapy Methods 0.000 description 1
- 230000037429 base substitution Effects 0.000 description 1
- 239000011230 binding agent Substances 0.000 description 1
- 230000004071 biological effect Effects 0.000 description 1
- 230000000903 blocking effect Effects 0.000 description 1
- 229960000517 boceprevir Drugs 0.000 description 1
- LHHCSNFAOIFYRV-DOVBMPENSA-N boceprevir Chemical compound O=C([C@@H]1[C@@H]2[C@@H](C2(C)C)CN1C(=O)[C@@H](NC(=O)NC(C)(C)C)C(C)(C)C)NC(C(=O)C(N)=O)CC1CCC1 LHHCSNFAOIFYRV-DOVBMPENSA-N 0.000 description 1
- 108010006025 bovine growth hormone Proteins 0.000 description 1
- 230000004663 cell proliferation Effects 0.000 description 1
- 230000005889 cellular cytotoxicity Effects 0.000 description 1
- 238000005119 centrifugation Methods 0.000 description 1
- 239000007795 chemical reaction product Substances 0.000 description 1
- 239000003153 chemical reaction reagent Substances 0.000 description 1
- 210000000349 chromosome Anatomy 0.000 description 1
- 238000003776 cleavage reaction Methods 0.000 description 1
- 238000010835 comparative analysis Methods 0.000 description 1
- 238000004590 computer program Methods 0.000 description 1
- 238000011109 contamination Methods 0.000 description 1
- 230000001276 controlling effect Effects 0.000 description 1
- 230000009260 cross reactivity Effects 0.000 description 1
- 238000011461 current therapy Methods 0.000 description 1
- 208000025729 dengue disease Diseases 0.000 description 1
- 230000001419 dependent effect Effects 0.000 description 1
- 238000001212 derivatisation Methods 0.000 description 1
- 238000001514 detection method Methods 0.000 description 1
- 230000009699 differential effect Effects 0.000 description 1
- 230000029087 digestion Effects 0.000 description 1
- 239000003085 diluting agent Substances 0.000 description 1
- 229940042399 direct acting antivirals protease inhibitors Drugs 0.000 description 1
- 230000008034 disappearance Effects 0.000 description 1
- 239000000890 drug combination Substances 0.000 description 1
- 230000000694 effects Effects 0.000 description 1
- 230000008030 elimination Effects 0.000 description 1
- 238000003379 elimination reaction Methods 0.000 description 1
- 239000000839 emulsion Substances 0.000 description 1
- 244000309457 enveloped RNA virus Species 0.000 description 1
- 210000003527 eukaryotic cell Anatomy 0.000 description 1
- 230000007717 exclusion Effects 0.000 description 1
- 230000001747 exhibiting effect Effects 0.000 description 1
- 238000000605 extraction Methods 0.000 description 1
- -1 for example Substances 0.000 description 1
- 238000009472 formulation Methods 0.000 description 1
- 230000003485 founder effect Effects 0.000 description 1
- 230000037433 frameshift Effects 0.000 description 1
- 108020001507 fusion proteins Proteins 0.000 description 1
- 102000037865 fusion proteins Human genes 0.000 description 1
- 238000001415 gene therapy Methods 0.000 description 1
- 230000013595 glycosylation Effects 0.000 description 1
- 238000006206 glycosylation reaction Methods 0.000 description 1
- 230000036541 health Effects 0.000 description 1
- 244000052637 human pathogen Species 0.000 description 1
- 230000028996 humoral immune response Effects 0.000 description 1
- 230000004727 humoral immunity Effects 0.000 description 1
- 230000001900 immune effect Effects 0.000 description 1
- 210000000987 immune system Anatomy 0.000 description 1
- 238000002649 immunization Methods 0.000 description 1
- 230000003053 immunization Effects 0.000 description 1
- 230000002998 immunogenetic effect Effects 0.000 description 1
- 230000002779 inactivation Effects 0.000 description 1
- 230000006698 induction Effects 0.000 description 1
- 230000001939 inductive effect Effects 0.000 description 1
- 230000015788 innate immune response Effects 0.000 description 1
- 229940047124 interferons Drugs 0.000 description 1
- 229940047122 interleukins Drugs 0.000 description 1
- 229940065638 intron a Drugs 0.000 description 1
- 230000000670 limiting effect Effects 0.000 description 1
- 230000033001 locomotion Effects 0.000 description 1
- 210000004698 lymphocyte Anatomy 0.000 description 1
- 210000004962 mammalian cell Anatomy 0.000 description 1
- 238000007726 management method Methods 0.000 description 1
- 238000004519 manufacturing process Methods 0.000 description 1
- 238000005259 measurement Methods 0.000 description 1
- 108091070501 miRNA Proteins 0.000 description 1
- 239000002679 microRNA Substances 0.000 description 1
- 230000003278 mimic effect Effects 0.000 description 1
- 238000002703 mutagenesis Methods 0.000 description 1
- 231100000350 mutagenesis Toxicity 0.000 description 1
- 230000036438 mutation frequency Effects 0.000 description 1
- 230000000869 mutational effect Effects 0.000 description 1
- 238000007857 nested PCR Methods 0.000 description 1
- 230000003472 neutralizing effect Effects 0.000 description 1
- 239000002547 new drug Substances 0.000 description 1
- 230000009871 nonspecific binding Effects 0.000 description 1
- 230000008506 pathogenesis Effects 0.000 description 1
- 230000001717 pathogenic effect Effects 0.000 description 1
- 230000007918 pathogenicity Effects 0.000 description 1
- 230000026731 phosphorylation Effects 0.000 description 1
- 238000006366 phosphorylation reaction Methods 0.000 description 1
- 210000002381 plasma Anatomy 0.000 description 1
- 230000004481 post-translational protein modification Effects 0.000 description 1
- 230000001323 posttranslational effect Effects 0.000 description 1
- 239000003755 preservative agent Substances 0.000 description 1
- 230000037452 priming Effects 0.000 description 1
- 230000002797 proteolythic effect Effects 0.000 description 1
- 230000006337 proteolytic cleavage Effects 0.000 description 1
- 238000000746 purification Methods 0.000 description 1
- 238000003908 quality control method Methods 0.000 description 1
- 239000011541 reaction mixture Substances 0.000 description 1
- 230000002829 reductive effect Effects 0.000 description 1
- 230000001850 reproductive effect Effects 0.000 description 1
- 238000011160 research Methods 0.000 description 1
- 230000004043 responsiveness Effects 0.000 description 1
- 238000012552 review Methods 0.000 description 1
- 210000003935 rough endoplasmic reticulum Anatomy 0.000 description 1
- 230000007017 scission Effects 0.000 description 1
- 238000011451 sequencing strategy Methods 0.000 description 1
- 210000002966 serum Anatomy 0.000 description 1
- 239000003381 stabilizer Substances 0.000 description 1
- 230000004936 stimulating effect Effects 0.000 description 1
- 230000035892 strand transfer Effects 0.000 description 1
- 239000000725 suspension Substances 0.000 description 1
- 229960002935 telaprevir Drugs 0.000 description 1
- 108010017101 telaprevir Proteins 0.000 description 1
- BBAWEDCPNXPBQM-GDEBMMAJSA-N telaprevir Chemical compound N([C@H](C(=O)N[C@H](C(=O)N1C[C@@H]2CCC[C@@H]2[C@H]1C(=O)N[C@@H](CCC)C(=O)C(=O)NC1CC1)C(C)(C)C)C1CCCCC1)C(=O)C1=CN=CC=N1 BBAWEDCPNXPBQM-GDEBMMAJSA-N 0.000 description 1
- 230000002123 temporal effect Effects 0.000 description 1
- 229940124597 therapeutic agent Drugs 0.000 description 1
- 229940126585 therapeutic drug Drugs 0.000 description 1
- 230000001225 therapeutic effect Effects 0.000 description 1
- 238000002560 therapeutic procedure Methods 0.000 description 1
- 229940021747 therapeutic vaccine Drugs 0.000 description 1
- 210000001519 tissue Anatomy 0.000 description 1
- 229960000187 tissue plasminogen activator Drugs 0.000 description 1
- 230000005030 transcription termination Effects 0.000 description 1
- 230000014621 translational initiation Effects 0.000 description 1
- 238000013520 translational research Methods 0.000 description 1
- LENZDBCJOHFCAS-UHFFFAOYSA-N tris Chemical compound OCC(N)(CO)CO LENZDBCJOHFCAS-UHFFFAOYSA-N 0.000 description 1
- 241001430294 unidentified retrovirus Species 0.000 description 1
- 238000011870 unpaired t-test Methods 0.000 description 1
- 210000002700 urine Anatomy 0.000 description 1
- 230000008957 viral persistence Effects 0.000 description 1
- 230000006394 virus-host interaction Effects 0.000 description 1
- XLYOFNOQVPJJNP-UHFFFAOYSA-N water Substances O XLYOFNOQVPJJNP-UHFFFAOYSA-N 0.000 description 1
- 229940051021 yellow-fever virus Drugs 0.000 description 1
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N7/00—Viruses; Bacteriophages; Compositions thereof; Preparation or purification thereof
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/70—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving virus or bacteriophage
- C12Q1/701—Specific hybridization probes
- C12Q1/706—Specific hybridization probes for hepatitis
- C12Q1/707—Specific hybridization probes for hepatitis non-A, non-B Hepatitis, excluding hepatitis D
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
- A61K39/00—Medicinal preparations containing antigens or antibodies
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2770/00—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA ssRNA viruses positive-sense
- C12N2770/00011—Details
- C12N2770/24011—Flaviviridae
- C12N2770/24211—Hepacivirus, e.g. hepatitis C virus, hepatitis G virus
- C12N2770/24221—Viruses as such, e.g. new isolates, mutants or their genomic sequences
Definitions
- the invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection.
- the invention provides methods for identifying nucleotide sequences corresponding to specific transmitted HCV genomes, including structural gene nucleotide sequences of global HCV genotypes, subtypes, and drug resistant variants that actively mediate viral infection.
- the invention further provides methods of administering a vaccine comprising transmitted full-length hepatitis C virus (HCV) genomes or portions thereof.
- Hepatitis C Virus is a positive strand, non-segmented, enveloped RNA virus of approximately 9.6 kb in length.
- the virus is classified in the genus Hepacivirus within the larger family of Flavivirus, which includes the human pathogens West Nile virus, yellow fever virus and dengue fever virus among others.
- a common feature among the Flaviviridae is their dependence on a virally-encoded RNA-dependent RNA polymerase (RdRp) for replication.
- RdRp is error-prone, and HCV is notable for its quasispecies complexity and broad genotypic diversity. Globally, there are seven major genotypes of HCV that differ by approximately 30- 35% in nucleotide sequence.
- HCV The extraordinary diversity of HCV has important implications for clinical management and for basic and translational research aimed at elucidating viral natural history, pathogenesis and susceptibility to novel therapeutics and vaccines.
- DAA direct acting antiviral
- the development of new drugs and drug combinations that effectively suppress HCV replication and prevent the emergence of DAA resistance is a major challenge given the high rates of virus replication and variation.
- HCV diversity poses similar challenges to the development of effective vaccines and to the elucidation of virus biology, immunopathogenesis, and gene structure-function relationships.
- Acute HCV infection defined as the period between virus transmission and antibody seroconversion which occurs about 6 - 10 weeks later, sets in motion viral-host interactions that largely dictate the natural history of the infection.
- viral genotype and host immunogenetic factors most importantly IL28B alleles, a proportion of newly infected individuals spontaneously control or eliminate the virus.
- IL28B alleles a proportion of newly infected individuals spontaneously control or eliminate the virus.
- a greater number of patients can be cured if the infection is treated with interferon and ribavirin alone or in combination with DAA drugs. Mechanistically, how this occurs is unknown, and how the emergence of DAA drug resistance in communities of chronically infected subjects or in acutely infected subjects who initiate early therapy will affect treatment responses is uncertain, but again there are parallels with HIV-1.
- the invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection.
- HCV hepatitis C virus
- One application of this invention is to use these transmitted full-length genomes for development of more effective vaccines and treatments.
- the transmitted full-length HCV genome comprises the polynucleotide of SEQ ID NO: 776.
- the HCV genome mediates viral transmission.
- the polypeptide sequence comprises env and core genes. Nucleotide sequence of transmitted HCV genomes is set forth in Table 5.
- the invention provides methods for identifying full-length transmitted HCV genomes, the methods comprising collecting a patient sample, isolating and preparing viral RNA for sequencing, sequencing viral RNA that includes HCV genomes of circulating virus, performing sequence alignment of selected HCV genome regions, analyzing phylogenetically selected sequence alignments; and identifying full-length HCV genomes of transmitted virus.
- the HCV genome polypeptide sequence comprises SEQ ID NO. 776.
- the invention provides an immunogenic composition comprising transmitted full-length HCV genome or portions thereof.
- the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
- the invention provides methods for administering vaccines comprising transmitted full-length HCV genomes or portions thereof, wherein an immune response is induced in the patient following vaccination.
- the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
- Figure 1 illustrates HCV R A kinetics in 16 acute infection subjects.
- the shaded area represents that the HCV RNA is below the linear range of quantitation (43-69,000,000 IU/ml). Plus signs indicate the samples are Anti-HCV antibody positive.
- the circles represent the time points that were sequence analyzed from each subject.
- Figure 2 is a Maximum-likelihood tree (ML) of 5' quarter genome ⁇ core, El and E2) sequences from 17 acutely and 14 chronically infected subjects. The sequences from acutely infected subjects are shown in red and the sequences from chronically infected subjects are shown in blue. The reference sequences of genotype 1 to 7 are in gray. Bootstrap values (>70%) are indicated for intra-subject clusters. The horizontal scale bar represents 3% genetic distance.
- ML Maximum-likelihood tree
- Figure 3 illustrates 5' half-genome sequence diversity in two subjects, one with chronic infection (WIMI4025 (Fig 3 A)) and one with acute infection (10051 Fig 3(B)). Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. Bootstrap values (>70%) are indicated for intra-linerage clusters. The horizontal scale bars represent 0.5% (24.5 nt) and 0.02% (1 nt) genetic distance for subject WIMI and 10051, respectively.
- Figure 4 illustrates 5' quarter-genome ⁇ Core, El and E2) sequence diversity in four acutely infected subjects. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. Sequences from multiple time points are color coded with orange, green, blue and black in chronological order. Sequences from subject 10021 (Fig 4A) and 10025 (Fig 4B) each showed productive infection by a single virus. Sequences from subject 10012 (Fig 4C) and 10062 (Fig 4D) showed productive infection by at least three viruses, repectively. The horizotal scale bars represent 0.04% (1 nt) genetic distance for Figs. 4A, 4B and 4D and 0.4% (10 nt) for 4C. Bootstrap values (>70%) are indicated for within patient intra-lineage clusters in the ML tree.
- Figure 5 illustrates 5' quarter-genome ⁇ Core, El and E2) sequence divsersity in an acutely infected subject. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the ML tree and highlighter plot of 5' quarter genome ⁇ Core, El and E2) sequences generated from multiple time points showed productive infection by at least 9 transmitted/founder variants.
- the close genetic diversity between variant 1 and 2, variant 6 and 7 is indicative of transmission from an acutely infected donor.
- the horizotal scale bars represent 0.2% (5 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra- lineage clusters.
- Figure 6 illustrates HCV diversity in acutely infected subject 10024. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the ML tree and highlighter plot of 5' quarter genome sequences derived from multiple time points showed productive infection by at least 7 transmitted/founder variants.
- the horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
- Figure 7 illustrates HCV diversity in acutely infected subject 6123. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 3 transmitted/founder variants.
- the horizotal scale bars represent 0.2% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
- Figure 8 illustrates HCV diversity in acutely infected subject 6222. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 4 transmitted/founder variants.
- the horizotal scale bars represent 0.02% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
- Figure 9 illustrates HCV diversity in acutely infected subject 10004. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the ML tree and highlighter plot of the 5' quarter genome sequences showed productive infection by at least 3 transmitted/founder variants.
- the horizotal scale bars represent 0.36% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
- Figure 10 illustrates HCV diversity in acutely infected subject 10002. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 13 transmitted/founder variants.
- the horizotal scale bars represent 0.2% (10 nt) genetic distance.
- Bootstrap values (>70%) are indicated for patient intra- lineage clusters.
- Figure 11 illustrates HCV diversity in acutely infected subject 10017. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the ML tree and highlighter plot of 5' quarter genome sequences derived from multiple time points showed productive infection by at least 6 transmitted/founder variants.
- the horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
- Figure 12 illustrates HCV diversity in acutely infected subject 9055. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' half genome sequences derived from multiple time points showed productive infection by a single transmitted/founder variants. The horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
- Figure 13 illustrates 5' quarter-genome (Core, El and E2) sequence divsersity in acute- to-acute transmission.
- 5' quarter genome (Core, El and E2) sequences from subject 10020 (Fig 13 A) and 10016 (Fig 13B), and 5' half genome (core, El, E2, p7, NS2 and NS3A) sequences from subject 10003 (Fig 13C) are depicted by ML trees and by highlighter plots. Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the first sequence in red is the consensus sequence derived from the entire sequence dataset.
- the horizotal scale bars represent 0.04%> (1 nt), 0.036%) (1 nt) and 0.02%> (1 nt) genetic distance for Figs. 13 A, 13B and 13C, repectively. Bootstrap values (>70%) are indicated.
- Figure 14 illustrates HCV RNA kinectics and diversity in subject 106889.
- the inset shows the HCV RNA kinetics.
- the time point that was sequence analyzed is circled.
- Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
- the 5' half genome (core, El, E2, p7, NS2 and NS3A) sequences formed 10 distinct lineages (labeled with letter L) with high statistical support in the ML tree. Each cluster is color coded.
- the T/F variants was identified within each lineage. A total of 33 T/F variants were found that were responsible for productive infection.
- the horizotal scale bars represent 0.04%) (1 nt) genetic distance. Bootstrap values (>70%>) are indicated.
- Figure 15 illustrates strand transfers at stem loop structures of HCV RNA.
- Figures 17A- E illustrate four examples of stem loop structures in double-stranded HCV RNA leading to template switching by the HCV RNA polymerase.
- Fig 17A SEQ ID NO: 800 First 5' to 3' horizontal sequence; SEQ ID NO: 801 Second 3' to 5' horizontal sequence; SEQ ID NO: 802 Third 5' to 3' horizontal sequence; SEQ ID NO: 800 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 802 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 801 Sixth 3' to 5' haripin sequence.
- Fig 17B SEQ ID NO: 803 First 5' to 3' horizontal sequence; SEQ ID NO: 804 Second 3' to 5' horizontal sequence; SEQ ID NO: 805 Third 5' to 3' horizontal sequence; SEQ ID NO: 803 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 805 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 804 Sixth 3' to 5' hairpin sequence.
- Fig 17C SEQ ID NO: 806 First 5' to 3' horizontal sequence; SEQ ID NO: 807 Second 3' to 5' horizontal sequence; SEQ ID NO: 808 Third 5' to 3' horizontal sequence; SEQ ID NO: 806 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 808 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 807 Sixth 3' to 5' hairpin sequence.
- Fig 17D SEQ ID NO: 809 First 5' to 3' horizontal sequence; SEQ ID NO: 810 Second 3' to 5' horizontal sequence; SEQ ID NO: 811 Third 5' to 3' horizontal sequence; SEQ ID NO: 809 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 811 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 810 Sixth 3' to 5' hairpin sequence.
- Fig 17E SEQ ID NO: 812 First 5' to 3' horizontal sequence; SEQ ID NO: 813 Second 3' to 5' horizontal sequence; SEQ ID NO: 814 Third 5' to 3' horizontal sequence; SEQ ID NO: 812 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 814 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 813 Sixth 3' to 5' hairpin sequence.
- Figure 16 illsutrates the strategy for identification of full-length HCV genome by single gneome amplification (SGA).
- the complete HCV genome of subject 10021 was determined by amplifying five overlapping genome fragments: fragment I (nt 1 - 852) contains the complete 5' UTR and partial Core; fragment II (nt 391 - 5297) contains Core, El, E2, NS2 and NS3; fragment III (nt 5168 - 9374) contains NS4A, NS4B, NS5A and NS5B; fragment IV (nt 9082 - 9582) contains partial NS5B, variable region and the complete poly U/UC tract; and fragment V (nt 9563 - 9646) contains the complete x-tail.
- the nucleotide positions are based reference sequence H77.
- Figure 17 illustrates the strategy for generating a full-length T/F HCV molcular clone from subject 10021.
- Figure 18 illustrates HCV diversity in subject 10003.
- ML tree and Highlighter plot of 5' halfgenome sequences reveal many sets of closely related sequences distinguished by unique shared mutations.
- Figure 19 illustrates HCV diversity analysis in subject 10016 suggests acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 5' quarter 1 genome sequences. Visualization of 15 potential T/F viral lineages distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveal 10 and 4 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
- Figure 20 illustrates maximum diversity of discrete HCV sequence lineages from acute infection subjects versus maximum sequence diversity in chronic subjects.
- Primary data are derived from Tables 1 and 2.
- Mean ( ⁇ 95% CI) values are represented by horizontal lines. Differences between the two groups were highly significant (p ⁇ 0.0001; unpaired T-test with Welch's correction), reflecting the recent and remote diversification histories of acute and chronic sequences, respectively.
- Figure 21 shows HCV diversity in acute subject 10051. 5' quarter 1 genomesequences are color coded in orange, green, blue and black in chronological order to reflect sampling time points in Figure 1 and are represented in a ML tree and Highlighter plot. Sequences show evidence of productive clinical infection by a single virus. The horizontal scale bar indicates genetic distance.
- Figure 22 illustrates nonsynonymous and synonymous mutations in HCV sequences from acute subjects 10003, 10020 and 10016. Highlighter plotsof 5' half or quarter 1 genomesequences are color coded to denote nonynonymous (red) and synonymous (green) mutations for subjects 10003 (panel A), 10020 (panel B) and 10016 (panel C).
- Figure 23 illustrates amino acid alignment of the HCV Env coding region of acute subject 10003.
- the H77 reference sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
- Figure 24 illustrates amino acid alignment of the HCV Env coding region of acute subject 10020.
- the H77 reference sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
- Figure 25 illustrates amino acid alignment of the HCV Env coding region of acute subject 10016.
- the H77 sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
- Figure 26 shows HCV diversity analysis in subject 10020 to suggest acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 5' quarter 1 genome sequences. Visualization of 10 potential T/F viral sequences distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveals 5 and 3 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
- Figure 27 shows HCV diversity analysis in subject 10003 to suggest acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 59 half genome sequences. Visualization of 37 potential T/F viral sequences distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveals 15 and 8 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
- Figure 28 illustrates nonsynonymous and synonymous mutations in HCV sequences from acute subject 9055. A Highlighter plot (panel A) of 5' half genomesequences is color coded to denote nonynonymous (red) and synonymous (green) mutations.
- the boxed area reveals a temporal expansion of sequences with concentrated amino acid subtitutions in NS3.
- amino acid selection is evident in a previously identified CTL epitope highlighted in red.
- the top-most sequence represents the genotype 3a consensus.
- the invention provides full-length transmitted HCV genomes that mediate viral infection and transmission. Prior to the present invention the identification of these genomes was unavailable due to the following: (i) a significant time period from weeks to months between the moment of transmission and the first appearance of HCV in the blood (Bowen. D.G. and Walker, CM., 2005, Nature 436: 946-952; Moradpour, D. et al, 2007, Nature Rev Micro 5:453- 463); (ii) HCV is genetically highly variable in its nucleotide sequence due to its error-prone RNA-dependent RNA polymerase, and as a result, it exists in individuals as a complex mixture of sequences commonly referred to as a 'quasispecies' (Moradpour, D.
- polynucleotide is intended to encompass a singular nucleic acid or nucleic acid fragment as well as plural nucleic acids or nucleic acid fragments, and refers to an isolated molecule or construct, e.g., a virus genome (e.g., vRNA), messenger RNA (mRNA), plasmid DNA (pDNA), or derivatives of pDNA (e.g., minicircles as described in (Darquet, A-M et al., 1997, Gene Therapy, 4: 1341-1349) comprising a polynucleotide.
- virus genome e.g., vRNA
- mRNA messenger RNA
- pDNA plasmid DNA
- derivatives of pDNA e.g., minicircles as described in (Darquet, A-M et al., 1997, Gene Therapy, 4: 1341-1349
- a polynucleotide may comprise a conventional phosphodiester bond or a non-conventional bond (e.g., an amide bond, such as found in peptide nucleic acids (PNA)).
- PNA peptide nucleic acids
- nucleic acid or nucleic acid fragment refer to any one or more nucleic acid segments, e.g., DNA or R A fragments, present in a polynucleotide or construct.
- a nucleic acid or fragment thereof may be provided in linear (e.g., mR A) or circular (e.g., plasmid) form as well as double-stranded or single-stranded forms.
- isolated nucleic acid or polynucleotide is intended a nucleic acid molecule, DNA or RNA, which has been removed from its native environment.
- a recombinant polynucleotide contained in a vector is considered isolated or “cloned” for the purposes of the present invention.
- Further examples of an isolated polynucleotide include recombinant polynucleotides maintained in heterologous host cells or purified (partially or substantially) polynucleotides in solution.
- Isolated RNA molecules include in vivo or in vitro RNA transcripts of the polynucleotides of the present invention. Isolated polynucleotides or nucleic acids according to the present invention further include such molecules produced synthetically.
- fragment when referring to HCV polypeptides of the present invention include any polypeptides which retain at least some of the immunogenicity or antigenicity of the corresponding native polypeptide. Fragments of HCV polypeptides of the present invention include proteolytic fragments, deletion fragments and in particular, fragments of HCV polypeptides which exhibit increased secretion from the cell or higher immunogenicity or reduced pathogenicity when delivered to an animal. Polypeptide fragments further include any portion of the polypeptide which comprises an antigenic or immunogenic epitope of the native polypeptide, including linear as well as three-dimensional epitopes.
- Variants of HCV polypeptides of the present invention include fragments, and also polypeptides with altered amino acid sequences due to amino acid substitutions, deletions, or insertions. Variants may occur naturally, such as an allelic variant.
- allelic variant is intended alternate forms of a gene occupying a given locus on a chromosome or genome of an organism or virus. Genes II, Lewin, B., ed., John Wiley & Sons, New York (1985), which is incorporated herein by reference.
- variations in a given gene product is a "variant".
- Naturally or non-naturally occurring variations such as amino acid deletions, insertions or substitutions may occur.
- Non-naturally occurring variants may be produced using art-known mutagenesis techniques.
- Variant polypeptides may comprise conservative or non-conservative amino acid substitutions, deletions or additions.
- Derivatives of HCV polypeptides of the present invention are polypeptides which have been altered so as to exhibit additional features not found on the native polypeptide. Examples include fusion proteins.
- An analog is another form of an HCV polypeptide of the present invention.
- An example is a proprotein which can be activated by cleavage of the proprotein to produce an active mature polypeptide.
- the polynucleotide, nucleic acid, or nucleic acid fragment is DNA.
- a polynucleotide comprising a nucleic acid which encodes a polypeptide normally also comprises a promoter and/or other transcription or translation control elements operably associated with the polypeptide-encoding nucleic acid fragment.
- An operable association is when a nucleic acid fragment encoding a gene product, e.g., a polypeptide, is associated with one or more regulatory sequences in such a way as to place expression of the gene product under the influence or control of the regulatory sequence(s).
- Two DNA fragments are "operably associated” if induction of promoter function results in the transcription of mRNA encoding the desired gene product and if the nature of the linkage between the two DNA fragments does not (1) result in the introduction of a frame-shift mutation, (2) interfere with the ability of the expression regulatory sequences to direct the expression of the gene product, or (3) interfere with the ability of the DNA template to be transcribed.
- a promoter region would be operably associated with a nucleic acid fragment encoding a polypeptide if the promoter was capable of effecting transcription of that nucleic acid fragment.
- the promoter may be a cell-specific promoter that directs substantial transcription of the DNA only in predetermined cells.
- Other transcription control elements besides a promoter, for example enhancers, operators, repressors, and transcription termination signals, can be operably associated with the polynucleotide to direct cell-specific transcription. Suitable promoters and other transcription control regions are disclosed herein.
- transcription control regions are known to those skilled in the art. These include, without limitation, transcription control regions which function in vertebrate cells, such as, but not limited to, promoter and enhancer segments from cytomegaloviruses (the immediate early promoter, in conjunction with intron-A), simian virus 40 (the early promoter), and retroviruses (such as Rous sarcoma virus).
- Other transcription control regions include those derived from vertebrate genes such as actin, heat shock protein, bovine growth hormone and rabbit ⁇ -globin, as well as other sequences capable of controlling gene expression in eukaryotic cells. Additional suitable transcription control regions include tissue-specific promoters and enhancers as well as lymphokine-inducible promoters (e.g., promoters inducible by interferons or interleukins).
- a DNA polynucleotide of the present invention may be a circular or linearized plasmid or vector, or other linear DNA which may also be non-infectious and nonintegrating (i.e., does not integrate into the genome of vertebrate cells).
- a linearized plasmid is a plasmid that was previously circular but has been linearized, for example, by digestion with a restriction endonuclease.
- RNA for example, in the form of messenger RNA (mRNA).
- mRNA messenger RNA
- Polynucleotides, nucleic acids, and nucleic acid fragments of the present invention may be associated with additional nucleic acids which encode secretory or signal peptides, which direct the secretion of a polypeptide encoded by a nucleic acid fragment or polynucleotide of the present invention.
- proteins secreted by mammalian cells have a signal peptide or secretory leader sequence which is cleaved from the mature protein once export of the growing protein chain across the rough endoplasmic reticulum has been initiated.
- polypeptides secreted by vertebrate cells generally have a signal peptide fused to the N-terminus of the polypeptide, which is cleaved from the complete or "full length" polypeptide to produce a secreted, or "mature” form of the polypeptide.
- the native leader sequence is used, or a functional derivative of that sequence that retains the ability to direct the secretion of the polypeptide that is operably associated with it.
- a heterologous mammalian leader sequence, or a functional derivative thereof may be used.
- the wild-type leader sequence may be substituted with the leader sequence of human tissue plasminogen activator (TP A) or mouse beta-glucuronidase.
- RNA messenger-RNA
- polypeptide is intended to encompass a singular "polypeptide” as well as plural “polypeptides,” and comprises any chain or chains of two or more amino acids.
- polypeptide terms including, but not limited to “peptide,” “dipeptide,” “tripeptide,” “protein,” “amino acid chain,” or any other term used to refer to a chain or chains of two or more amino acids, are included in the definition of a “polypeptide,” and the term “polypeptide” can be used instead of, or interchangeably with any of these terms.
- the term further includes polypeptides which have undergone post-translational modifications, for example, glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting/blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids.
- polypeptides of the present invention are fragments, derivatives, analogs, or variants of the foregoing polypeptides, and any combination thereof.
- Polypeptides, and fragments, derivatives, analogs, or variants thereof of the present invention can be antigenic and immunogenic polypeptides related to HCV polypeptides, which are used to prevent or treat, i.e., cure, ameliorate, lessen the severity of, or prevent or reduce contagion of infectious disease caused by the HCV.
- an "antigenic polypeptide” or an “immunogenic polypeptide” is a polypeptide which, when introduced into a vertebrate, reacts with the vertebrate's immune system molecules, i.e., is antigenic, and/or induces an immune response in the vertebrate, i.e., is immunogenic. It is quite likely that an immunogenic polypeptide will also be antigenic, but an antigenic polypeptide, because of its size or conformation, may not necessarily be immunogenic.
- Isolated antigenic and immunogenic polypeptides of the present invention in addition to those encoded by polynucleotides of the invention, may be provided as a recombinant protein, a purified subunit, a viral vector expressing the protein, or may be provided in the form of an inactivated HCV vaccine, e.g., a live-attenuated virus vaccine, a heat-killed virus vaccine, etc.
- an "isolated" HCV polypeptide or a fragment, variant, or derivative thereof is intended an HCV polypeptide or protein that is not in its natural form. No particular level of purification is required.
- an isolated HCV polypeptide can be removed from its native or natural environment.
- HCV polypeptides and proteins expressed in host cells are considered isolated for purposed of the invention, as are native or recombinant HCV polypeptides which have been separated, fractionated, or partially or substantially purified by any suitable technique, including the separation of HCV virions from culture cells in which they have been propagated.
- an isolated HCV polypeptide or protein can be provided as a live or inactivated viral vector expressing an isolated HCV polypeptide and can include those found in inactivated HCV vaccine compositions.
- isolated HCV polypeptides and proteins can be provided as, for example, recombinant HCV polypeptides, a purified subunit of HCV, a viral vector expressing an isolated HCV polypeptide, or in the form of an inactivated or attenuated HCV vaccine.
- epitopes refers to portions of a polypeptide having antigenic or immunogenic activity in a vertebrate, for example a human.
- An "immunogenic epitope,” as used herein, is defined as a portion of a protein that elicits an immune response in an animal, as determined by any method known in the art.
- antigenic epitope is defined as a portion of a protein to which an antibody or T-cell receptor can immunospecifically bind as determined by any method well known in the art. Immunospecific binding excludes nonspecific binding but does not exclude cross-reactivity with other antigens. Where all immunogenic epitopes are antigenic, antigenic epitopes need not be immunogenic.
- peptides or polypeptides bearing an antigenic epitope e.g., that contain a region of a protein molecule to which an antibody or T cell receptor can bind
- relatively short synthetic peptides that mimic part of a protein sequence are routinely capable of eliciting an antiserum that reacts with the partially mimicked protein. See, e.g., Sutcliffe, J. G., et al., 1983, Science 219:660-666.
- the present invention also provides methods for identifying transmitted full-length HCV genomes that mediate viral infection and transmission.
- the methods comprise: (a) collecting a patient sample; (b) isolating viral RNA from said sample; (c) sequencing said viral RNA, wherein viral RNA sequencing includes HCV genomes of circulating virus; (d) performing sequence alignment of selected HCV genome regions; (e) analyzing phylogenetically selected sequence alignments; and (f) identifying HCV genomes of transmitted virus.
- a "patient” or “subject” to be utilized by the disclosed methods can mean a human, chimpanzee or non-human primate. In certain embodiments the patient is a non-human mammal.
- patient sample includes but is not limited to a blood, serum, plasma, or urine sample obtained from a patient.
- the patient sample is plasma.
- analyzing phylogentically represents the analysis of clinical viral isolates, regardless of the particular methodology employed. Comparative analysis of the genetic relatedness of any a collection of circulating viral isolates is used to select for nucleotide sequence of transmitted HCV genomes. This methodology is used to select for the actual HCV genomes present at the time of patient infection/viral transmission. This can be accomplished by any number of methods, including but not limited to: i) a novel mathematical model of random virus evolution disclosed herein; ii) star phylogeny; iii) Baysian analysis; or any other method of phylogenetic analysis known to one of skill in the art.
- performing sequence alignment is meant to include aligning genetic sequences by any of a number of different procedures that produce a sufficient match between the corresponding residue in the sequences. Typically, Smith- Waterman or Needleman- Wunsch algorithms are used. However, other procedures such as BLAST, FASTA, PSI-BLAST can be used.
- isolated viral RNA includes those methods well known in the art for extraction of viral RNA from a sample, cDNA synthesis from viral RNA template, optionally cloning of cDNA fragments, and/or amplification of polynucleotide sequence. Such methods are described in more detail, but such experimental procedures are provided in Sambrook and Russell, 2001, Molecular Cloning: A Laboratory Manual, Third Ed., Cold Spring Harbor Laboratory Press, Woodbury, N.Y.
- nucleotide sequence amplification such as polymerase chain reaction (PCR) and modifications thereof (including for example reverse transcription (RT)-PCR, and stem-loop PCR), as well as reverse transcription and in vitro transcription.
- PCR polymerase chain reaction
- RT reverse transcription
- these methods utilize one or a pair of oligonucleotide primers having sequence complimentary to sequences 5' and 3' to the sequence of interest, and in the use of these primers they are hybridized to a nucleotide sequence and extended during the practice of PCR amplification using DNA polymerase (preferably using a thermal-stable polymerase such as Taq polymerase).
- RT-PCR may be performed on miRNA or mRNA with a specific 5' primer or random primers and appropriate reverse transcription enzymes such as avian (AMV-RT) or murine (MMLV-RT) reverse transcriptase enzymes.
- AMV-RT avian
- MMLV-RT murine reverse transcriptase enzymes.
- over time represents a period of time between the collection of patient samples. For example, a patient sample is collected at time point A and then subsequently as a later date at time point B. The period between samplings is over time. In certain embodiments, samples are taken from the same patient. In alternative embodiments, samples may be taken from different patients. Performing the method of the invention at different time points followed by the differential analysis of HCV genomes at the different time points provides a means for assessing the evolution of HCV genomes both inter- and intra- patient. The identification of highly variable genomic regions is useful for the generation of effective therapeutics and vaccines to HCV.
- transmitted refers to viral genotype and phenotype at the time of viral infection.
- transmitted virus as used here is in reference to actual HCV virus that mediates patient infection.
- transmitted viral sequence as used herein is meant to include nucleotide or amino acid sequence of HCV genomes of transmitted virus, an in a particular embodiment sequence corresponding to acute infection stage HCV.
- circulating refers to HCV virus collected from patient post-infection. In general, patient samples comprise circulating virus because HCV infection has occurred at a prior time point. The half-life of plasma virus is less than 1 day, thus the collection of transmitted virus from a patient sample would be a rare event.
- selected as used in the phrases “selected sequence alignments” or “selected HCV genome regions” is meant to include the identification and utilization of particular subsets of nucleotide sequence and/or genome regions for subsequent analysis. In certain embodiments, full-length HCV genomes are identified, however discrete portions of the genome are utilized or “selected” for further analysis.
- a or “an” entity refers to one or more of that entity; for example, “a polynucleotide,” is understood to represent one or more polynucleotides.
- the terms “a” (or “an”), “one or more,” and “at least one” can be used interchangeably herein.
- the present invention also provides immunogenic compositions and methods for delivery of transmitted full-length HCV polynucleotide or polypeptide sequences to a vertebrate with optimal expression and safety conferred.
- the transmitted full-length HCV genome comprises the polynucleotide of SEQ ID NO. 776.
- These immunogenic compositions may be prepared and administered in such a manner that the encoded gene products are optimally expressed in the vertebrate of interest. As a result, these compositions and methods are useful in stimulating an immune response against HCV infection.
- expression systems and delivery systems are also included in the invention.
- the identified polynucleotides or polypeptides encoded by the polynucleotides of the invention may be in any form, and polypeptides are generated using techniques well known in the art. Examples include isolated HCV proteins produced recombinantly or proteins delivered in the form of an inactivated HCV vaccine, such as conventional vaccines.
- an isolated HCV polynucleotide or polypeptide or fragment, variant or derivative thereof is administered in an immunologically effective amount.
- the effective amount of conventional vaccines is determinable by one of ordinary skill in the art based upon several factors, including the antigen being expressed, the age and weight of the subject, and the precise condition requiring treatment and its severity, and route of administration.
- the combination of conventional antigen vaccine compositions with optimized nucleic acid or polypeptide compositions provides for therapeutically beneficial effects at dose sparing concentrations.
- immunological responses sufficient for a therapeutically beneficial effect in patients predetermined for an approved commercial product, such as for the conventional product described above, can be attained by using less of the approved commercial product when supplemented or enhanced with the appropriate amount of nucleic acid or polypeptide.
- a desirable level of an immunological response afforded by a DNA based pharmaceutical alone may be attained with less DNA by including an aliquot of a conventional vaccine.
- using a combination of conventional and DNA based pharmaceuticals may allow both materials to be used in lesser amounts while still affording the desired level of immune response arising from administration of either component alone in higher amounts (e.g. one may use less of either immunological product when they are used in combination). This may be manifest not only by using lower amounts of materials being delivered at any time, but also to reducing the number of administrations points in a vaccination regime (e.g. 2 versus 3 or 4 injections), and/or to reducing the kinetics of the immunological response (e.g. desired response levels are attained in 3 weeks instead of 6 after immunization).
- Determining the precise amounts of DNA based pharmaceutical and conventional antigen is based on a number of factors as described above, and is readily determined by one of ordinary skill in the art.
- an adjuvant to increase the immune response to an antigen is typically manifested by a significant increase in immune -mediated protection.
- an increase in humoral immunity is typically manifested by a significant increase in the titer of antibodies raised to the antigen
- an increase in T-cell activity is typically manifested in increased cell proliferation, or cellular cytotoxicity, or cytokine secretion.
- Nucleic acid molecules and/or polynucleotides of the present invention may be solubilized in any of various buffers.
- Suitable buffers include, for example, phosphate buffered saline (PBS), normal saline, Tris buffer, and sodium phosphate (e.g., 150 mM sodium phosphate).
- PBS phosphate buffered saline
- Tris buffer Tris buffer
- sodium phosphate e.g. 150 mM sodium phosphate
- Insoluble polynucleotides may be solubilized in a weak acid or weak base, and then diluted to the desired volume with a buffer. The pH of the buffer may be adjusted as appropriate.
- a pharmaceutically acceptable additive can be used to provide an appropriate osmolarity.
- compositions of the present invention can be formulated according to known methods. Suitable preparation methods are described, for example, in Remington's Pharmaceutical Sciences, 16th Edition, A. Osol, ed., Mack Publishing Co., Easton, Pa. (1980), and Remington's Pharmaceutical Sciences, 19th Edition, A. R. Gennaro, ed., Mack Publishing Co., Easton, Pa. (1995).
- composition may be administered as an aqueous solution, it can also be formulated as an emulsion, gel, solution, suspension, lyophilized form, or any other form known in the art.
- composition may contain pharmaceutically acceptable additives including, for example, diluents, binders, stabilizers, and preservatives.
- Example 1 is illustrative of specific embodiments of the invention, and various uses thereof. They are set forth for explanatory purposes only, and are not to be taken as limiting the invention.
- Example 1 is illustrative of specific embodiments of the invention, and various uses thereof. They are set forth for explanatory purposes only, and are not to be taken as limiting the invention.
- Example 1 is illustrative of specific embodiments of the invention, and various uses thereof. They are set forth for explanatory purposes only, and are not to be taken as limiting the invention.
- Plasma samples were obtained from 17 subjects with acute or very recent HCV infection representing subtypes la, lb, 2 or 3. These consisted of twice-weekly serial collections from source plasma donors who became HCV infected during the course of their plasma donations. The donors were untreated and asymptomatic throughout the collection period. Plasma samples from 14 subjects with chronic HCV infection from the U.S. served as controls. All patients were treatment-na ' ive. All subjects gave informed consent, and plasma collections were performed with institutional review board and other regulatory approvals. Plasma samples were tested for HCV RNA and viral specific antigen and antibodies by a battery of commercial tests (Abbott and Roche) (Fig. 1 and Table 2).
- a total of 154 samples were tested including a median of 8 sequential specimens per acutely infected subject (range 5-11).
- the initial samples from acute subjects were HCV vRNA and antibody negative, followed by a sharp rise in vRNA levels.
- Five of 17 acutely infected subjects developed HCV antibodies by the last sampling time point.
- RNA isolation and cDNA synthesis was performed as follows. Samples which contained approximately 100,000 viral RNA copies were extracted using the Qiagen BioRobot EZ1 Workstation with EZ1 Virus Mini Kit v2.0 (Qiagen, Valencia, CA). RNA was eluted in 60 ⁇ and immediately subjected to cDNA synthesis. Reverse transcription of RNA to single stranded cDNA was performed using Superscript III reverse transcriptase (Invitrogen Life Technologies, Carlsbad, CA).
- a mixture of 3 ⁇ (0.5 mM) of 10 mM deoxynucleoside triphosphate, 1.5 ⁇ (0.25 ⁇ ) of 10 ⁇ anti-sense primer and 34.5 ⁇ of vRNA was incubated at 65°C for 5 min and then was placed on ice for at least 1 min. The contents of the tube were collected by brief centrifugation. Next, 12 ⁇ of 5X reaction buffer, 3 ⁇ of 0.1 M DTT (5 mM), 3 ⁇ of RNAseOUT Recombinant RNASE Inhibitor, (40 units/ ⁇ , Invitrogen) and 3 ⁇ of Superscript IIITM Reverse Transcriptase (2001 ⁇ / ⁇ 1) were added to the mixture.
- the final reaction was incubated at 50°C for 60 min followed by an increase in temperature to 55°C for an additional 60 minutes.
- the reaction was heat-inactivated at 70°C for 15 minutes and then treated with RNaseH at 37°C for 20 minutes.
- SEQ ID NOS The antisense primers were designed specifically for different genotype.
- Single genome amplification was performed from prepared viral cDNA.
- cDNA was serially diluted and distributed among wells of replicate 96-well plates so as to identify a dilution where PCR positive wells constituted less than 30% of the total number of reactions. At this dilution, most wells contained amplicons derived from a single cDNA molecule. This was confirmed in every positive well by direct sequencing of the amplicon and inspection of the sequence for mixed bases (double peaks), which would be evidence of priming from more than one original template or the introduction of PCR error in early cycles. Any sequence with evidence of mixed bases was excluded from further analysis.
- PCR amplification was carried out in the presence of 10 x Taq High Fidelity Platinum PCR buffer, 2 mM MgS0 4 , 0.2 mM of each deoxynucleoside triphosphate, 0.2 ⁇ of each primer, and 0.1 ⁇ (0.5 Unit) Platinum Taq High Fidelity polymerase in a 20 ⁇ reaction (Invitrogen, Carlsbad, CA).
- genotype 1 1 st round sense primer l .core.Fl 5'- ATGAGCACGAATCCTAAACCTCAAAGA-3' (SEQ ID NO:761)(nt 342-368 H77) and 1 st round antisense primer 1.NS4A.R1 5'-GCACTCTTCCATCTCATCGAACTC-3' (SEQ ID NO:763) (nt 5451-5474 H77), 2 nd round sense primer l .core.F2 5 ' -TC AAAG AAAAAC C AAA CGTAACACCAACCG-3' (SEQ ID NO:764) (nt 362-391 H77) and 2 nd round antisense primer 1.NS3A4A.R2 5'-AGGTGCTCGTGACGACCTCCAGG-3' (SEQ ID NO:766) (nt 5297-5319 H77); (2) 5' quarter genome of
- PCR was performed in MicroAmp 96-well reaction plates (Applied Biosystems, Foster City, CA) with the following PCR parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94 °C for 15 s, an annealing step of 58°C for 30 s, an extension step of 68 °C for 5 min, followed by a final extension of 68 °C for 10 min.
- the product of the 1 st round PCR was subsequently used as a template in the 2 nd round PCR under same conditions but with a total of 45 cycles.
- sequences of 2850 were unambiguous at every position. Sequences chromatograms from 150 amplicons contained one or more "double peaks" representing mixed bases. Because mixed bases generally represented only a minority of polymorphisms, it was inferred that such ambiguities had resulted from Taq polymerase errors in the initial PCR cycles and not from amplification from more than one initial target template; in such cases a correction of the assignment of the base was made. In rare instances where the only polymorphic site(s) present were represented by double peaks, possibly due to mixed initial templates, the sequences were discarded and not included in the analysis. This amounted to 30 sequences or less than 1% of total amplicons generated.
- Example 2 Example 2
- a total of 38 transmitted/founder lineages from 17 acutely infected subjects were used to calculate the mutation rate.
- Each sequence within the lineage was compared with the T/F virus sequence of that lineage.
- the insertion, deletion, transition and transversion frequencies were counted by self-developed computer program and by Highlighter tool.
- the rates were calculated by taking the ratio of each frequency number and the total number of nucleotides of all the sequences within that lineage.
- the SNAP program (www.HIV.lanl.gov) was applied to the codon-aligned sequences of each T/F lineage. Within each lineage, the accumulations of synonymous and nonsynonymous substitutions were counted by comparing to the transmitted/founder viral sequences. The Jukes-Cantor corrected accumulation rates of synonymous substitutions per potential synonymous site (ds) and nonsynonymous substitution per potential nonsynonymous site (dn) were compared to screen for positive selection.
- Nucleotide polymorphisms in subject 10051 were essentially random, corresponding to a star-like phylogeny and a Poisson distribution of low frequency events.
- Two sequences (2C3 and 2C2) contained a single common polymorphism at position -2200, and two others (2A2 and 2B34) contained a different common polymorphism at position -4375, indicating for each pair shared recent ancestry.
- Figures 4C and 4D depict sequences from subjects having evidence of productive infection by more than one genetically distinct virus. Each of the subjects was sampled on 4 occasions and sequences again showed increasing diversity over time. The consensus sequences of each low diversity lineage in these subjects were interpreted to correspond to a unique T/F HCV genome. Thus, each of these subjects was productively infected by at least three viruses. The proportion of sequences represented in each lineage was approximately the same in subject 10012 (Fig 4C) but not in subject 10062 (Fig 4D). This could be due either to different replication rates of the different T/F viral viruses, differential effects of early innate or adaptive immune responses, different times of infection by different transmitted viruses, or early stochastic events in the infection process.
- a mathematical model was developed to estimate the adequacy of sampling given the observed distribution of T/F lineage sequences and total number of sequences analyzed.
- This model when applied to the 179 sequences analyzed for subject 10012, estimated that the actual number of T/F viruses was 3 at a 95 % confidence level.
- the model when applied to the 164 sequences analyzed for subject 10062, the model estimated that the actual number of T/F viruses could be as high as 5, again at a 95 % confidence level.
- Sequences from subject 10029 where low diversity sequence lineages corresponding to 9 T/F viruses could be unambiguously identified, but with even greater differences in sequence proportions.
- the model estimated that the actual number of T/F viruses could be as high as 13 at a 95 % confidence level.
- T/F viruses If an acutely-infected subject acquired multiple HCV genomes from a contact who himself was acutely infected by a single virus, then these T/F viruses would be expected to all be closely related to each other, since they must reflect the viral diversity in the donor. If that donor contact had been acutely infected by more than one genetically divergent virus and the acutely infected recipient acquired multiple progeny of these, then the expectation would be for T/F viral sequences to be represented by subsets of sequences, some closely related and some not. Evidence in five subjects of acute-to-acute transmission was found by these two scenarios.
- Figures 13A-C depict ML trees and Highlighter plots from acutely infected subjects 10020, 10016 and 10003, where the maximum diversity of T/F viruses from a consensus was 0.18 %, 0.16 % and 0.12 %, respectively.
- the phylogenetic patterns of these sequences were quite different from the single variant transmissions shown in Figs 3B and 4 A and 4B, which exhibited star-like phylogeny and conformed to a Poisson distribution of random mutations.
- the sequences in 13A-C instead, violated the Poisson model. They were comprised of distinct subsets of highly related sequences that differed from each other and from a common consensus by 1-6 nucleotides.
- the phylogenetic pattern of sequences from subject 106889 was far more complicated than that of any of the other 16 acutely infected subjects.
- Subject 106889 exhibited a typical acute infection viral kinetic profile with four sequential plasma samples negative for HCV vRNA followed by rapid vR A ramp-up to nearly 10 7 vRNA IU/ml (Fig 14 insert). Eighty-six 5 ' half genome sequences were obtained from the initial plasma vRNA positive time point.
- T/F viral sequences in 17 subjects provided a unique strategy for analyzing HCV sequence evolution in vivo in the critical period beginning at or near the moment of virus transmission and extending to the establishment of viral load setpoint and in some subjects HCV antibody seroconversion.
- a summary of this analysis is presented in Table 1.
- Maximum intra-lineage diversity for the 17 subjects ranged from 0.06 % to 0.22 %. Insertions (0.000001), deletions (0.00004) and stop codons (0.000004) were infrequent. Transitions outnumbered transversions by 8 to 1.
- the overall mutation frequency including all sampled time points was low (0.000145) given the number of possible virus replication cycles between the first infected hepatocyte and setpoint viremia 6-10 weeks later when as many as 7-20 % of hepatocytes may be productively infected.
- the dN/dS ratio was low at 0.33 as was its derivative pN/pS at 0.33 %.
- the diversity of sequences within the T/F lineages of each acutely infected subject corresponded to a star-like pattern of random diversification that conformed to a Poisson distribution mutations.
- Exceptions were of four types: (i) infrequent shared polymorphisms resulting from stochastic changes in the newly infected subjects (e.g., Fig 3B; 4A, 4B, 4C, 4D); (ii) transmission of multiple closely related variants (e.g., Figs 11, 13 A, 13B; 14); (iii) evidence of immune selection in later samples (Fig 12); and (iv) rare examples of short inverted repeats resulting from template switching between double-stranded RNA duplexed hairpin structures by the RNA dependent polymerase (Fig 13).
- Non- random amino acid polymorphisms in multiple T/F sequences in subjects 10020 (Fig 13 A) and 10016 (Fig 13B) also corresponded to known or predicted CTL epitopes and thus are likely to represent CTL escape or reversion in the respective donors (not in subjects 10020 or 10016).
- HCV-1 transmission for acutely infected subjects is believed to result from high virus loads, each of circulating neutralizing antibodies and restricted viral genetic diversity, an interpretation corroborated by acute infection studies of SIV in Indian rhesus macaques. All of these same factors could pertain to acute -to-acute HCV transmission.
- the complicated HCV transmission scenario for subject 106889 illustrates the sensitivity and specificity with which the SGA-direct amplicon sequencing strategy can detect and discriminate T/F viral genomes and their progeny.
- SGA-direct sequencing includes in vitro recombination artifacts, avoids Taq polymerase-mediated nucleotide substitution errors in finished sequences, precludes founder effects due to disproportionate target amplification, and avoids cloning bias. This allows for a proportional representation of target molecules in finished sequences. It also allows for genetic linkages to be preserved in finished sequences specific mutations, in this case NS3 DAA resistance mutations V36M and R155K to be precisely identified.
- each T/F virus contained the V36M/R155K double mutation indicating high level NS3 protease resistance.
- Subject 106889 as well as the transmitting partner, SGA-direct sequencing is ideally suited to identifying genetically-linked DAA resistant mutants within single or multiple genes (e.g., NS3, NS5A and NS5B) of transmitted viruses and characterizing their persistence or disappearance over time.
- the transmitted/founder virus sequence of the full length genome of subject 10021 was obtained by analysis of five overlapped genome fragments: 5' UTR (nt 1 to 852), 5' half genome (core, El, E2, p7, NS2 andNS3, nt 391 to 5297), 3' half genome (NS4A, NS4B, NS5A andNS5B, nt 5168 to 9374), 3' U/C tract (nt 9082 to 9582) and 3' x-tail (nt 9563 to 9646).
- the primers designed for cDNA synthesis and PCR amplification are listed below.
- the nucleotide numbering system is based on reference sequence H77 (accession number NC 004102). The sequences of all five genome regions were analyzed and assembled.
- the transmitted/founder sequence was inferred based on phylogenetic inference and mathematical modeling.
- the cDNA synthesis and single genome amplification of '5 and '3 half genome was performed as described supra in Example 2.
- the positive PCR reactions were subject to direct sequencing.
- the cDNA synthesis of 3' U/C tract of subject 10021 was also performed using Superscript IIITM Reverse Transcriptase.
- the mixture of 1 ⁇ (0.5 mM) of 10 mM deoxynucleoside triphosphate, 0.5 ⁇ (0.25 ⁇ ) of 10 ⁇ anti-sense primer 3UTR-R10 and 11.50 ⁇ of vRNA was heated at 65 °C for 5 min and then was placed on ice for at least 1 min.
- PCR amplification was carried out in the presence of 2 ⁇ of 10 x Taq High Fidelity Platinum PCR buffer, 0.8 ⁇ of 50 mM MgS0 4 (2 mM), 0.4 ⁇ of 10 mM deoxynucleoside triphosphate (0.2 mM), 0.2 ⁇ of each primer (0.2 ⁇ ), and 0.1 ⁇ (0.5 Unit) Platinum Taq High Fidelity polymerase in a 20 ⁇ reaction.
- the semi-nested PCR primers used are listed in Table 4.
- PCR was performed in MicroAmp 96-well reaction plates with the following PCR parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94°C for 15 s, an annealing step of 50 °C for 30 s, an extension step of 68°C for 1.5 min, followed by a final extension of 68 °C for 10 min.
- the product of the 1 st round PCR was subsequently used as a template in the 2 nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on precasted 1 % agarose E-gels 96 and 4 % Nusieve GTG Agarose gel (CAMBREX, Cat. No. 50080).
- the positive PCR reactions were TA cloned into pGEM-T Easy Vector System (Promega, Cat. No. A3610). The inserts were identified and sequenced using Ml 3 forward/reverse primers.
- the transmitted/founder virus sequence of 5' UTR of subject 10021 was obtained using FirstChoice RLM-RACE kit (Ambion, Cat. No. AM 1700), with slight modification of the methods recommended by the manufacturer.
- TAP Tobacco Acid Pyrophosphatase
- vRNA acquires the adapter sequence as its 5' end
- a core gene specific RT-PCR amplified the complete 5' UTR genome.
- the cDNA synthesis and PCR amplification were performed under same conditions as that described for 3' U/C tract except using 55 °C as the annealing temperature.
- the positive PCR products were sequenced directly.
- RNA oligonucleotide adaptor (Integrated DMA Technologies INC.) w r as iigated to the 3' end of the vRNA by using T4 RNA ligase.
- the adaptor was 5' phosphorylated and 3' Dideoxy-C blocked to prevent intra-molecular ligation.
- the 20 ul ligation reaction that contained 15 ⁇ of vRNA, 2 ⁇ of 10X ligation buffer, 1 ⁇ of 10 rnM ATP, 0.5 ⁇ of RNase Inhibitor (20 ⁇ / ⁇ 1), 1 ⁇ of 10 ⁇ oligo adaptor ( 0.5 ⁇ ) and 1 ⁇ of T4 RNA ligase (10 units) was incubated at 37 °C for 60 min.
- the whole ligation reaction was subject to cDNA synthesis directly using the oligo adaptor specific primer.
- the cDNA synthesis was carried at 50 °C for 60 min.
- PCR was performed with the following parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94 °C for 15 s, an annealing step of 55 °C for 30 s, an extension step of 68 °C for 30 sec, followed by a final extension of 68 °C for 10 min.
- the product of the 1 st round PCR was subsequently used as a template in the 2 nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on 4 % Nusieve GTG Agarose gel (CAMBREX). The positive PCR reactions were directly TA cloned into pGEM-T Easy Vector System (Promega, Cat. No. A3610). The inserts were identified and sequenced using Ml 3 forward/reverse primers.
- the inferred full-length transmitted/founder sequences was chemically synthesized in four subgenomic fragments (Blue Heron Biotechnology). The fragments were overlapped at the unique restriction sites, BsrGI (nt 3640), SnaBI (nt 6644) and Sfil (nt 9404), respectively.
- BsrGI nt 3640
- SnaBI nt 6644
- Sfil nt 9404
- the noncutter Notl was attached to the 5' terminal of the first fragment and noncutter Xbal was attached to the 3' terminal of all the fragments during chemical synthesis.
- the synthesized fragments were propagated, restriction enzyme digested and sequentially cloned into the MCS Notl-Xbal site of pBlueScriptKS(+). The final full-length clone was confirmed by sequence analysis. Table 1. Diversity and mutation analyses of HCV sequences in acute infection.
- a 5h-5' half genome contains core, El, E2, p7, NS2, and NS3;
- 5ql-5' quarter 1 genome contains core, El, E2, p7 and partial NS2;
- 5q2-5' quarter 2 genome contains partial NS2 and NS3.
- h Averages were calculated from total mutations in all transmitted/founder lineages from all subjects combined. Because of low numbers of sequences and mutations in some lineages, certain values (e.g. dN/dS for subjects 10062 v2 and 10004 v3) vary substantially from the mean.
- a obtained by applying jackknife to the biasd estimator in this case is the clusters found.
- the average cut-off method is the more conservative estimate of the minimum number of founder strains needed to explain the observed diversity.
- the maximum cut-off distinguishes more lineages and separates them into clusters from distinct founders.
- NS3.F2 (SEQ ID NO: 785) 5146-5168 +, 2nd round
- 3UTR-R10 (SEQ ID NO: 791) 9582-9604 -, RT and 1st round
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Organic Chemistry (AREA)
- Wood Science & Technology (AREA)
- Engineering & Computer Science (AREA)
- Zoology (AREA)
- Genetics & Genomics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Immunology (AREA)
- General Engineering & Computer Science (AREA)
- Communicable Diseases (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Biotechnology (AREA)
- Biochemistry (AREA)
- Virology (AREA)
- General Health & Medical Sciences (AREA)
- Microbiology (AREA)
- Physics & Mathematics (AREA)
- Biomedical Technology (AREA)
- Analytical Chemistry (AREA)
- Biophysics (AREA)
- Medicinal Chemistry (AREA)
- Molecular Biology (AREA)
- Medicines Containing Antibodies Or Antigens For Use As Internal Diagnostic Agents (AREA)
- Peptides Or Proteins (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
The invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection, and methods for identifying nucleotide sequences corresponding to specific transmitted HCV genomes.
Description
FULL-LENGTH TRANSMITTED HEPATITIS C VIRUS (HCV) GENOMES
IDENTIFIED BY SINGLE GENOME AMPLIFICATION
CROSS-REFERENCES TO RELATED APPLICATIONS
This application claims the benefit of priority to United States Provisional Patent Application Serial No. 61/630,161, filed on December 5, 2011, which is incorporated by reference in its entirety.
GOVERNMENT RIGHTS
This invention was supported in part with funding provided by NIH Grant Nos. NIH Chavi Award (NIH U01 AI 067854) and the University of Pennsylvania Center for AIDS Research (NIH P30 AI 045008), each awarded by the National Institutes of Health. The government may have certain rights to this invention.
FIELD OF THE INVENTION
The invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection. The invention provides methods for identifying nucleotide sequences corresponding to specific transmitted HCV genomes, including structural gene nucleotide sequences of global HCV genotypes, subtypes, and drug resistant variants that actively mediate viral infection. The invention further provides methods of administering a vaccine comprising transmitted full-length hepatitis C virus (HCV) genomes or portions thereof.
BACKGROUND OF THE INVENTION
Hepatitis C Virus (HCV) is a positive strand, non-segmented, enveloped RNA virus of approximately 9.6 kb in length. The virus is classified in the genus Hepacivirus within the larger family of Flavivirus, which includes the human pathogens West Nile virus, yellow fever virus and dengue fever virus among others. A common feature among the Flaviviridae is their dependence on a virally-encoded RNA-dependent RNA polymerase (RdRp) for replication. RdRp is error-prone, and HCV is notable for its quasispecies complexity and broad genotypic diversity. Globally, there are seven major genotypes of HCV that differ by approximately 30- 35% in nucleotide sequence.
The extraordinary diversity of HCV has important implications for clinical management and for basic and translational research aimed at elucidating viral natural history, pathogenesis and susceptibility to novel therapeutics and vaccines. Clinically, the different HCV genotypes exhibit variable natural history, responsiveness to interferon and ribavirin (mainstays of current therapy), and sensitivity to the many recently approved or still investigational direct acting
antiviral (DAA) drugs. The development of new drugs and drug combinations that effectively suppress HCV replication and prevent the emergence of DAA resistance is a major challenge given the high rates of virus replication and variation. HCV diversity poses similar challenges to the development of effective vaccines and to the elucidation of virus biology, immunopathogenesis, and gene structure-function relationships. It is of interest and practical relevance that the extraordinary diversity of HCV is mirrored by comparable diversity of HIV- 1 , and that a novel experimental strategy to identify transmitted/founder (T/F) HIV-1 genomes in acute infection has led to new insights into HIV-1 transmission, persistence, and evolution in the face of cellular and humoral immune responses and antiretroviral therapy.
Acute HCV infection, defined as the period between virus transmission and antibody seroconversion which occurs about 6 - 10 weeks later, sets in motion viral-host interactions that largely dictate the natural history of the infection. Depending on viral genotype and host immunogenetic factors, most importantly IL28B alleles, a proportion of newly infected individuals spontaneously control or eliminate the virus. A greater number of patients can be cured if the infection is treated with interferon and ribavirin alone or in combination with DAA drugs. Mechanistically, how this occurs is unknown, and how the emergence of DAA drug resistance in communities of chronically infected subjects or in acutely infected subjects who initiate early therapy will affect treatment responses is uncertain, but again there are parallels with HIV-1. From a vaccine perspective, the acute infection period is critical. Transmitted viruses are the targets of a vaccine and early stages of infection represent a period when the virus should be most vulnerable to elimination by vaccine-elicited immune responses. For these reasons, there is considerable interest in the molecular features of the initial population 'bottleneck' to HCV transmission and subsequent pathways of virus evolution leading to viral persistence.
Previous reports have described different experimental approaches to the analysis of the HCV transmission bottleneck. These can generally be divided into studies that employed a DNA heteroduplex gel shift analytical method; studies that used conventional polymerase chain reaction (PCR) methods to amplify, clone and sequence fragments of the HCV genome; and studies that used 454 pyrosequencing to analyze acute and early viral sequences more deeply. All of these studies documented a restriction in viral diversity associated with virus transmission. However, despite the use of increasingly sensitive methods, a precise quantitative and molecular description of HCV transmission and early diversification remain elusive. This is because previous studies employed methods that were unable to sufficiently resolve HCV diversity between transmitted genomes that can be as low as 0.01%; employed Tag polymerase
amplification from heterogeneous target templates, thereby allowing for artifactual in vitro recombination that confounds phylogenetic interpretation; employed 454 pyrosequencing, which yields short sequences lacking genetic linkage across genes and genomes; or evaluated few subjects.
Thus, there is a need in the art to allow for precise and unambiguous molecular identification of the complete full-length nucleotide sequences of HCV genome, particularly acute infection stage HCV, in order to gain a more comprehensive molecular understanding of HCV transmission and for development of effective HCV vaccines and treatments.
SUMMARY OF THE INVENTION
The invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection. One application of this invention is to use these transmitted full-length genomes for development of more effective vaccines and treatments. In one embodiment the transmitted full-length HCV genome comprises the polynucleotide of SEQ ID NO: 776. In certain embodiments, the HCV genome mediates viral transmission. In other embodiments, the polypeptide sequence comprises env and core genes. Nucleotide sequence of transmitted HCV genomes is set forth in Table 5.
In other aspects, the invention provides methods for identifying full-length transmitted HCV genomes, the methods comprising collecting a patient sample, isolating and preparing viral RNA for sequencing, sequencing viral RNA that includes HCV genomes of circulating virus, performing sequence alignment of selected HCV genome regions, analyzing phylogenetically selected sequence alignments; and identifying full-length HCV genomes of transmitted virus. In particular embodiments the HCV genome polypeptide sequence comprises SEQ ID NO. 776.
In other aspects, the invention provides an immunogenic composition comprising transmitted full-length HCV genome or portions thereof. In particular embodiments the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
In other aspects, the invention provides methods for administering vaccines comprising transmitted full-length HCV genomes or portions thereof, wherein an immune response is induced in the patient following vaccination. In particular embodiments the HCV genome comprises the polynucleotide of SEQ ID NO: 776. The advantages of administering the disclosed HCV sequences and/or polypeptides encoded by such HCV sequences is such vaccines would illicit an immune response that is specific to transmitted HCV genomes rather than circulating HCV genomes, thereby improving vaccine effectiveness.
BRIEF DESCRIPTION OF THE DRAWINGS
These and other objects and features of this invention will be better understood from the following detailed description taken in conjunction with the drawings, wherein:
Figure 1 illustrates HCV R A kinetics in 16 acute infection subjects. The shaded area represents that the HCV RNA is below the linear range of quantitation (43-69,000,000 IU/ml). Plus signs indicate the samples are Anti-HCV antibody positive. The circles represent the time points that were sequence analyzed from each subject.
Figure 2 is a Maximum-likelihood tree (ML) of 5' quarter genome {core, El and E2) sequences from 17 acutely and 14 chronically infected subjects. The sequences from acutely infected subjects are shown in red and the sequences from chronically infected subjects are shown in blue. The reference sequences of genotype 1 to 7 are in gray. Bootstrap values (>70%) are indicated for intra-subject clusters. The horizontal scale bar represents 3% genetic distance.
Figure 3 illustrates 5' half-genome sequence diversity in two subjects, one with chronic infection (WIMI4025 (Fig 3 A)) and one with acute infection (10051 Fig 3(B)). Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. Bootstrap values (>70%) are indicated for intra-linerage clusters. The horizontal scale bars represent 0.5% (24.5 nt) and 0.02% (1 nt) genetic distance for subject WIMI and 10051, respectively.
Figure 4 illustrates 5' quarter-genome {Core, El and E2) sequence diversity in four acutely infected subjects. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. Sequences from multiple time points are color coded with orange, green, blue and black in chronological order. Sequences from subject 10021 (Fig 4A) and 10025 (Fig 4B) each showed productive infection by a single virus. Sequences from subject 10012 (Fig 4C) and 10062 (Fig 4D) showed productive infection by at least three viruses, repectively. The horizotal scale bars represent 0.04% (1 nt) genetic distance for Figs. 4A, 4B and 4D and 0.4% (10 nt) for 4C. Bootstrap values (>70%) are indicated for within patient intra-lineage clusters in the ML tree.
Figure 5 illustrates 5' quarter-genome {Core, El and E2) sequence divsersity in an acutely infected subject. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' quarter genome {Core, El and E2) sequences generated from multiple time points showed productive infection by at least 9 transmitted/founder variants. The close genetic diversity between variant 1 and 2, variant 6 and 7 is indicative of transmission from an acutely infected
donor. The horizotal scale bars represent 0.2% (5 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra- lineage clusters.
Figure 6 illustrates HCV diversity in acutely infected subject 10024. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' quarter genome sequences derived from multiple time points showed productive infection by at least 7 transmitted/founder variants. The horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 7 illustrates HCV diversity in acutely infected subject 6123. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 3 transmitted/founder variants. The horizotal scale bars represent 0.2% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 8 illustrates HCV diversity in acutely infected subject 6222. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 4 transmitted/founder variants. The horizotal scale bars represent 0.02% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 9 illustrates HCV diversity in acutely infected subject 10004. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of the 5' quarter genome sequences showed productive infection by at least 3 transmitted/founder variants. The horizotal scale bars represent 0.36% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 10 illustrates HCV diversity in acutely infected subject 10002. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 13 transmitted/founder variants. The horizotal scale bars represent 0.2% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra- lineage clusters.
Figure 11 illustrates HCV diversity in acutely infected subject 10017. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure,
respectively. The ML tree and highlighter plot of 5' quarter genome sequences derived from multiple time points showed productive infection by at least 6 transmitted/founder variants. The horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 12 illustrates HCV diversity in acutely infected subject 9055. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' half genome sequences derived from multiple time points showed productive infection by a single transmitted/founder variants. The horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 13 illustrates 5' quarter-genome (Core, El and E2) sequence divsersity in acute- to-acute transmission. 5' quarter genome (Core, El and E2) sequences from subject 10020 (Fig 13 A) and 10016 (Fig 13B), and 5' half genome (core, El, E2, p7, NS2 and NS3A) sequences from subject 10003 (Fig 13C) are depicted by ML trees and by highlighter plots. Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The first sequence in red is the consensus sequence derived from the entire sequence dataset. The horizotal scale bars represent 0.04%> (1 nt), 0.036%) (1 nt) and 0.02%> (1 nt) genetic distance for Figs. 13 A, 13B and 13C, repectively. Bootstrap values (>70%) are indicated.
Figure 14 illustrates HCV RNA kinectics and diversity in subject 106889. The inset shows the HCV RNA kinetics. The time point that was sequence analyzed is circled. Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The 5' half genome (core, El, E2, p7, NS2 and NS3A) sequences formed 10 distinct lineages (labeled with letter L) with high statistical support in the ML tree. Each cluster is color coded. The T/F variants was identified within each lineage. A total of 33 T/F variants were found that were responsible for productive infection. The horizotal scale bars represent 0.04%) (1 nt) genetic distance. Bootstrap values (>70%>) are indicated.
Figure 15 illustrates strand transfers at stem loop structures of HCV RNA. Figures 17A- E illustrate four examples of stem loop structures in double-stranded HCV RNA leading to template switching by the HCV RNA polymerase. Fig 17A, SEQ ID NO: 800 First 5' to 3' horizontal sequence; SEQ ID NO: 801 Second 3' to 5' horizontal sequence; SEQ ID NO: 802 Third 5' to 3' horizontal sequence; SEQ ID NO: 800 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 802 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 801 Sixth 3' to 5' haripin sequence. Fig 17B, SEQ ID NO: 803 First 5' to 3' horizontal sequence; SEQ ID NO: 804 Second 3' to 5'
horizontal sequence; SEQ ID NO: 805 Third 5' to 3' horizontal sequence; SEQ ID NO: 803 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 805 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 804 Sixth 3' to 5' hairpin sequence. Fig 17C, SEQ ID NO: 806 First 5' to 3' horizontal sequence; SEQ ID NO: 807 Second 3' to 5' horizontal sequence; SEQ ID NO: 808 Third 5' to 3' horizontal sequence; SEQ ID NO: 806 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 808 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 807 Sixth 3' to 5' hairpin sequence. Fig 17D, SEQ ID NO: 809 First 5' to 3' horizontal sequence; SEQ ID NO: 810 Second 3' to 5' horizontal sequence; SEQ ID NO: 811 Third 5' to 3' horizontal sequence; SEQ ID NO: 809 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 811 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 810 Sixth 3' to 5' hairpin sequence. Fig 17E, SEQ ID NO: 812 First 5' to 3' horizontal sequence; SEQ ID NO: 813 Second 3' to 5' horizontal sequence; SEQ ID NO: 814 Third 5' to 3' horizontal sequence; SEQ ID NO: 812 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 814 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 813 Sixth 3' to 5' hairpin sequence.
Figure 16 illsutrates the strategy for identification of full-length HCV genome by single gneome amplification (SGA). The complete HCV genome of subject 10021 was determined by amplifying five overlapping genome fragments: fragment I (nt 1 - 852) contains the complete 5' UTR and partial Core; fragment II (nt 391 - 5297) contains Core, El, E2, NS2 and NS3; fragment III (nt 5168 - 9374) contains NS4A, NS4B, NS5A and NS5B; fragment IV (nt 9082 - 9582) contains partial NS5B, variable region and the complete poly U/UC tract; and fragment V (nt 9563 - 9646) contains the complete x-tail. The nucleotide positions are based reference sequence H77.
Figure 17 illustrates the strategy for generating a full-length T/F HCV molcular clone from subject 10021.
Figure 18 illustrates HCV diversity in subject 10003. ML tree and Highlighter plot of 5' halfgenome sequences reveal many sets of closely related sequences distinguished by unique shared mutations.
Figure 19 illustrates HCV diversity analysis in subject 10016 suggests acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 5' quarter 1 genome sequences. Visualization of 15 potential T/F viral lineages distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveal 10 and 4 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
Figure 20 illustrates maximum diversity of discrete HCV sequence lineages from acute infection subjects versus maximum sequence diversity in chronic subjects. Primary data are
derived from Tables 1 and 2. Mean (±95% CI) values are represented by horizontal lines. Differences between the two groups were highly significant (p<0.0001; unpaired T-test with Welch's correction), reflecting the recent and remote diversification histories of acute and chronic sequences, respectively.
Figure 21 shows HCV diversity in acute subject 10051. 5' quarter 1 genomesequences are color coded in orange, green, blue and black in chronological order to reflect sampling time points in Figure 1 and are represented in a ML tree and Highlighter plot. Sequences show evidence of productive clinical infection by a single virus. The horizontal scale bar indicates genetic distance.
Figure 22 illustrates nonsynonymous and synonymous mutations in HCV sequences from acute subjects 10003, 10020 and 10016. Highlighter plotsof 5' half or quarter 1 genomesequences are color coded to denote nonynonymous (red) and synonymous (green) mutations for subjects 10003 (panel A), 10020 (panel B) and 10016 (panel C).
Figure 23 illustrates amino acid alignment of the HCV Env coding region of acute subject 10003. The H77 reference sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
Figure 24 illustrates amino acid alignment of the HCV Env coding region of acute subject 10020. The H77 reference sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
Figure 25 illustrates amino acid alignment of the HCV Env coding region of acute subject 10016. The H77 sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
Figure 26 shows HCV diversity analysis in subject 10020 to suggest acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 5' quarter 1 genome sequences. Visualization of 10 potential T/F viral sequences distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveals 5 and 3 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
Figure 27 shows HCV diversity analysis in subject 10003 to suggest acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 59 half genome sequences. Visualization of 37 potential T/F viral sequences distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveals 15 and 8 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
Figure 28 illustrates nonsynonymous and synonymous mutations in HCV sequences from acute subject 9055. A Highlighter plot (panel A) of 5' half genomesequences is color coded to denote nonynonymous (red) and synonymous (green) mutations. The boxed area reveals a temporal expansion of sequences with concentrated amino acid subtitutions in NS3. In panel B, amino acid selection is evident in a previously identified CTL epitope highlighted in red. The top-most sequence represents the genotype 3a consensus.
DETAILED DESCRIPTION
The invention provides full-length transmitted HCV genomes that mediate viral infection and transmission. Prior to the present invention the identification of these genomes was unavailable due to the following: (i) a significant time period from weeks to months between the moment of transmission and the first appearance of HCV in the blood (Bowen. D.G. and Walker, CM., 2005, Nature 436: 946-952; Moradpour, D. et al, 2007, Nature Rev Micro 5:453- 463); (ii) HCV is genetically highly variable in its nucleotide sequence due to its error-prone RNA-dependent RNA polymerase, and as a result, it exists in individuals as a complex mixture of sequences commonly referred to as a 'quasispecies' (Moradpour, D. et al., 2007, Nature Rev Micro 5:453-463); (iii) conventional experimental approaches to sequencing the HCV genome from clinical samples introduced addition variation into the sequences as a consequence of Taq polymerase induced recombination and nucleotide misincorporation errors (Salazar-Gonzalez, J.F. et al., 2008, J Virol 82:3952-3970); and (iv) identifying from these myriad of sequences which sequences corresponded to actual transmitted viruses, or even viruses that were replication-competent and responsible for ongoing virus replication and persistence in the infected subject was not achievable. The identification of the full-length transmitted viral sequence provides specific nucleotide and protein regions responsible for mediating viral infection. In certain embodiments the polypeptide sequence comprises env and core genes. Utilizing these specific regions in the preparation of vaccines and therapeutic provides for more effective HCV treatments.
The term "polynucleotide" is intended to encompass a singular nucleic acid or nucleic acid fragment as well as plural nucleic acids or nucleic acid fragments, and refers to an isolated molecule or construct, e.g., a virus genome (e.g., vRNA), messenger RNA (mRNA), plasmid DNA (pDNA), or derivatives of pDNA (e.g., minicircles as described in (Darquet, A-M et al., 1997, Gene Therapy, 4: 1341-1349) comprising a polynucleotide. A polynucleotide may comprise a conventional phosphodiester bond or a non-conventional bond (e.g., an amide bond, such as found in peptide nucleic acids (PNA)).
The terms "nucleic acid" or "nucleic acid fragment" refer to any one or more nucleic acid segments, e.g., DNA or R A fragments, present in a polynucleotide or construct. A nucleic acid or fragment thereof may be provided in linear (e.g., mR A) or circular (e.g., plasmid) form as well as double-stranded or single-stranded forms. By "isolated" nucleic acid or polynucleotide is intended a nucleic acid molecule, DNA or RNA, which has been removed from its native environment. For example, a recombinant polynucleotide contained in a vector is considered isolated or "cloned" for the purposes of the present invention. Further examples of an isolated polynucleotide include recombinant polynucleotides maintained in heterologous host cells or purified (partially or substantially) polynucleotides in solution. Isolated RNA molecules include in vivo or in vitro RNA transcripts of the polynucleotides of the present invention. Isolated polynucleotides or nucleic acids according to the present invention further include such molecules produced synthetically.
The terms "fragment," "variant," and "derivative" when referring to HCV polypeptides of the present invention include any polypeptides which retain at least some of the immunogenicity or antigenicity of the corresponding native polypeptide. Fragments of HCV polypeptides of the present invention include proteolytic fragments, deletion fragments and in particular, fragments of HCV polypeptides which exhibit increased secretion from the cell or higher immunogenicity or reduced pathogenicity when delivered to an animal. Polypeptide fragments further include any portion of the polypeptide which comprises an antigenic or immunogenic epitope of the native polypeptide, including linear as well as three-dimensional epitopes. Variants of HCV polypeptides of the present invention include fragments, and also polypeptides with altered amino acid sequences due to amino acid substitutions, deletions, or insertions. Variants may occur naturally, such as an allelic variant. By an "allelic variant" is intended alternate forms of a gene occupying a given locus on a chromosome or genome of an organism or virus. Genes II, Lewin, B., ed., John Wiley & Sons, New York (1985), which is incorporated herein by reference. For example, as used herein, variations in a given gene product is a "variant". Naturally or non-naturally occurring variations such as amino acid deletions, insertions or substitutions may occur. Non-naturally occurring variants may be produced using art-known mutagenesis techniques. Variant polypeptides may comprise conservative or non-conservative amino acid substitutions, deletions or additions. Derivatives of HCV polypeptides of the present invention, are polypeptides which have been altered so as to exhibit additional features not found on the native polypeptide. Examples include fusion proteins. An analog is another form of an HCV polypeptide of the present invention. An example is a proprotein which can be activated by cleavage of the proprotein to produce an active mature polypeptide.
In certain embodiments, the polynucleotide, nucleic acid, or nucleic acid fragment is DNA. In the case of DNA, a polynucleotide comprising a nucleic acid which encodes a polypeptide normally also comprises a promoter and/or other transcription or translation control elements operably associated with the polypeptide-encoding nucleic acid fragment. An operable association is when a nucleic acid fragment encoding a gene product, e.g., a polypeptide, is associated with one or more regulatory sequences in such a way as to place expression of the gene product under the influence or control of the regulatory sequence(s). Two DNA fragments (such as a polypeptide-encoding nucleic acid fragment and a promoter associated with the 5' end of the nucleic acid fragment) are "operably associated" if induction of promoter function results in the transcription of mRNA encoding the desired gene product and if the nature of the linkage between the two DNA fragments does not (1) result in the introduction of a frame-shift mutation, (2) interfere with the ability of the expression regulatory sequences to direct the expression of the gene product, or (3) interfere with the ability of the DNA template to be transcribed. Thus, a promoter region would be operably associated with a nucleic acid fragment encoding a polypeptide if the promoter was capable of effecting transcription of that nucleic acid fragment. The promoter may be a cell-specific promoter that directs substantial transcription of the DNA only in predetermined cells. Other transcription control elements, besides a promoter, for example enhancers, operators, repressors, and transcription termination signals, can be operably associated with the polynucleotide to direct cell-specific transcription. Suitable promoters and other transcription control regions are disclosed herein.
A variety of transcription control regions are known to those skilled in the art. These include, without limitation, transcription control regions which function in vertebrate cells, such as, but not limited to, promoter and enhancer segments from cytomegaloviruses (the immediate early promoter, in conjunction with intron-A), simian virus 40 (the early promoter), and retroviruses (such as Rous sarcoma virus). Other transcription control regions include those derived from vertebrate genes such as actin, heat shock protein, bovine growth hormone and rabbit β-globin, as well as other sequences capable of controlling gene expression in eukaryotic cells. Additional suitable transcription control regions include tissue-specific promoters and enhancers as well as lymphokine-inducible promoters (e.g., promoters inducible by interferons or interleukins).
Similarly, a variety of translation control elements are known to those of ordinary skill in the art. These include, but are not limited to ribosome binding sites, translation initiation and termination codons, elements from picornaviruses (particularly an internal ribosome entry site, or IRES, also referred to as a CITE sequence).
A DNA polynucleotide of the present invention may be a circular or linearized plasmid or vector, or other linear DNA which may also be non-infectious and nonintegrating (i.e., does not integrate into the genome of vertebrate cells). A linearized plasmid is a plasmid that was previously circular but has been linearized, for example, by digestion with a restriction endonuclease. Linear DNA may be advantageous in certain situations as discussed, e.g., in Cherng, J. Y., et al., 1999, J Control. Release 60:343-53, and Chen, Z. Y., et al, 2001, Mol Ther 3:403-10. As used herein, the terms plasmid and vector can be used interchangeably. In other embodiments, a polynucleotide of the present invention is RNA, for example, in the form of messenger RNA (mRNA). Methods for introducing RNA sequences into vertebrate cells are described in U.S. Pat. No. 5,580,859.
Polynucleotides, nucleic acids, and nucleic acid fragments of the present invention may be associated with additional nucleic acids which encode secretory or signal peptides, which direct the secretion of a polypeptide encoded by a nucleic acid fragment or polynucleotide of the present invention. According to the signal hypothesis, proteins secreted by mammalian cells have a signal peptide or secretory leader sequence which is cleaved from the mature protein once export of the growing protein chain across the rough endoplasmic reticulum has been initiated. Those of ordinary skill in the art are aware that polypeptides secreted by vertebrate cells generally have a signal peptide fused to the N-terminus of the polypeptide, which is cleaved from the complete or "full length" polypeptide to produce a secreted, or "mature" form of the polypeptide. In certain embodiments, the native leader sequence is used, or a functional derivative of that sequence that retains the ability to direct the secretion of the polypeptide that is operably associated with it. Alternatively, a heterologous mammalian leader sequence, or a functional derivative thereof, may be used. For example, the wild-type leader sequence may be substituted with the leader sequence of human tissue plasminogen activator (TP A) or mouse beta-glucuronidase.
The term "expression" refers to the biological production of a product encoded by a coding sequence. In most cases a DNA sequence, including the coding sequence, is transcribed to form a messenger-RNA (mRNA). The messenger-RNA is then translated to form a polypeptide product which has a relevant biological activity. Also, the process of expression may involve further processing steps to the RNA product of transcription, such as splicing to remove introns, and/or post-translational processing of a polypeptide product. As used herein, the term "polypeptide" is intended to encompass a singular "polypeptide" as well as plural "polypeptides," and comprises any chain or chains of two or more amino acids. Thus, as used herein, terms including, but not limited to "peptide," "dipeptide," "tripeptide," "protein," "amino
acid chain," or any other term used to refer to a chain or chains of two or more amino acids, are included in the definition of a "polypeptide," and the term "polypeptide" can be used instead of, or interchangeably with any of these terms. The term further includes polypeptides which have undergone post-translational modifications, for example, glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting/blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids.
Also included as polypeptides of the present invention are fragments, derivatives, analogs, or variants of the foregoing polypeptides, and any combination thereof. Polypeptides, and fragments, derivatives, analogs, or variants thereof of the present invention can be antigenic and immunogenic polypeptides related to HCV polypeptides, which are used to prevent or treat, i.e., cure, ameliorate, lessen the severity of, or prevent or reduce contagion of infectious disease caused by the HCV.
As used herein, an "antigenic polypeptide" or an "immunogenic polypeptide" is a polypeptide which, when introduced into a vertebrate, reacts with the vertebrate's immune system molecules, i.e., is antigenic, and/or induces an immune response in the vertebrate, i.e., is immunogenic. It is quite likely that an immunogenic polypeptide will also be antigenic, but an antigenic polypeptide, because of its size or conformation, may not necessarily be immunogenic. Isolated antigenic and immunogenic polypeptides of the present invention in addition to those encoded by polynucleotides of the invention, may be provided as a recombinant protein, a purified subunit, a viral vector expressing the protein, or may be provided in the form of an inactivated HCV vaccine, e.g., a live-attenuated virus vaccine, a heat-killed virus vaccine, etc. By an "isolated" HCV polypeptide or a fragment, variant, or derivative thereof is intended an HCV polypeptide or protein that is not in its natural form. No particular level of purification is required. For example, an isolated HCV polypeptide can be removed from its native or natural environment. Recombinantly produced HCV polypeptides and proteins expressed in host cells are considered isolated for purposed of the invention, as are native or recombinant HCV polypeptides which have been separated, fractionated, or partially or substantially purified by any suitable technique, including the separation of HCV virions from culture cells in which they have been propagated. In addition, an isolated HCV polypeptide or protein can be provided as a live or inactivated viral vector expressing an isolated HCV polypeptide and can include those found in inactivated HCV vaccine compositions. Thus, isolated HCV polypeptides and proteins can be provided as, for example, recombinant HCV polypeptides, a purified subunit of HCV, a viral vector expressing an isolated HCV polypeptide, or in the form of an inactivated or attenuated HCV vaccine.
The term "epitopes," as used herein, refers to portions of a polypeptide having antigenic or immunogenic activity in a vertebrate, for example a human. An "immunogenic epitope," as used herein, is defined as a portion of a protein that elicits an immune response in an animal, as determined by any method known in the art. The term "antigenic epitope," as used herein, is defined as a portion of a protein to which an antibody or T-cell receptor can immunospecifically bind as determined by any method well known in the art. Immunospecific binding excludes nonspecific binding but does not exclude cross-reactivity with other antigens. Where all immunogenic epitopes are antigenic, antigenic epitopes need not be immunogenic.
As to the selection of peptides or polypeptides bearing an antigenic epitope (e.g., that contain a region of a protein molecule to which an antibody or T cell receptor can bind), it is well known in that art that relatively short synthetic peptides that mimic part of a protein sequence are routinely capable of eliciting an antiserum that reacts with the partially mimicked protein. See, e.g., Sutcliffe, J. G., et al., 1983, Science 219:660-666.
The present invention also provides methods for identifying transmitted full-length HCV genomes that mediate viral infection and transmission. In one embodiment, the methods comprise: (a) collecting a patient sample; (b) isolating viral RNA from said sample; (c) sequencing said viral RNA, wherein viral RNA sequencing includes HCV genomes of circulating virus; (d) performing sequence alignment of selected HCV genome regions; (e) analyzing phylogenetically selected sequence alignments; and (f) identifying HCV genomes of transmitted virus.
As used herein, a "patient" or "subject" to be utilized by the disclosed methods can mean a human, chimpanzee or non-human primate. In certain embodiments the patient is a non-human mammal.
The term "patient sample" as used herein includes but is not limited to a blood, serum, plasma, or urine sample obtained from a patient. In particular embodiments, the patient sample is plasma.
The phrase "analyzing phylogentically" as used herein represents the analysis of clinical viral isolates, regardless of the particular methodology employed. Comparative analysis of the genetic relatedness of any a collection of circulating viral isolates is used to select for nucleotide sequence of transmitted HCV genomes. This methodology is used to select for the actual HCV genomes present at the time of patient infection/viral transmission. This can be accomplished by any number of methods, including but not limited to: i) a novel mathematical model of random virus evolution disclosed herein; ii) star phylogeny; iii) Baysian analysis; or any other method of phylogenetic analysis known to one of skill in the art.
The phrase "performing sequence alignment" as used herein is meant to include aligning genetic sequences by any of a number of different procedures that produce a sufficient match between the corresponding residue in the sequences. Typically, Smith- Waterman or Needleman- Wunsch algorithms are used. However, other procedures such as BLAST, FASTA, PSI-BLAST can be used.
The phrase "isolating viral RNA" as used herein includes those methods well known in the art for extraction of viral RNA from a sample, cDNA synthesis from viral RNA template, optionally cloning of cDNA fragments, and/or amplification of polynucleotide sequence. Such methods are described in more detail, but such experimental procedures are provided in Sambrook and Russell, 2001, Molecular Cloning: A Laboratory Manual, Third Ed., Cold Spring Harbor Laboratory Press, Woodbury, N.Y.
The practice of this invention can involve procedures well known in the art, including for example nucleotide sequence amplification, such as polymerase chain reaction (PCR) and modifications thereof (including for example reverse transcription (RT)-PCR, and stem-loop PCR), as well as reverse transcription and in vitro transcription. Generally these methods utilize one or a pair of oligonucleotide primers having sequence complimentary to sequences 5' and 3' to the sequence of interest, and in the use of these primers they are hybridized to a nucleotide sequence and extended during the practice of PCR amplification using DNA polymerase (preferably using a thermal-stable polymerase such as Taq polymerase). RT-PCR may be performed on miRNA or mRNA with a specific 5' primer or random primers and appropriate reverse transcription enzymes such as avian (AMV-RT) or murine (MMLV-RT) reverse transcriptase enzymes.
The phrase "over time" as used herein represents a period of time between the collection of patient samples. For example, a patient sample is collected at time point A and then subsequently as a later date at time point B. The period between samplings is over time. In certain embodiments, samples are taken from the same patient. In alternative embodiments, samples may be taken from different patients. Performing the method of the invention at different time points followed by the differential analysis of HCV genomes at the different time points provides a means for assessing the evolution of HCV genomes both inter- and intra- patient. The identification of highly variable genomic regions is useful for the generation of effective therapeutics and vaccines to HCV.
The term "transmitted" as used herein refers to viral genotype and phenotype at the time of viral infection. The term "transmitted virus" as used here is in reference to actual HCV virus that mediates patient infection. The phrase "transmitted viral sequence" as used herein is meant
to include nucleotide or amino acid sequence of HCV genomes of transmitted virus, an in a particular embodiment sequence corresponding to acute infection stage HCV. The term "circulating" refers to HCV virus collected from patient post-infection. In general, patient samples comprise circulating virus because HCV infection has occurred at a prior time point. The half-life of plasma virus is less than 1 day, thus the collection of transmitted virus from a patient sample would be a rare event.
The term "selected" as used in the phrases "selected sequence alignments" or "selected HCV genome regions" is meant to include the identification and utilization of particular subsets of nucleotide sequence and/or genome regions for subsequent analysis. In certain embodiments, full-length HCV genomes are identified, however discrete portions of the genome are utilized or "selected" for further analysis.
It is to be noted that the term "a" or "an" entity refers to one or more of that entity; for example, "a polynucleotide," is understood to represent one or more polynucleotides. As such, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein.
The present invention also provides immunogenic compositions and methods for delivery of transmitted full-length HCV polynucleotide or polypeptide sequences to a vertebrate with optimal expression and safety conferred. In particular instances the transmitted full-length HCV genome comprises the polynucleotide of SEQ ID NO. 776. These immunogenic compositions may be prepared and administered in such a manner that the encoded gene products are optimally expressed in the vertebrate of interest. As a result, these compositions and methods are useful in stimulating an immune response against HCV infection. Also included in the invention are expression systems and delivery systems.
The identified polynucleotides or polypeptides encoded by the polynucleotides of the invention may be in any form, and polypeptides are generated using techniques well known in the art. Examples include isolated HCV proteins produced recombinantly or proteins delivered in the form of an inactivated HCV vaccine, such as conventional vaccines.
When utilized, an isolated HCV polynucleotide or polypeptide or fragment, variant or derivative thereof is administered in an immunologically effective amount. The effective amount of conventional vaccines is determinable by one of ordinary skill in the art based upon several factors, including the antigen being expressed, the age and weight of the subject, and the precise condition requiring treatment and its severity, and route of administration.
In the instant invention, the combination of conventional antigen vaccine compositions with optimized nucleic acid or polypeptide compositions provides for therapeutically beneficial effects at dose sparing concentrations. For example, immunological responses sufficient for a
therapeutically beneficial effect in patients predetermined for an approved commercial product, such as for the conventional product described above, can be attained by using less of the approved commercial product when supplemented or enhanced with the appropriate amount of nucleic acid or polypeptide.
A desirable level of an immunological response afforded by a DNA based pharmaceutical alone may be attained with less DNA by including an aliquot of a conventional vaccine. Further, using a combination of conventional and DNA based pharmaceuticals may allow both materials to be used in lesser amounts while still affording the desired level of immune response arising from administration of either component alone in higher amounts (e.g. one may use less of either immunological product when they are used in combination). This may be manifest not only by using lower amounts of materials being delivered at any time, but also to reducing the number of administrations points in a vaccination regime (e.g. 2 versus 3 or 4 injections), and/or to reducing the kinetics of the immunological response (e.g. desired response levels are attained in 3 weeks instead of 6 after immunization).
Determining the precise amounts of DNA based pharmaceutical and conventional antigen is based on a number of factors as described above, and is readily determined by one of ordinary skill in the art.
The ability of an adjuvant to increase the immune response to an antigen is typically manifested by a significant increase in immune -mediated protection. For example, an increase in humoral immunity is typically manifested by a significant increase in the titer of antibodies raised to the antigen, and an increase in T-cell activity is typically manifested in increased cell proliferation, or cellular cytotoxicity, or cytokine secretion.
Nucleic acid molecules and/or polynucleotides of the present invention, e.g., plasmid DNA, mRNA, linear DNA or oligonucleotides, may be solubilized in any of various buffers. Suitable buffers include, for example, phosphate buffered saline (PBS), normal saline, Tris buffer, and sodium phosphate (e.g., 150 mM sodium phosphate). Insoluble polynucleotides may be solubilized in a weak acid or weak base, and then diluted to the desired volume with a buffer. The pH of the buffer may be adjusted as appropriate. In addition, a pharmaceutically acceptable additive can be used to provide an appropriate osmolarity. Such additives are within the purview of one skilled in the art. For aqueous compositions used in vivo, sterile pyrogen-free water can be used. Such formulations will contain an effective amount of a polynucleotide together with a suitable amount of an aqueous solution in order to prepare pharmaceutically acceptable compositions suitable for administration to a human.
Compositions of the present invention can be formulated according to known methods. Suitable preparation methods are described, for example, in Remington's Pharmaceutical Sciences, 16th Edition, A. Osol, ed., Mack Publishing Co., Easton, Pa. (1980), and Remington's Pharmaceutical Sciences, 19th Edition, A. R. Gennaro, ed., Mack Publishing Co., Easton, Pa. (1995). Although the composition may be administered as an aqueous solution, it can also be formulated as an emulsion, gel, solution, suspension, lyophilized form, or any other form known in the art. In addition, the composition may contain pharmaceutically acceptable additives including, for example, diluents, binders, stabilizers, and preservatives.
The invention illustratively described herein suitably can be practiced in the absence of any element or elements, limitation or limitations that are not specifically disclosed herein. Thus, for example, in each instance herein any of the terms "comprising", "consisting essentially of, and "consisting of may be replaced with either of the other two terms, while retaining their ordinary meanings. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention that in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by embodiments, optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the description and the appended claims.
This invention is more particularly described below and the Examples set forth herein are intended as illustrative only, as numerous modifications and variations therein will be apparent to those skilled in the art. As used in the description herein and throughout the claims that follow, the meaning of "a", "an", and "the" includes plural reference unless the context clearly dictates otherwise. The terms used in the specification generally have their ordinary meanings in the art, within the context of the invention, and in the specific context where each term is used. Some terms have been more specifically defined below to provide additional guidance to the practitioner regarding the description of the invention.
Examples
The Examples which follow are illustrative of specific embodiments of the invention, and various uses thereof. They are set forth for explanatory purposes only, and are not to be taken as limiting the invention.
Example 1
HCV Patient Selection and Analysis
HCV patients were selected and samples collected as described. Plasma samples were obtained from 17 subjects with acute or very recent HCV infection representing subtypes la, lb, 2 or 3. These consisted of twice-weekly serial collections from source plasma donors who became HCV infected during the course of their plasma donations. The donors were untreated and asymptomatic throughout the collection period. Plasma samples from 14 subjects with chronic HCV infection from the U.S. served as controls. All patients were treatment-na'ive. All subjects gave informed consent, and plasma collections were performed with institutional review board and other regulatory approvals. Plasma samples were tested for HCV RNA and viral specific antigen and antibodies by a battery of commercial tests (Abbott and Roche) (Fig. 1 and Table 2). In particular, a total of 154 samples were tested including a median of 8 sequential specimens per acutely infected subject (range 5-11). The initial samples from acute subjects were HCV vRNA and antibody negative, followed by a sharp rise in vRNA levels. Five of 17 acutely infected subjects developed HCV antibodies by the last sampling time point. Chronic subjects had a median vRNA load of 1,975,569 IU/ml (range = 24,000 - 6,400,000 IU/ml) and all were HCV antibody positive. Viral load measurements were typical of acute and chronic HCV infection.
Viral RNA isolation and cDNA synthesis was performed as follows. Samples which contained approximately 100,000 viral RNA copies were extracted using the Qiagen BioRobot EZ1 Workstation with EZ1 Virus Mini Kit v2.0 (Qiagen, Valencia, CA). RNA was eluted in 60 μΐ and immediately subjected to cDNA synthesis. Reverse transcription of RNA to single stranded cDNA was performed using Superscript III reverse transcriptase (Invitrogen Life Technologies, Carlsbad, CA). For a 60 μΐ of reaction, a mixture of 3 μΐ (0.5 mM) of 10 mM deoxynucleoside triphosphate, 1.5 μΐ (0.25 μΜ) of 10 μΜ anti-sense primer and 34.5 μΐ of vRNA was incubated at 65°C for 5 min and then was placed on ice for at least 1 min. The contents of the tube were collected by brief centrifugation. Next, 12 μΐ of 5X reaction buffer, 3 μΐ of 0.1 M DTT (5 mM), 3 μΐ of RNAseOUT Recombinant RNASE Inhibitor, (40 units/μΐ, Invitrogen) and 3 μΐ of Superscript III™ Reverse Transcriptase (2001Ι/μ1) were added to the mixture. The final reaction was incubated at 50°C for 60 min followed by an increase in temperature to 55°C for an additional 60 minutes. The reaction was heat-inactivated at 70°C for 15 minutes and then treated with RNaseH at 37°C for 20 minutes. SEQ ID NOS The antisense primers were designed specifically for different genotype. 1.NS4A-R1 5'- GCACTCTTCC ATCTCATCGAACTC-3 ' (SEQ ID NO: 758) (nt 5451-5474 H77 (accession
number NC 004102)) for genotype 1, 2NS2-R1 5'-CCCCAGACGATGACTTTCTTCTCCAT- 3' (SEQ ID NO: 777) (nt 5445-5467 H77) for genotype 2 and 3aNS3-R2V2 5'- TTACTTCC AGATCAGCTGAC A-3 ' (SEQ ID NO: 772) for genotype 3. The newly synthesized cDNA was used immediately or kept frozen at -80°C.
Single genome amplification was performed from prepared viral cDNA. cDNA was serially diluted and distributed among wells of replicate 96-well plates so as to identify a dilution where PCR positive wells constituted less than 30% of the total number of reactions. At this dilution, most wells contained amplicons derived from a single cDNA molecule. This was confirmed in every positive well by direct sequencing of the amplicon and inspection of the sequence for mixed bases (double peaks), which would be evidence of priming from more than one original template or the introduction of PCR error in early cycles. Any sequence with evidence of mixed bases was excluded from further analysis. PCR amplification was carried out in the presence of 10 x Taq High Fidelity Platinum PCR buffer, 2 mM MgS04, 0.2 mM of each deoxynucleoside triphosphate, 0.2 μΜ of each primer, and 0.1 μ(0.5 Unit) Platinum Taq High Fidelity polymerase in a 20 μΐ reaction (Invitrogen, Carlsbad, CA).
The nested or hemi-nested primers for generating 5' half and 5' quarter genome from different genotypes included: (1) genotype 1 : 1st round sense primer l .core.Fl 5'- ATGAGCACGAATCCTAAACCTCAAAGA-3' (SEQ ID NO:761)(nt 342-368 H77) and 1st round antisense primer 1.NS4A.R1 5'-GCACTCTTCCATCTCATCGAACTC-3' (SEQ ID NO:763) (nt 5451-5474 H77), 2nd round sense primer l .core.F2 5 ' -TC AAAG AAAAAC C AAA CGTAACACCAACCG-3' (SEQ ID NO:764) (nt 362-391 H77) and 2nd round antisense primer 1.NS3A4A.R2 5'-AGGTGCTCGTGACGACCTCCAGG-3' (SEQ ID NO:766) (nt 5297-5319 H77); (2) 5' quarter genome of genotype 2: 1st round sense primer 2.core.Fl 5'- ATGAGCACA AATCCTAAACCTCAAAGA-3' (SEQ ID NO: 767) (nt 342-368 H77) and 1st round antisense primer 2.NS2.R1 5'-CCCCACACAATGACCTTCTTCTCCATTG-3' (SEQ ID NO: 778) (nt 5445-5467 H77), 2nd round sense primer 2.core.F2 5'-AATCCTAAACCTCAAAGAAAAACC AAA-3' (SEQ ID NO: 769) (nt 351-377 H77) and 2nd round antisense primer 2.NS2.R2 5'-GG GGAGAGGTGGTCATAGATGTAA -3 '(SEQ ID NO 779); (3) 5' half genome of genotype 3 : 1st round sense primer 3a.core.Fl 5'-ATGAGCACACTTCCTAAACCTCAAAGA-3' (SEQ ID NO: 771) and 1st round antisense primer 3aNS3-R2V2 5 ' -TTACTTCCAGATCAGCTGACA- 3 '(SEQ ID NO: 760), 2nd round sense primer 3a.core.F2 5 ' -TC AAAG AAAAACC AAAAGAAA CACCATCCG-3' (SEQ ID NO: 773) and 2nd round antisense primer PCR 3a.NS3-R2V2 5'-TT ACTTCCAGATCAGCTGACA -3 '(SEQ ID NO. 774). Strain specific primers were used to amplify the quarter genomes from the early time point samples. The positive PCR reactions
were subject to direct sequencing. PCR was performed in MicroAmp 96-well reaction plates (Applied Biosystems, Foster City, CA) with the following PCR parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94 °C for 15 s, an annealing step of 58°C for 30 s, an extension step of 68 °C for 5 min, followed by a final extension of 68 °C for 10 min. The product of the 1st round PCR was subsequently used as a template in the 2nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on precasted 1% agarose E- gels 96 (Invitrogen Life Technologies, Carlsbad, CA). All PCR procedures were carried out under PCR clean room conditions using procedural safeguards against sample contamination, including pre-aliquoting of all reagents, use of dedicated equipment, and physical separation of sample processing from pre- and post-PCR amplification steps.
For DNA sequencing, 5' half-genome amplicons were directly sequenced by cycle- sequencing using BigDye terminator chemistry and protocols recommended by the manufacturer (Applied Biosystems; Foster City, CA). Sequencing reaction products were analyzed with an ABI 3730x1 genetic analyzer (Applied Biosystems; Foster City, CA). Both DNA strands were sequenced using partially overlapping fragments. Individual sequence fragments for each amplicon were assembled and edited using the Sequencher program 4.7 (Gene Codes; Ann Arbor, MI). Inspection of individual chromatograms allowed for the identification of amplicons derived from single versus multiple templates. The absence of mixed bases at each nucleotide position throughout the entire 5 ' half-genome sequences was taken as evidence of single genome amplification from a single viral RNA/cDNA template. This quality control measure allowed the exclusion of amplicons that resulted from PCR-generated in vitro recombination events or Taq polymerase errors and to obtain multiple individual sequences that proportionately represented those circulating HCV virions. All the sequences alignments were initially made with ClustalW and then hand-checked using MacClade 4.08 to improve the alignments according to the codon translation.
Among the 3000 amplicons generated, the sequences of 2850 were unambiguous at every position. Sequences chromatograms from 150 amplicons contained one or more "double peaks" representing mixed bases. Because mixed bases generally represented only a minority of polymorphisms, it was inferred that such ambiguities had resulted from Taq polymerase errors in the initial PCR cycles and not from amplification from more than one initial target template; in such cases a correction of the assignment of the base was made. In rare instances where the only polymorphic site(s) present were represented by double peaks, possibly due to mixed initial templates, the sequences were discarded and not included in the analysis. This amounted to 30 sequences or less than 1% of total amplicons generated.
Example 2
Phylogenetic Analysis: Identification and Enumeration of Transmitted Viruses
A total of 3000 Sequences corresponding to the core antigen genes from the initial time point from all 17 acutely-infected and 14 chronically-infected subjects were analyzed using neighbor-joining (NJ) phylogenetic tree methods together with a sequence visualization tool, Highlighter (www.HIV.lanl.gov). Phylogenetic trees were generated by the Maximum Likelihood (ML) method. For the subjects with multiple transmitted/founder viral lineages, the lineages that contained more than 5 identical or near identical sequences were included in the subsequent diversity analyses. The maximum sequence diversity within subject or lineage was calculated using Poisson Fitter program (www.HIV.lanl.gov).
A total of 38 transmitted/founder lineages from 17 acutely infected subjects were used to calculate the mutation rate. Each sequence within the lineage was compared with the T/F virus sequence of that lineage. The insertion, deletion, transition and transversion frequencies were counted by self-developed computer program and by Highlighter tool. The rates were calculated by taking the ratio of each frequency number and the total number of nucleotides of all the sequences within that lineage.
To determine the Synonymous and non-synonymous substitution rate, the SNAP program (www.HIV.lanl.gov) was applied to the codon-aligned sequences of each T/F lineage. Within each lineage, the accumulations of synonymous and nonsynonymous substitutions were counted by comparing to the transmitted/founder viral sequences. The Jukes-Cantor corrected accumulation rates of synonymous substitutions per potential synonymous site (ds) and nonsynonymous substitution per potential nonsynonymous site (dn) were compared to screen for positive selection.
To determine the likelihood of missing infrequent transmitted variants, a power study was performed to investigate the probability of sampling limitations. With a sample of at least n = 20 plasma vRNA sequences, a 95 % confidence level existed that a given missed variant comprised less than 15 % of the virus population. For samples for which n > 30, a 95 % confidence level existed that the results did not miss any variant that comprised at least 10 % of the total viral population.
Sequences from each subject formed monophyletic clades (bootstraps 99-100 %) corresponding to HCV genotypes la (n = 23), lb (n = 4), 2b (n = 2) and 3a (n = 2). In no case was a subject infected by more than one virus clade and there was no intermixing of sequences between subjects in the phylogenetic tree. Sequences from chronic subjects (blue shading) and acute subjects (red shading) showed variable degrees of within- subject diversity. Acute
sequences were distinct, however, in exhibiting one or more discrete sublineages characterized by generally large numbers of sequences with extremely low diversity.
Sequences from chronic subject WIMI4025 showed broad genotypic heterogeneity with a maximum diversity of 3.97 % (median 1.24 %; range 0.12-3.97 %) (Fig 3A). Sequences from acute subject 10051 revealed a very different pattern of diversification (Fig 3B). These sequences, which were derived from the last sampled time point 21 days after the beginning of documented viremia, were very homogeneous with a maximum diversity of 0.18 % (median 0 %; range 0-0.18 %). The maximum diversities of sequence lineages from all 17 acutely infected subjects were similarly low, ranging from 0.04 % to 0.22 % (median = 0.1 1 %). This extremely limited diversity within sequence lineages from acutely infected subjects was significantly lower than the maximum viral diversity in chronic subjects (0.56 % to 3.97 %; median = 2.51 %; p<0.000X) (Tables 1 and 2). Nucleotide polymorphisms in subject 10051 were essentially random, corresponding to a star-like phylogeny and a Poisson distribution of low frequency events. Two sequences (2C3 and 2C2) contained a single common polymorphism at position -2200, and two others (2A2 and 2B34) contained a different common polymorphism at position -4375, indicating for each pair shared recent ancestry. Rare shared polymorphisms like these in newly infected subjects can be explained as having been generated subsequent to virus transmission as a consequence of RdRp errors very early in infection when the effective number (Ne) of productively infected cells is relatively low. The 60 sequences depicted in Fig. 3B coalesced to a single unambiguous consensus that was inferred, based on models of random virus evolution and empirical testing to represent the T/F virus in this subject. However, to obtain additional evidence that the sequences coalesced to a virus at or near the moment of transmission and not an intermediate time point, 243 additional sequences from three earlier time points were sampled. Whether sequences from each time point were considered separately or altogether, they coalesced to the same T/F genome (Table 1). Power calculations indicate that a sample size of 60 sequences as shown in Fig. 3B provide 95 % likelihood of detecting variants present at 5% in the population. A sample size of 300 sequences provides 95 % likelihood of detecting variants present at 1 % in the population. Thus, it was concluded that subject 10051 was productively infected by a single virus.
For subjects 10021 and 10025, at the initial time points, 19 of 24 (79 %) of 10021 sequences and 32 of 43 (74 %) of 10025 sequences were identical with the remainder differing from the respective consensus sequences by only 1 or 2 nucleotides. At the second and third time points 1 - 4 weeks later, an increasing proportion of sequences differed from the respective consensus sequences but only by 3 or 4 nucleotides. In each subject, sequences varied
essentially randomly in a star-like fashion that fit a Poisson model of low frequency independent mutations. For each subject it was concluded that the consensus sequence corresponded to a single, unique T/F HCV genome. Altogether, of the 17 acute subjects in the present study, 4 had evidence of productive clinical infection by single viruses (Table 1). Figures 4C and 4D depict sequences from subjects having evidence of productive infection by more than one genetically distinct virus. Each of the subjects was sampled on 4 occasions and sequences again showed increasing diversity over time. The consensus sequences of each low diversity lineage in these subjects were interpreted to correspond to a unique T/F HCV genome. Thus, each of these subjects was productively infected by at least three viruses. The proportion of sequences represented in each lineage was approximately the same in subject 10012 (Fig 4C) but not in subject 10062 (Fig 4D). This could be due either to different replication rates of the different T/F viral viruses, differential effects of early innate or adaptive immune responses, different times of infection by different transmitted viruses, or early stochastic events in the infection process. A mathematical model was developed to estimate the adequacy of sampling given the observed distribution of T/F lineage sequences and total number of sequences analyzed. This model, when applied to the 179 sequences analyzed for subject 10012, estimated that the actual number of T/F viruses was 3 at a 95 % confidence level. In contrast, when applied to the 164 sequences analyzed for subject 10062, the model estimated that the actual number of T/F viruses could be as high as 5, again at a 95 % confidence level. Sequences from subject 10029 where low diversity sequence lineages corresponding to 9 T/F viruses could be unambiguously identified, but with even greater differences in sequence proportions. The model estimated that the actual number of T/F viruses could be as high as 13 at a 95 % confidence level. Comparable patterns of virus diversification following multivariant HCV transmission were observed in six additional subjects (Figs 6-12). Importantly, in contrast to HIV-1 where viral recombination in acute and early infection is exceedingly common and widespread, there was no evidence of viral recombination in any subject acutely infected by HCV. In addition, it was not observed in any of 13 subjects whose plasma virus was sequenced at multiple time points, the appearance of a distinct lineage of virus different from those sampled at the initial time points, indicating that there was no evidence of superinfection or substantial undersampling at the initial time points.
A mathematical approximation of early HCV evolution based on estimated parameters of RdRp error rate, virus generation time, infected cell lifespan, and viral reproductive ratio suggests that maximum viral diversity 10 weeks post-infection by a single virus is 0.4 % (95 % C.I. = 0.25 - 0.55 %), which is well within empirical determinations in these experiments (Table 1). With this as an upper bound, the diversity among T/F sequences from acutely infected
subjects were evaluated to look for evidence of virus transmission from subjects who themselves were acutely infected. A similar approach has been successful in detecting acute-to-acute transmission of HIV- 1. If an acutely-infected subject acquired multiple HCV genomes from a contact who himself was acutely infected by a single virus, then these T/F viruses would be expected to all be closely related to each other, since they must reflect the viral diversity in the donor. If that donor contact had been acutely infected by more than one genetically divergent virus and the acutely infected recipient acquired multiple progeny of these, then the expectation would be for T/F viral sequences to be represented by subsets of sequences, some closely related and some not. Evidence in five subjects of acute-to-acute transmission was found by these two scenarios. Figures 13A-C depict ML trees and Highlighter plots from acutely infected subjects 10020, 10016 and 10003, where the maximum diversity of T/F viruses from a consensus was 0.18 %, 0.16 % and 0.12 %, respectively. The phylogenetic patterns of these sequences were quite different from the single variant transmissions shown in Figs 3B and 4 A and 4B, which exhibited star-like phylogeny and conformed to a Poisson distribution of random mutations. The sequences in 13A-C, instead, violated the Poisson model. They were comprised of distinct subsets of highly related sequences that differed from each other and from a common consensus by 1-6 nucleotides. If, however, the sequence subsets within each subject were considered independently, then their diversity conformed to the Poisson model. Thus, it was concluded that these three subjects were each productively infected by as many as 9-16 genetically distinct T/F viruses. The other scenario of acute -to-acute transmission is virus acquisition from a donor who is acutely infected by multiple viruses, where the expectation is for a multimodal distribution of diversity among T/F viruses in the recipient. This was observed in two subjects (Figs. 5 and 12).
The phylogenetic pattern of sequences from subject 106889 (Fig. 14) was far more complicated than that of any of the other 16 acutely infected subjects. Subject 106889 exhibited a typical acute infection viral kinetic profile with four sequential plasma samples negative for HCV vRNA followed by rapid vR A ramp-up to nearly 107 vRNA IU/ml (Fig 14 insert). Eighty-six 5 ' half genome sequences were obtained from the initial plasma vRNA positive time point. The recent evolutionary history of these sequences and the clinical circumstances of virus transmission could be inferred based on the ML tree, the Highlighter plot, and two additional lines of evidence: First, subject 106889 was a persistently HCV antibody-negative and HCV RNA-negative source plasma donor who had undergone regular biweekly plasmaphereses. Thus, when the subject became HCV viremic, it was clear that this was due to a new infection and not recrudescence of an old one. Second, 85 of the 86 sequences from subject 106889 depicted in Fig. 14 contained two signature mutations (V36M and R155R) in the NS3 protease that
conferred high level resistance to the protease inhibitors Boceprevir and Telaprevir. A single sequence (02B11) contained only one of these mutations (V36M). It is clinically and biologically implausible for subject 106889 to have been diagnosed with acute HCV infection, to have been treated with an NS3 protease inhibitor (which was investigational at the time), and to have developed drug resistance, all in the span of one week. Instead, the transmitting partner to subject 106889 must have been chronically infected by HCV and treated with an NS3 protease inhibitor leading to narrowing of viral diversity and emergence of DAA drug resistance. Then, virus transmission to subject 106889 occurred, with the ML tree and Highlighter plot in Fig. 14 showing evidence of more than 30 T/F viruses. Mathematical modeling based on the Jackknife estimation tool provided a 95 % upper limit on the estimated number of T/F viruses in this subject at 48 (Table 3). The T/F viruses in subject 106889 are evident as discrete sublineages (designated as variants, v) within lineages designated Ll-10. Thus, in lineage 1 (LI), the progeny of as many as nine closely related T/F variants (vl-9) are evident. In lineage 2 (L2), the progeny of as many as six closely related T/F variants (vl-6) are evident (Fig. 14). This phylogenetic pattern of the progeny of multiple T/F variants within any one lineage is indistinguishable from T/F virus progeny in acutely infected subjects 10020, 10016 and 10003 (Figs. 13A-C). Lineage 5 (L5) in subject 106889 is comprised of sequences arising from a single T/F variant (vl), which is indistinguishable from examples of single variant transmission (subject 10051; Fig. 3B). Importantly, T/F lineages with sufficient numbers of sequences for model testing (e.g., L2-vl and L5-vl), conformed to a star-like phylogeny with a Poisson distribution mutations (Fig. 14; Table 1).
The identification of T/F viral sequences in 17 subjects provided a unique strategy for analyzing HCV sequence evolution in vivo in the critical period beginning at or near the moment of virus transmission and extending to the establishment of viral load setpoint and in some subjects HCV antibody seroconversion. A summary of this analysis is presented in Table 1. Maximum intra-lineage diversity for the 17 subjects ranged from 0.06 % to 0.22 %. Insertions (0.000001), deletions (0.00004) and stop codons (0.000004) were infrequent. Transitions outnumbered transversions by 8 to 1. The overall mutation frequency including all sampled time points was low (0.000145) given the number of possible virus replication cycles between the first infected hepatocyte and setpoint viremia 6-10 weeks later when as many as 7-20 % of hepatocytes may be productively infected. The dN/dS ratio was low at 0.33 as was its derivative pN/pS at 0.33 %.
The present study builds on a body of previous work aimed at characterizing the genetic bottleneck to HCV transmission and early and late patterns of virus diversification. Although
previous studies demonstrated a transmission bottleneck, they could not identify with precision actual T/F viral genomes that were responsible for productive clinical infection, nor could they discriminate between T/F genomes that differed by as few as 1 to 3 nucleotides in 10,000 (0.01- 0.03% diversity). In addition, they could not evaluate virus diversity unencumbered by Taq polymerase-mediated nucleotide substitution errors or recombination artifacts that prevented mutational linkage to be analyzed across genes and genomes. The present study addressed these objectives by testing the hypothesis that SGA-sequencing could enable an unambiguous molecular identification and enumeration of T/F HCV genomes and their progeny. Although these goals had been previously reached for HIV-1 and simian immunodeficiency virus (SIV), it was unclear if a similar result could be achieved for HCV given the differences in viral replication strategies of the two viruses, the variable durations of the eclipse phases, differences in viral target cells (long-lived hepatocytes versus short-lived lymphocytes), the higher estimates of HCV RdRp error, and the possibility that higher multiplicities of infection by HCV in injection drug users could confound the identification of T/F viral lineages. Despite these challenges and uncertainties, the empirical findings of the present study make clear that even in complicated cases of multivariant HCV transmission in the setting of community-acquired infection, T/F viral genomes can be unambiguously and routinely identified.
With few exceptions, the diversity of sequences within the T/F lineages of each acutely infected subject corresponded to a star-like pattern of random diversification that conformed to a Poisson distribution mutations. Exceptions were of four types: (i) infrequent shared polymorphisms resulting from stochastic changes in the newly infected subjects (e.g., Fig 3B; 4A, 4B, 4C, 4D); (ii) transmission of multiple closely related variants (e.g., Figs 11, 13 A, 13B; 14); (iii) evidence of immune selection in later samples (Fig 12); and (iv) rare examples of short inverted repeats resulting from template switching between double-stranded RNA duplexed hairpin structures by the RNA dependent polymerase (Fig 13). The first three of these exceptions are also found in early HIV-1 and SIV diversification and can be explained by models of virus replication and diversification. The last exception, however, was unique to HCV. Among the 2670 amplicons sequenced, four sequences exhibited short concentrated stretches of 3 to 20 nucleotide substitutions compared with other sequences from the same subjects, including the T/F consensus sequences. In each of the four instances, the nucleotide mismatches represented perfect inverted repeats. This implies that double-stranded viral RNA served as the replication template for these sequences, although it cannot determine if they are the result of HCV RdRp-mediated replication in vivo or the product of cDNA synthesis from in vitro generated reverse transcription products of the MuLV reverse transcriptase. Finally, in any
of the sequences analyzed no evidence of HCV recombination in vivo was observed, which would have been plainly evident in those subjects infected by multiple genetically diverse viral genomes (Fig 4C, D; 5; 6-11). Absence of viral recombination in vivo distinguishes HCV from HIV-1 and is consistent with epidemiological data showing that HCV recombination is rare.
Because early HCV sequences could be mapped precisely to T/F sequences, the genetic pathways of early HCV evolution were examined in a manner not previously possible, both at the level of nucleotide substitution and amino acid selection. For the proteome, one subject (9055) had evidence of a non-random accumulation of changes in a previously identified HLA- restricted cytotoxic T-cell (CTL) epitope. Since the HLA type of this subject was not known, it could not verified if this peptide sequence represented a CTL epitope, but it is likely based on parallel observation in HIV-1 that this region of HCV represents CTL escape or reversion. CTL responses and virus escape or reversion are generally believed to occur later in infection but early mutations could have been overlooked if sequences were not mapped to T/F proteomes. Again, this was a scenario observed in acute HIV-1 infection where identification of T/F genomes and proteomes enabled the detection of earliest CTL recognition and escape. Non- random amino acid polymorphisms in multiple T/F sequences in subjects 10020 (Fig 13 A) and 10016 (Fig 13B) also corresponded to known or predicted CTL epitopes and thus are likely to represent CTL escape or reversion in the respective donors (not in subjects 10020 or 10016). At the genome level, it was determined the frequency of infections, deletions, and termination codons were low (between 0.000001-0.000004) and the dN/dS ratio to be similarly low at 0.33, consistent with purifying or negative selective early in infection. Interestingly, the frequency of transitions to exceeded that of transversions by a factor of 8: 1, again consistent with a bias toward negative selection. Gotte et al have recently provided a biochemical explanation for this observation based on a strong bias for G:U/U:G mismatches by the HCV RdRp. The frequency of nucleotide substitutions in evolved sequences compared with T/F sequences was overall 0.000145. This value is derived from all sequences from all sampling time points and thus substantially overestimates the error rate per replication cycle of the HCV RdRp. In addition, it includes base substitutions resulting from the MuLV RT reverse transcription of HCV RNA.
An unexpected finding of the present study was evidence of acute-to-acute HCV transmission in a high proportion of subjects (5 of 17). A limitation of the current study is that virus from donor-recipient HCV transmission partners was not available. However, in cases of acute -to-acute HIV-1 transmission when this has been done, the phylogenetic patterns of closely related sequences are virtually indistinguishable from those found here for HCV (Fig 5; 11, 12, 13 A, 13B). Transmission of HIV-1 during the acute infection period is believed to contribute
substantially to the epidemic in high prevalence and high incidence regions including Africa. Enhanced HCV-1 transmission for acutely infected subjects is believed to result from high virus loads, each of circulating neutralizing antibodies and restricted viral genetic diversity, an interpretation corroborated by acute infection studies of SIV in Indian rhesus macaques. All of these same factors could pertain to acute -to-acute HCV transmission.
The complicated HCV transmission scenario for subject 106889 (Fig 14) illustrates the sensitivity and specificity with which the SGA-direct amplicon sequencing strategy can detect and discriminate T/F viral genomes and their progeny. SGA-direct sequencing includes in vitro recombination artifacts, avoids Taq polymerase-mediated nucleotide substitution errors in finished sequences, precludes founder effects due to disproportionate target amplification, and avoids cloning bias. This allows for a proportional representation of target molecules in finished sequences. It also allows for genetic linkages to be preserved in finished sequences specific mutations, in this case NS3 DAA resistance mutations V36M and R155K to be precisely identified. In subject 106889, each T/F virus contained the V36M/R155K double mutation indicating high level NS3 protease resistance. Subject 106889 as well as the transmitting partner, SGA-direct sequencing is ideally suited to identifying genetically-linked DAA resistant mutants within single or multiple genes (e.g., NS3, NS5A and NS5B) of transmitted viruses and characterizing their persistence or disappearance over time.
Finally, this study demonstrates that that SGA-direct amplicon sequence allows for a precise and unambiguous molecular identification of the complete nucleotide sequences of T/F HCV genomes that are responsible for productive HCV infection of humans. These sequences, when synthesized, molecularly cloned, and expressed eukaryotically, recapitulate all of the features of naturally-occurring T/F HCV viruses. Because their nucleotide sequences match exactly those of viruses responsible for transmission and productive clinical infection of humans, the sequences of T/F viral genomes by definition contain all of the genetic elements necessary for a pathogenic infection of humans. These molecular HCV genomes should prove to be a valuable resource for in vitro and in vivo testing of therapeutic agents, drugs, and vaccines.
Example 3
Molecular Identification, Chemical Synthesis and Molecular Cloning of a Full-Length, Replication-Competent, Infectious HCV Genome
The transmitted/founder virus sequence of the full length genome of subject 10021 was obtained by analysis of five overlapped genome fragments: 5' UTR (nt 1 to 852), 5' half genome (core, El, E2, p7, NS2 andNS3, nt 391 to 5297), 3' half genome (NS4A, NS4B, NS5A andNS5B,
nt 5168 to 9374), 3' U/C tract (nt 9082 to 9582) and 3' x-tail (nt 9563 to 9646). The primers designed for cDNA synthesis and PCR amplification are listed below. The nucleotide numbering system is based on reference sequence H77 (accession number NC 004102). The sequences of all five genome regions were analyzed and assembled. The transmitted/founder sequence was inferred based on phylogenetic inference and mathematical modeling.
For viral RNA extraction, the plasma sample contained approximately 100,000 viral RNA copies and was extracted using the Qiagen BioRobot EZ1 Workstation with EZ1 Vrius Mini Kit v2.0 (Qiagen, Valencia, CA). RNA was eluted and immediately subjected to cDNA synthesis or stored at -80°C. The cDNA synthesis and single genome amplification of '5 and '3 half genome was performed as described supra in Example 2. The positive PCR reactions were subject to direct sequencing.
The cDNA synthesis of 3' U/C tract of subject 10021 was also performed using Superscript III™ Reverse Transcriptase. For a 20 μΐ of reaction, the mixture of 1 μΐ (0.5 mM) of 10 mM deoxynucleoside triphosphate, 0.5 μΐ (0.25 μΜ) of 10 μΜ anti-sense primer 3UTR-R10 and 11.50 μΐ of vRNA was heated at 65 °C for 5 min and then was placed on ice for at least 1 min. Then 4 μΐ of 5X reaction buffer, Ιμΐ of 0.1 M DTT (5 mM final), 1 μΐ of RNAseOUT Recombinant RNASE Inhibitor and 1 μΐ of Superscript III™ Reverse Transcriptase (2001Ι/μ1) were added. The final reaction mix was incubated at 50 °C for 60 min followed by a 5 min heat inactivation at 85°C. PCR amplification was carried out in the presence of 2 μΐ of 10 x Taq High Fidelity Platinum PCR buffer, 0.8 μΐ of 50 mM MgS04 (2 mM), 0.4 μΐ of 10 mM deoxynucleoside triphosphate (0.2 mM), 0.2 μΐ of each primer (0.2 μΜ), and 0.1 μΐ (0.5 Unit) Platinum Taq High Fidelity polymerase in a 20 μΐ reaction. The semi-nested PCR primers used are listed in Table 4. PCR was performed in MicroAmp 96-well reaction plates with the following PCR parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94°C for 15 s, an annealing step of 50 °C for 30 s, an extension step of 68°C for 1.5 min, followed by a final extension of 68 °C for 10 min. The product of the 1st round PCR was subsequently used as a template in the 2nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on precasted 1 % agarose E-gels 96 and 4 % Nusieve GTG Agarose gel (CAMBREX, Cat. No. 50080). The positive PCR reactions were TA cloned into pGEM-T Easy Vector System (Promega, Cat. No. A3610). The inserts were identified and sequenced using Ml 3 forward/reverse primers.
The transmitted/founder virus sequence of 5' UTR of subject 10021 was obtained using FirstChoice RLM-RACE kit (Ambion, Cat. No. AM 1700), with slight modification of the methods recommended by the manufacturer. First, vRNA was treated with Tobacco Acid
Pyrophosphatase (TAP) to remove two 5' P04 from 5' -triphosphate of the full-length vRNA, leaving one 5'-monophosphate. Then a 45 base RNA Adapter oligonucleotide was iigated to the R A population using T4 RNA ligase. During the ligation reaction, the full-length vRNA acquires the adapter sequence as its 5' end, A core gene specific RT-PCR amplified the complete 5' UTR genome. The cDNA synthesis and PCR amplification were performed under same conditions as that described for 3' U/C tract except using 55 °C as the annealing temperature. The positive PCR products were sequenced directly.
To determine the complete x tail sequences, a synthetic RNA oligonucleotide adaptor (Integrated DMA Technologies INC.) wras iigated to the 3' end of the vRNA by using T4 RNA ligase. The adaptor was 5' phosphorylated and 3' Dideoxy-C blocked to prevent intra-molecular ligation. The 20 ul ligation reaction that contained 15 μΐ of vRNA, 2 μΐ of 10X ligation buffer, 1 μΐ of 10 rnM ATP, 0.5 μΐ of RNase Inhibitor (20υ/μ1), 1 μΐ of 10 μΜ oligo adaptor ( 0.5 μΜ) and 1 μΐ of T4 RNA ligase (10 units) was incubated at 37 °C for 60 min. The whole ligation reaction was subject to cDNA synthesis directly using the oligo adaptor specific primer. The cDNA synthesis was carried at 50 °C for 60 min. PCR was performed with the following parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94 °C for 15 s, an annealing step of 55 °C for 30 s, an extension step of 68 °C for 30 sec, followed by a final extension of 68 °C for 10 min. The product of the 1st round PCR was subsequently used as a template in the 2nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on 4 % Nusieve GTG Agarose gel (CAMBREX). The positive PCR reactions were directly TA cloned into pGEM-T Easy Vector System (Promega, Cat. No. A3610). The inserts were identified and sequenced using Ml 3 forward/reverse primers.
The inferred full-length transmitted/founder sequences was chemically synthesized in four subgenomic fragments (Blue Heron Biotechnology). The fragments were overlapped at the unique restriction sites, BsrGI (nt 3640), SnaBI (nt 6644) and Sfil (nt 9404), respectively. The noncutter Notl was attached to the 5' terminal of the first fragment and noncutter Xbal was attached to the 3' terminal of all the fragments during chemical synthesis. The synthesized fragments were propagated, restriction enzyme digested and sequentially cloned into the MCS Notl-Xbal site of pBlueScriptKS(+). The final full-length clone was confirmed by sequence analysis.
Table 1. Diversity and mutation analyses of HCV sequences in acute infection.
Table 1. Continued
a 5h-5' half genome contains core, El, E2, p7, NS2, and NS3; 5ql-5' quarter 1 genome contains core, El, E2, p7 and partial NS2; 5q2-5' quarter 2 genome contains partial NS2 and NS3.
b rate calculations were derived from sequences from each discrete transmitted/founder lineage.
c sequences from 1st sampled time point with >4 sequences per lineage were analyzed. If 5' half genomes were not available, quarter genomes
1 and 2 were analyzed with each result shown.
d sequences from 2nd sampled time point were analyzed due to insufficient nos. of sequences or sequence diversity from 1st sampled time point.
e Insufficient numbers of sequences from each time point to calculate fit to Poisson or star-like phylogeny.
Transmitted/founder lineage identified by this sequence in respective ML tree and Highlighter plot.
g N/A, not applicable due to multiple transmitted/founder virus genomes.
h Averages were calculated from total mutations in all transmitted/founder lineages from all subjects combined. Because of low numbers of sequences and mutations in some lineages, certain values (e.g. dN/dS for subjects 10062 v2 and 10004 v3) vary substantially from the mean.
Table 2. Diversity analysis of 5' half HCV genome sequences from 14 chronically infected
*HCV infection only
Table 3. Jackknife estimation of number of T/F variants in acute infection
a obtained by applying jackknife to the biasd estimator, in this case is the clusters found.
b obtained by bootstrapping the fixed cluster membership.
c power calculation of the prevalence of the unseen variants based on the number of sequences obtained.
Table 4. Estimates of numbers of T/F viruses in acute HCV infection using empirical and model based methods
a manual estimates were based on phylogenies and Highlighter plots of all time points combined.
b power calculation estimating an upper bound on the prevalence of unseen variants given the number of sequences analyzed. This estimate is based on the total number of sequences from all time points.
Table 5. Estimation of numbers of T/F viruses by time point using model-based methods with different cut-offs
Days Avg.
Maximum cut-off1'
Sampling No. of Seq. since cut-off"
Subject Genome
time pt. seq. length negative Total mutation Model- Model- sample cut-off based based
7 5* half 44 4992 6 5 1 1
9055 9 5* half 78 4993 15 5 1 1
11 5* half 35 4993 37 9 1 1
10021 8 5' quarter 1 31 2172 7 3 1 1
Days Avg.
Maximum cut-off1'
Sampling No. of Seq. since cut-off"
Subject Genome
time pt. seq. length negative Total mutation Model- Model- sample cut-off based based
8 5' quarter 2 24 2773 7 3 1 1
10 5* half 30 4960 14 5 1 1
14 5* half 66 4960 35 9 1 1
8 5' quarter 1 40 2318 3 3 1 1
8 5' quarter 2 43 2645 3 3 1 1
10025
9 5* half 38 4964 15 5 1 1
11 5* half 54 4964 31 8 1 1
9 5' quarter 1 46 2212 8 3 1 1
9 5' quarter 2 45 2733 8 3 1 1
10 5' quarter 1 54 2212 13 3 1 1
10051
10 5' quarter 2 33 2733 13 3 1
11 5* half 65 4948 15 5 1 1
14 5* half 60 4948 27 7 1 1
7 5* half 120 4995 7 5 19 9
10003 9 5* half 7 4995 14 5
12 5* half 6 4995 26 7
10 5' quarter 53 2851 4 3 9 3
10016
12 5' quarter 19 2851 28 4 5 3
6 5' quarter 58 2194 9 3 5 2
10020 8 5* half 24 4987 16 5 6 1
13 5* half 40 4987 41 10 1 1
6213 10 5* half 41 4905 29 7 3 3
6222 8 5* half 17 4985 38 10 4 4
4 5* half 5 4987 11 5
10002
7 5* half 26 4987 24 6 13 11
10004c 6 5' quarter 36 2849 12 3 3 3
6 5' quarter 1 52 2367 7 3 3 3
6 5' quarter 2 49 2596 7 3 3 3
10012c 8 5* half 36 4963 14 5 3 3
10 5* half 49 4963 21 5 3 3
13 5* half 44 4964 33 8 3 3
9 5' quarter 1 35 2288 11 3 3 3
9 5' quarter 2 22 2681 11 3 2 1
10 5* half 63 4970 16 5 3 3
10017
12 5* half 48 4969 24 6 2 2
14 5* half 54 4970 31 8 2 2
16 5* half 27 4970 42 11 1 1
6 5' quarter 1 40 2318 16 3 4 4
6 5' quarter 2 70 2700 16 3 3 3
10024
7 5* half 63 4964 20 5 3 3
8 5* half 49 4964 22 6 3 3
8 5' quarter 1 68 2358 6 3 7 7
10029 8 5' quarter 2 53 2599 6 3 7 6
9 5* half 68 4957 13 5 6 7
Days Avg.
Maximum cut-off1'
Sampling No. of Seq. since cut-off"
Subject Genome
time pt. seq. length negative Total mutation Model- Model- sample cut-off based based
11 5* half 75 4957 20 5 7 7
15 5* half 58 4957 34 9 6 5
3 5' quarter 1 24 2425 4 3 2 2
3 5' quarter 2 24 2514 4 3 2 2
10062 4 5* half 54 4938 7 5 3 3
5 5* half 59 4938 11 5 3 3
8 5* half 27 4938 42 11 2 2
106889 5 5* half 87 4984 11 5 28 16 a total mutation cut-off was based on the time of sampling relative to the last negative time point and the average diversity in a cluster and was used to define distinct founders.
b two methods were used to implement the automated clustering algorithm. The average cut-off method is the more conservative estimate of the minimum number of founder strains needed to explain the observed diversity. The maximum cut-off distinguishes more lineages and separates them into clusters from distinct founders.
c the model assumes absence of homoplasy. In these subjects infected by multiple divergent viruses, one mutation at one site could be explained as a homoplasy and hence the
corresponding column was removed from the alignment.
Table 6. Oligonucleotides for ligation, cDNA synthesis and PCR amplification of full genome
Genome Name Position Sense and usage
1.NS4A.R1 (SEQ ID NO: 758) 5451-5474 -, RT and 1st round l .core.Fl (SEQ ID NO: 761) 342-368 +, 1st round
5' half
1.NS3.4A.R2 (SEQ ID NO: 780) 5297-5319 -, 2nd round l .core.F2 (SEQ ID NO: 781) 362-391 +, 2nd round la.NS5Bend.Rla (SEQ ID NO: 782) 9374-9398 -, RT and 1st round
10021.NS3.F1 (SEQ ID NO: 783) 5039-5067 +, 1st round
3' half
la.NS5B.R2 (SEQ ID NO: 784) 9315-9341 -, 2nd round
10021. NS3.F2 (SEQ ID NO: 785) 5146-5168 +, 2nd round
5* RACE Adaptor (SEQ ID NO: 786)
l .Core.Rl (SEQ ID NO: 787) 852-874 -, RT and 1st round
5' UTR 5* RACE Outer (SEQ ID NO: 788) +, 1st round
l .Core.R2 (SEQ ID NO: 789) 822-844 -, 2nd round
5* RACE Inner (SEQ ID NO: 790) +, 2nd round
3UTR-R10 (SEQ ID NO: 791) 9582-9604 -, RT and 1st round
3' U/C laNS5B-F1.2 (SEQ ID NO: 792) 9028-9055 +, 1st round tract 3UTR-R10 (SEQ ID NO: 793) 9582-9604 -, 2nd round
laNS5B-F2 (SEQ ID NO: 794) 9060-9082 +, 2nd round
HL-3* adaptor (SEQ ID NO: 795)
3*-termi-Rl (SEQ ID NO: 796) -, RT and 1st round
3' xtail
10021 3*raceFl (SEQ ID NO: 797) 9537-9559 +, 1st round
3*-termi-R2 (SEQ ID NO: 798) -, 2nd round
Genome Name Position Sense and usage
10021 3*raceF2 (SEQ ID NO: 799) 9544-9563 +, 2nd round
The examples given above are merely illustrative and are not meant to be an exhaustive list of all possible embodiments, applications or modifications of the invention. Thus, various modifications and variations of the described methods and systems of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in connection with specific embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention which are obvious to those skilled in molecular biology, immunology, chemistry, biochemistry or in the relevant fields are intended to be within the scope of the appended claims.
Claims
1. A transmitted full-length HCV genome comprising the polynucleotide of SEQ ID NO: 776.
2. The HCV genome of claim 1, wherein the genome mediates viral transmission.
3. The HCV genome of claim 1, wherein the polynucleotide sequence comprises env and core genes.
4. An immunogenic composition comprising transmitted full-length HCV genome or portions thereof.
5. The immunogenic composition of claim 4, wherein the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
6. A method of administering a vaccine, the method comprising treating a patient with transmitted full-length HCV genome or portions thereof, wherein a patient immune response is induced.
7. The method of claim 6, wherein the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
8. A method for identifying full-length transmitted HCV genome identified by the method comprising:
(a) collecting a patient sample;
(b) isolating viral RNA from said sample;
(c) sequencing said viral RNA, wherein viral RNA sequencing includes HCV genomes of circulating virus;
(d) performing sequence alignment of selected HCV genome regions;
(e) analyzing phylogenetically selected sequence alignments; and
(f) identifying full-length HCV genomes of transmitted virus.
9. The HCV genome as identified by the method of claim 8, wherein the polynucleotide sequence comprises SEQ ID NO. 776.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US201161630161P | 2011-12-05 | 2011-12-05 | |
| PCT/US2012/067731 WO2013085889A1 (en) | 2011-12-05 | 2012-12-04 | Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP2788490A1 true EP2788490A1 (en) | 2014-10-15 |
Family
ID=48574801
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP12854779.1A Withdrawn EP2788490A1 (en) | 2011-12-05 | 2012-12-04 | Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP2788490A1 (en) |
| WO (1) | WO2013085889A1 (en) |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| ATE526423T1 (en) * | 2000-04-18 | 2011-10-15 | Virco Bvba | METHOD FOR MEASURING DRUG RESISTANCE TO HCV |
| US7659103B2 (en) * | 2004-02-20 | 2010-02-09 | Tokyo Metropolitan Organization For Medical Research | Nucleic acid construct containing fulllength genome of human hepatitis C virus, recombinant fulllength virus genome-replicating cells having the nucleic acid construct transferred thereinto and method of producing hepatitis C virus Particle |
| US20140023683A1 (en) * | 2010-09-08 | 2014-01-23 | The Uab Research Foundation | Identification of transmitted hepatitis c virus (hcv) genomes by single genome amplification |
-
2012
- 2012-12-04 EP EP12854779.1A patent/EP2788490A1/en not_active Withdrawn
- 2012-12-04 WO PCT/US2012/067731 patent/WO2013085889A1/en not_active Ceased
Non-Patent Citations (1)
| Title |
|---|
| See references of WO2013085889A1 * |
Also Published As
| Publication number | Publication date |
|---|---|
| WO2013085889A8 (en) | 2013-11-28 |
| WO2013085889A1 (en) | 2013-06-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Cecchinato et al. | Avian metapneumovirus (AMPV) attachment protein involvement in probable virus evolution concurrent with mass live vaccine introduction | |
| US7320857B2 (en) | Characterization of the earliest stages of the severe acute respiratory syndrome (SARS) virus and uses thereof | |
| WO2006086188A2 (en) | Use of consensus sequence as vaccine antigen to enhance recognition of virulent viral variants | |
| Costa-Mattioli et al. | Evidence of recombination in natural populations of hepatitis A virus | |
| Wu et al. | Recombination of hepatitis D virus RNA sequences and its implications. | |
| Leary et al. | Three adjacent nucleotide changes spanning two residues in SARS-CoV-2 nucleoprotein: possible homologous recombination from the transcription-regulating sequence | |
| AU2003218111A1 (en) | Methods and compositions for identifying and characterizing hepatitis c | |
| Han et al. | Comparison of multiplex restriction fragment mass polymorphism and sequencing analyses for detecting entecavir resistance in chronic hepatitis B | |
| Franco et al. | Complete nucleotide sequence of genotype 4 hepatitis C viruses isolated from patients co-infected with human immunodeficiency virus type 1 | |
| US20140023683A1 (en) | Identification of transmitted hepatitis c virus (hcv) genomes by single genome amplification | |
| Peres-da-Silva et al. | Genetic diversity of NS3 protease from Brazilian HCV isolates and possible implications for therapy with direct-acting antiviral drugs | |
| Li et al. | Molecular epidemiology of hepatitis C genotype 6a from patients with chronic hepatitis C from Hong Kong | |
| Bracho et al. | Complete genome of a European hepatitis C virus subtype 1g isolate: phylogenetic and genetic analyses | |
| WO2020184730A1 (en) | Dengue virus vaccine | |
| EP2788490A1 (en) | Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification | |
| Harasawa et al. | Evidence for Pestivirus Infection in Free‐Living Japanese Serows, Capricornis crispus | |
| González-Horta et al. | Analysis of hepatitis C virus core encoding sequences in chronically infected patients reveals mutability, predominance, genetic history and potential impact on therapy of Cuban genotype 1b isolates | |
| KR102297300B1 (en) | A live virus banked from an attenuated dengue virus strain, and a dengue vaccine using them as an antigen | |
| JP2017500886A (en) | HCV genotyping algorithm | |
| Dinu et al. | Screening of protease inhibitors resistance mutations in hepatitis c virus isolates infecting romanian patients unexposed to triple therapy | |
| Mello et al. | Conservation of hepatitis C virus nonstructural protein 3 amino acid sequence in viral isolates during liver transplantation | |
| Larsson | Hepatitis B virus replication and integration | |
| US10815523B2 (en) | Indexing based deep DNA sequencing to identify rare sequences | |
| Ndjomou et al. | Functional domains of the human immunodeficiency virus type 1 Nef protein are conserved among different clades in Cameroon | |
| US20140271726A1 (en) | Compositions and methods for predicting hcv susceptibility to antiviral agents |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| 17P | Request for examination filed |
Effective date: 20140707 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR |
|
| DAX | Request for extension of the european patent (deleted) | ||
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN |
|
| 18D | Application deemed to be withdrawn |
Effective date: 20150701 |