EP2788490A1 - Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification - Google Patents

Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification

Info

Publication number
EP2788490A1
EP2788490A1 EP12854779.1A EP12854779A EP2788490A1 EP 2788490 A1 EP2788490 A1 EP 2788490A1 EP 12854779 A EP12854779 A EP 12854779A EP 2788490 A1 EP2788490 A1 EP 2788490A1
Authority
EP
European Patent Office
Prior art keywords
hcv
sequences
virus
seq
sequence
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Withdrawn
Application number
EP12854779.1A
Other languages
German (de)
French (fr)
Inventor
George M. Shaw
Hui Li
Beatrice H. Hahn
Barton F. Haynes
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Duke University
Original Assignee
Duke University
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Duke University filed Critical Duke University
Publication of EP2788490A1 publication Critical patent/EP2788490A1/en
Withdrawn legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N7/00Viruses; Bacteriophages; Compositions thereof; Preparation or purification thereof
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12QMEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
    • C12Q1/00Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
    • C12Q1/70Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving virus or bacteriophage
    • C12Q1/701Specific hybridization probes
    • C12Q1/706Specific hybridization probes for hepatitis
    • C12Q1/707Specific hybridization probes for hepatitis non-A, non-B Hepatitis, excluding hepatitis D
    • AHUMAN NECESSITIES
    • A61MEDICAL OR VETERINARY SCIENCE; HYGIENE
    • A61KPREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
    • A61K39/00Medicinal preparations containing antigens or antibodies
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2770/00MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA ssRNA viruses positive-sense
    • C12N2770/00011Details
    • C12N2770/24011Flaviviridae
    • C12N2770/24211Hepacivirus, e.g. hepatitis C virus, hepatitis G virus
    • C12N2770/24221Viruses as such, e.g. new isolates, mutants or their genomic sequences

Definitions

  • the invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection.
  • the invention provides methods for identifying nucleotide sequences corresponding to specific transmitted HCV genomes, including structural gene nucleotide sequences of global HCV genotypes, subtypes, and drug resistant variants that actively mediate viral infection.
  • the invention further provides methods of administering a vaccine comprising transmitted full-length hepatitis C virus (HCV) genomes or portions thereof.
  • Hepatitis C Virus is a positive strand, non-segmented, enveloped RNA virus of approximately 9.6 kb in length.
  • the virus is classified in the genus Hepacivirus within the larger family of Flavivirus, which includes the human pathogens West Nile virus, yellow fever virus and dengue fever virus among others.
  • a common feature among the Flaviviridae is their dependence on a virally-encoded RNA-dependent RNA polymerase (RdRp) for replication.
  • RdRp is error-prone, and HCV is notable for its quasispecies complexity and broad genotypic diversity. Globally, there are seven major genotypes of HCV that differ by approximately 30- 35% in nucleotide sequence.
  • HCV The extraordinary diversity of HCV has important implications for clinical management and for basic and translational research aimed at elucidating viral natural history, pathogenesis and susceptibility to novel therapeutics and vaccines.
  • DAA direct acting antiviral
  • the development of new drugs and drug combinations that effectively suppress HCV replication and prevent the emergence of DAA resistance is a major challenge given the high rates of virus replication and variation.
  • HCV diversity poses similar challenges to the development of effective vaccines and to the elucidation of virus biology, immunopathogenesis, and gene structure-function relationships.
  • Acute HCV infection defined as the period between virus transmission and antibody seroconversion which occurs about 6 - 10 weeks later, sets in motion viral-host interactions that largely dictate the natural history of the infection.
  • viral genotype and host immunogenetic factors most importantly IL28B alleles, a proportion of newly infected individuals spontaneously control or eliminate the virus.
  • IL28B alleles a proportion of newly infected individuals spontaneously control or eliminate the virus.
  • a greater number of patients can be cured if the infection is treated with interferon and ribavirin alone or in combination with DAA drugs. Mechanistically, how this occurs is unknown, and how the emergence of DAA drug resistance in communities of chronically infected subjects or in acutely infected subjects who initiate early therapy will affect treatment responses is uncertain, but again there are parallels with HIV-1.
  • the invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection.
  • HCV hepatitis C virus
  • One application of this invention is to use these transmitted full-length genomes for development of more effective vaccines and treatments.
  • the transmitted full-length HCV genome comprises the polynucleotide of SEQ ID NO: 776.
  • the HCV genome mediates viral transmission.
  • the polypeptide sequence comprises env and core genes. Nucleotide sequence of transmitted HCV genomes is set forth in Table 5.
  • the invention provides methods for identifying full-length transmitted HCV genomes, the methods comprising collecting a patient sample, isolating and preparing viral RNA for sequencing, sequencing viral RNA that includes HCV genomes of circulating virus, performing sequence alignment of selected HCV genome regions, analyzing phylogenetically selected sequence alignments; and identifying full-length HCV genomes of transmitted virus.
  • the HCV genome polypeptide sequence comprises SEQ ID NO. 776.
  • the invention provides an immunogenic composition comprising transmitted full-length HCV genome or portions thereof.
  • the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
  • the invention provides methods for administering vaccines comprising transmitted full-length HCV genomes or portions thereof, wherein an immune response is induced in the patient following vaccination.
  • the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
  • Figure 1 illustrates HCV R A kinetics in 16 acute infection subjects.
  • the shaded area represents that the HCV RNA is below the linear range of quantitation (43-69,000,000 IU/ml). Plus signs indicate the samples are Anti-HCV antibody positive.
  • the circles represent the time points that were sequence analyzed from each subject.
  • Figure 2 is a Maximum-likelihood tree (ML) of 5' quarter genome ⁇ core, El and E2) sequences from 17 acutely and 14 chronically infected subjects. The sequences from acutely infected subjects are shown in red and the sequences from chronically infected subjects are shown in blue. The reference sequences of genotype 1 to 7 are in gray. Bootstrap values (>70%) are indicated for intra-subject clusters. The horizontal scale bar represents 3% genetic distance.
  • ML Maximum-likelihood tree
  • Figure 3 illustrates 5' half-genome sequence diversity in two subjects, one with chronic infection (WIMI4025 (Fig 3 A)) and one with acute infection (10051 Fig 3(B)). Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. Bootstrap values (>70%) are indicated for intra-linerage clusters. The horizontal scale bars represent 0.5% (24.5 nt) and 0.02% (1 nt) genetic distance for subject WIMI and 10051, respectively.
  • Figure 4 illustrates 5' quarter-genome ⁇ Core, El and E2) sequence diversity in four acutely infected subjects. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. Sequences from multiple time points are color coded with orange, green, blue and black in chronological order. Sequences from subject 10021 (Fig 4A) and 10025 (Fig 4B) each showed productive infection by a single virus. Sequences from subject 10012 (Fig 4C) and 10062 (Fig 4D) showed productive infection by at least three viruses, repectively. The horizotal scale bars represent 0.04% (1 nt) genetic distance for Figs. 4A, 4B and 4D and 0.4% (10 nt) for 4C. Bootstrap values (>70%) are indicated for within patient intra-lineage clusters in the ML tree.
  • Figure 5 illustrates 5' quarter-genome ⁇ Core, El and E2) sequence divsersity in an acutely infected subject. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the ML tree and highlighter plot of 5' quarter genome ⁇ Core, El and E2) sequences generated from multiple time points showed productive infection by at least 9 transmitted/founder variants.
  • the close genetic diversity between variant 1 and 2, variant 6 and 7 is indicative of transmission from an acutely infected donor.
  • the horizotal scale bars represent 0.2% (5 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra- lineage clusters.
  • Figure 6 illustrates HCV diversity in acutely infected subject 10024. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the ML tree and highlighter plot of 5' quarter genome sequences derived from multiple time points showed productive infection by at least 7 transmitted/founder variants.
  • the horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
  • Figure 7 illustrates HCV diversity in acutely infected subject 6123. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 3 transmitted/founder variants.
  • the horizotal scale bars represent 0.2% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
  • Figure 8 illustrates HCV diversity in acutely infected subject 6222. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 4 transmitted/founder variants.
  • the horizotal scale bars represent 0.02% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
  • Figure 9 illustrates HCV diversity in acutely infected subject 10004. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the ML tree and highlighter plot of the 5' quarter genome sequences showed productive infection by at least 3 transmitted/founder variants.
  • the horizotal scale bars represent 0.36% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
  • Figure 10 illustrates HCV diversity in acutely infected subject 10002. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 13 transmitted/founder variants.
  • the horizotal scale bars represent 0.2% (10 nt) genetic distance.
  • Bootstrap values (>70%) are indicated for patient intra- lineage clusters.
  • Figure 11 illustrates HCV diversity in acutely infected subject 10017. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the ML tree and highlighter plot of 5' quarter genome sequences derived from multiple time points showed productive infection by at least 6 transmitted/founder variants.
  • the horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
  • Figure 12 illustrates HCV diversity in acutely infected subject 9055. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' half genome sequences derived from multiple time points showed productive infection by a single transmitted/founder variants. The horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
  • Figure 13 illustrates 5' quarter-genome (Core, El and E2) sequence divsersity in acute- to-acute transmission.
  • 5' quarter genome (Core, El and E2) sequences from subject 10020 (Fig 13 A) and 10016 (Fig 13B), and 5' half genome (core, El, E2, p7, NS2 and NS3A) sequences from subject 10003 (Fig 13C) are depicted by ML trees and by highlighter plots. Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the first sequence in red is the consensus sequence derived from the entire sequence dataset.
  • the horizotal scale bars represent 0.04%> (1 nt), 0.036%) (1 nt) and 0.02%> (1 nt) genetic distance for Figs. 13 A, 13B and 13C, repectively. Bootstrap values (>70%) are indicated.
  • Figure 14 illustrates HCV RNA kinectics and diversity in subject 106889.
  • the inset shows the HCV RNA kinetics.
  • the time point that was sequence analyzed is circled.
  • Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively.
  • the 5' half genome (core, El, E2, p7, NS2 and NS3A) sequences formed 10 distinct lineages (labeled with letter L) with high statistical support in the ML tree. Each cluster is color coded.
  • the T/F variants was identified within each lineage. A total of 33 T/F variants were found that were responsible for productive infection.
  • the horizotal scale bars represent 0.04%) (1 nt) genetic distance. Bootstrap values (>70%>) are indicated.
  • Figure 15 illustrates strand transfers at stem loop structures of HCV RNA.
  • Figures 17A- E illustrate four examples of stem loop structures in double-stranded HCV RNA leading to template switching by the HCV RNA polymerase.
  • Fig 17A SEQ ID NO: 800 First 5' to 3' horizontal sequence; SEQ ID NO: 801 Second 3' to 5' horizontal sequence; SEQ ID NO: 802 Third 5' to 3' horizontal sequence; SEQ ID NO: 800 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 802 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 801 Sixth 3' to 5' haripin sequence.
  • Fig 17B SEQ ID NO: 803 First 5' to 3' horizontal sequence; SEQ ID NO: 804 Second 3' to 5' horizontal sequence; SEQ ID NO: 805 Third 5' to 3' horizontal sequence; SEQ ID NO: 803 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 805 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 804 Sixth 3' to 5' hairpin sequence.
  • Fig 17C SEQ ID NO: 806 First 5' to 3' horizontal sequence; SEQ ID NO: 807 Second 3' to 5' horizontal sequence; SEQ ID NO: 808 Third 5' to 3' horizontal sequence; SEQ ID NO: 806 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 808 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 807 Sixth 3' to 5' hairpin sequence.
  • Fig 17D SEQ ID NO: 809 First 5' to 3' horizontal sequence; SEQ ID NO: 810 Second 3' to 5' horizontal sequence; SEQ ID NO: 811 Third 5' to 3' horizontal sequence; SEQ ID NO: 809 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 811 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 810 Sixth 3' to 5' hairpin sequence.
  • Fig 17E SEQ ID NO: 812 First 5' to 3' horizontal sequence; SEQ ID NO: 813 Second 3' to 5' horizontal sequence; SEQ ID NO: 814 Third 5' to 3' horizontal sequence; SEQ ID NO: 812 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 814 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 813 Sixth 3' to 5' hairpin sequence.
  • Figure 16 illsutrates the strategy for identification of full-length HCV genome by single gneome amplification (SGA).
  • the complete HCV genome of subject 10021 was determined by amplifying five overlapping genome fragments: fragment I (nt 1 - 852) contains the complete 5' UTR and partial Core; fragment II (nt 391 - 5297) contains Core, El, E2, NS2 and NS3; fragment III (nt 5168 - 9374) contains NS4A, NS4B, NS5A and NS5B; fragment IV (nt 9082 - 9582) contains partial NS5B, variable region and the complete poly U/UC tract; and fragment V (nt 9563 - 9646) contains the complete x-tail.
  • the nucleotide positions are based reference sequence H77.
  • Figure 17 illustrates the strategy for generating a full-length T/F HCV molcular clone from subject 10021.
  • Figure 18 illustrates HCV diversity in subject 10003.
  • ML tree and Highlighter plot of 5' halfgenome sequences reveal many sets of closely related sequences distinguished by unique shared mutations.
  • Figure 19 illustrates HCV diversity analysis in subject 10016 suggests acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 5' quarter 1 genome sequences. Visualization of 15 potential T/F viral lineages distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveal 10 and 4 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
  • Figure 20 illustrates maximum diversity of discrete HCV sequence lineages from acute infection subjects versus maximum sequence diversity in chronic subjects.
  • Primary data are derived from Tables 1 and 2.
  • Mean ( ⁇ 95% CI) values are represented by horizontal lines. Differences between the two groups were highly significant (p ⁇ 0.0001; unpaired T-test with Welch's correction), reflecting the recent and remote diversification histories of acute and chronic sequences, respectively.
  • Figure 21 shows HCV diversity in acute subject 10051. 5' quarter 1 genomesequences are color coded in orange, green, blue and black in chronological order to reflect sampling time points in Figure 1 and are represented in a ML tree and Highlighter plot. Sequences show evidence of productive clinical infection by a single virus. The horizontal scale bar indicates genetic distance.
  • Figure 22 illustrates nonsynonymous and synonymous mutations in HCV sequences from acute subjects 10003, 10020 and 10016. Highlighter plotsof 5' half or quarter 1 genomesequences are color coded to denote nonynonymous (red) and synonymous (green) mutations for subjects 10003 (panel A), 10020 (panel B) and 10016 (panel C).
  • Figure 23 illustrates amino acid alignment of the HCV Env coding region of acute subject 10003.
  • the H77 reference sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
  • Figure 24 illustrates amino acid alignment of the HCV Env coding region of acute subject 10020.
  • the H77 reference sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
  • Figure 25 illustrates amino acid alignment of the HCV Env coding region of acute subject 10016.
  • the H77 sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
  • Figure 26 shows HCV diversity analysis in subject 10020 to suggest acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 5' quarter 1 genome sequences. Visualization of 10 potential T/F viral sequences distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveals 5 and 3 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
  • Figure 27 shows HCV diversity analysis in subject 10003 to suggest acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 59 half genome sequences. Visualization of 37 potential T/F viral sequences distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveals 15 and 8 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
  • Figure 28 illustrates nonsynonymous and synonymous mutations in HCV sequences from acute subject 9055. A Highlighter plot (panel A) of 5' half genomesequences is color coded to denote nonynonymous (red) and synonymous (green) mutations.
  • the boxed area reveals a temporal expansion of sequences with concentrated amino acid subtitutions in NS3.
  • amino acid selection is evident in a previously identified CTL epitope highlighted in red.
  • the top-most sequence represents the genotype 3a consensus.
  • the invention provides full-length transmitted HCV genomes that mediate viral infection and transmission. Prior to the present invention the identification of these genomes was unavailable due to the following: (i) a significant time period from weeks to months between the moment of transmission and the first appearance of HCV in the blood (Bowen. D.G. and Walker, CM., 2005, Nature 436: 946-952; Moradpour, D. et al, 2007, Nature Rev Micro 5:453- 463); (ii) HCV is genetically highly variable in its nucleotide sequence due to its error-prone RNA-dependent RNA polymerase, and as a result, it exists in individuals as a complex mixture of sequences commonly referred to as a 'quasispecies' (Moradpour, D.
  • polynucleotide is intended to encompass a singular nucleic acid or nucleic acid fragment as well as plural nucleic acids or nucleic acid fragments, and refers to an isolated molecule or construct, e.g., a virus genome (e.g., vRNA), messenger RNA (mRNA), plasmid DNA (pDNA), or derivatives of pDNA (e.g., minicircles as described in (Darquet, A-M et al., 1997, Gene Therapy, 4: 1341-1349) comprising a polynucleotide.
  • virus genome e.g., vRNA
  • mRNA messenger RNA
  • pDNA plasmid DNA
  • derivatives of pDNA e.g., minicircles as described in (Darquet, A-M et al., 1997, Gene Therapy, 4: 1341-1349
  • a polynucleotide may comprise a conventional phosphodiester bond or a non-conventional bond (e.g., an amide bond, such as found in peptide nucleic acids (PNA)).
  • PNA peptide nucleic acids
  • nucleic acid or nucleic acid fragment refer to any one or more nucleic acid segments, e.g., DNA or R A fragments, present in a polynucleotide or construct.
  • a nucleic acid or fragment thereof may be provided in linear (e.g., mR A) or circular (e.g., plasmid) form as well as double-stranded or single-stranded forms.
  • isolated nucleic acid or polynucleotide is intended a nucleic acid molecule, DNA or RNA, which has been removed from its native environment.
  • a recombinant polynucleotide contained in a vector is considered isolated or “cloned” for the purposes of the present invention.
  • Further examples of an isolated polynucleotide include recombinant polynucleotides maintained in heterologous host cells or purified (partially or substantially) polynucleotides in solution.
  • Isolated RNA molecules include in vivo or in vitro RNA transcripts of the polynucleotides of the present invention. Isolated polynucleotides or nucleic acids according to the present invention further include such molecules produced synthetically.
  • fragment when referring to HCV polypeptides of the present invention include any polypeptides which retain at least some of the immunogenicity or antigenicity of the corresponding native polypeptide. Fragments of HCV polypeptides of the present invention include proteolytic fragments, deletion fragments and in particular, fragments of HCV polypeptides which exhibit increased secretion from the cell or higher immunogenicity or reduced pathogenicity when delivered to an animal. Polypeptide fragments further include any portion of the polypeptide which comprises an antigenic or immunogenic epitope of the native polypeptide, including linear as well as three-dimensional epitopes.
  • Variants of HCV polypeptides of the present invention include fragments, and also polypeptides with altered amino acid sequences due to amino acid substitutions, deletions, or insertions. Variants may occur naturally, such as an allelic variant.
  • allelic variant is intended alternate forms of a gene occupying a given locus on a chromosome or genome of an organism or virus. Genes II, Lewin, B., ed., John Wiley & Sons, New York (1985), which is incorporated herein by reference.
  • variations in a given gene product is a "variant".
  • Naturally or non-naturally occurring variations such as amino acid deletions, insertions or substitutions may occur.
  • Non-naturally occurring variants may be produced using art-known mutagenesis techniques.
  • Variant polypeptides may comprise conservative or non-conservative amino acid substitutions, deletions or additions.
  • Derivatives of HCV polypeptides of the present invention are polypeptides which have been altered so as to exhibit additional features not found on the native polypeptide. Examples include fusion proteins.
  • An analog is another form of an HCV polypeptide of the present invention.
  • An example is a proprotein which can be activated by cleavage of the proprotein to produce an active mature polypeptide.
  • the polynucleotide, nucleic acid, or nucleic acid fragment is DNA.
  • a polynucleotide comprising a nucleic acid which encodes a polypeptide normally also comprises a promoter and/or other transcription or translation control elements operably associated with the polypeptide-encoding nucleic acid fragment.
  • An operable association is when a nucleic acid fragment encoding a gene product, e.g., a polypeptide, is associated with one or more regulatory sequences in such a way as to place expression of the gene product under the influence or control of the regulatory sequence(s).
  • Two DNA fragments are "operably associated” if induction of promoter function results in the transcription of mRNA encoding the desired gene product and if the nature of the linkage between the two DNA fragments does not (1) result in the introduction of a frame-shift mutation, (2) interfere with the ability of the expression regulatory sequences to direct the expression of the gene product, or (3) interfere with the ability of the DNA template to be transcribed.
  • a promoter region would be operably associated with a nucleic acid fragment encoding a polypeptide if the promoter was capable of effecting transcription of that nucleic acid fragment.
  • the promoter may be a cell-specific promoter that directs substantial transcription of the DNA only in predetermined cells.
  • Other transcription control elements besides a promoter, for example enhancers, operators, repressors, and transcription termination signals, can be operably associated with the polynucleotide to direct cell-specific transcription. Suitable promoters and other transcription control regions are disclosed herein.
  • transcription control regions are known to those skilled in the art. These include, without limitation, transcription control regions which function in vertebrate cells, such as, but not limited to, promoter and enhancer segments from cytomegaloviruses (the immediate early promoter, in conjunction with intron-A), simian virus 40 (the early promoter), and retroviruses (such as Rous sarcoma virus).
  • Other transcription control regions include those derived from vertebrate genes such as actin, heat shock protein, bovine growth hormone and rabbit ⁇ -globin, as well as other sequences capable of controlling gene expression in eukaryotic cells. Additional suitable transcription control regions include tissue-specific promoters and enhancers as well as lymphokine-inducible promoters (e.g., promoters inducible by interferons or interleukins).
  • a DNA polynucleotide of the present invention may be a circular or linearized plasmid or vector, or other linear DNA which may also be non-infectious and nonintegrating (i.e., does not integrate into the genome of vertebrate cells).
  • a linearized plasmid is a plasmid that was previously circular but has been linearized, for example, by digestion with a restriction endonuclease.
  • RNA for example, in the form of messenger RNA (mRNA).
  • mRNA messenger RNA
  • Polynucleotides, nucleic acids, and nucleic acid fragments of the present invention may be associated with additional nucleic acids which encode secretory or signal peptides, which direct the secretion of a polypeptide encoded by a nucleic acid fragment or polynucleotide of the present invention.
  • proteins secreted by mammalian cells have a signal peptide or secretory leader sequence which is cleaved from the mature protein once export of the growing protein chain across the rough endoplasmic reticulum has been initiated.
  • polypeptides secreted by vertebrate cells generally have a signal peptide fused to the N-terminus of the polypeptide, which is cleaved from the complete or "full length" polypeptide to produce a secreted, or "mature” form of the polypeptide.
  • the native leader sequence is used, or a functional derivative of that sequence that retains the ability to direct the secretion of the polypeptide that is operably associated with it.
  • a heterologous mammalian leader sequence, or a functional derivative thereof may be used.
  • the wild-type leader sequence may be substituted with the leader sequence of human tissue plasminogen activator (TP A) or mouse beta-glucuronidase.
  • RNA messenger-RNA
  • polypeptide is intended to encompass a singular "polypeptide” as well as plural “polypeptides,” and comprises any chain or chains of two or more amino acids.
  • polypeptide terms including, but not limited to “peptide,” “dipeptide,” “tripeptide,” “protein,” “amino acid chain,” or any other term used to refer to a chain or chains of two or more amino acids, are included in the definition of a “polypeptide,” and the term “polypeptide” can be used instead of, or interchangeably with any of these terms.
  • the term further includes polypeptides which have undergone post-translational modifications, for example, glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting/blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids.
  • polypeptides of the present invention are fragments, derivatives, analogs, or variants of the foregoing polypeptides, and any combination thereof.
  • Polypeptides, and fragments, derivatives, analogs, or variants thereof of the present invention can be antigenic and immunogenic polypeptides related to HCV polypeptides, which are used to prevent or treat, i.e., cure, ameliorate, lessen the severity of, or prevent or reduce contagion of infectious disease caused by the HCV.
  • an "antigenic polypeptide” or an “immunogenic polypeptide” is a polypeptide which, when introduced into a vertebrate, reacts with the vertebrate's immune system molecules, i.e., is antigenic, and/or induces an immune response in the vertebrate, i.e., is immunogenic. It is quite likely that an immunogenic polypeptide will also be antigenic, but an antigenic polypeptide, because of its size or conformation, may not necessarily be immunogenic.
  • Isolated antigenic and immunogenic polypeptides of the present invention in addition to those encoded by polynucleotides of the invention, may be provided as a recombinant protein, a purified subunit, a viral vector expressing the protein, or may be provided in the form of an inactivated HCV vaccine, e.g., a live-attenuated virus vaccine, a heat-killed virus vaccine, etc.
  • an "isolated" HCV polypeptide or a fragment, variant, or derivative thereof is intended an HCV polypeptide or protein that is not in its natural form. No particular level of purification is required.
  • an isolated HCV polypeptide can be removed from its native or natural environment.
  • HCV polypeptides and proteins expressed in host cells are considered isolated for purposed of the invention, as are native or recombinant HCV polypeptides which have been separated, fractionated, or partially or substantially purified by any suitable technique, including the separation of HCV virions from culture cells in which they have been propagated.
  • an isolated HCV polypeptide or protein can be provided as a live or inactivated viral vector expressing an isolated HCV polypeptide and can include those found in inactivated HCV vaccine compositions.
  • isolated HCV polypeptides and proteins can be provided as, for example, recombinant HCV polypeptides, a purified subunit of HCV, a viral vector expressing an isolated HCV polypeptide, or in the form of an inactivated or attenuated HCV vaccine.
  • epitopes refers to portions of a polypeptide having antigenic or immunogenic activity in a vertebrate, for example a human.
  • An "immunogenic epitope,” as used herein, is defined as a portion of a protein that elicits an immune response in an animal, as determined by any method known in the art.
  • antigenic epitope is defined as a portion of a protein to which an antibody or T-cell receptor can immunospecifically bind as determined by any method well known in the art. Immunospecific binding excludes nonspecific binding but does not exclude cross-reactivity with other antigens. Where all immunogenic epitopes are antigenic, antigenic epitopes need not be immunogenic.
  • peptides or polypeptides bearing an antigenic epitope e.g., that contain a region of a protein molecule to which an antibody or T cell receptor can bind
  • relatively short synthetic peptides that mimic part of a protein sequence are routinely capable of eliciting an antiserum that reacts with the partially mimicked protein. See, e.g., Sutcliffe, J. G., et al., 1983, Science 219:660-666.
  • the present invention also provides methods for identifying transmitted full-length HCV genomes that mediate viral infection and transmission.
  • the methods comprise: (a) collecting a patient sample; (b) isolating viral RNA from said sample; (c) sequencing said viral RNA, wherein viral RNA sequencing includes HCV genomes of circulating virus; (d) performing sequence alignment of selected HCV genome regions; (e) analyzing phylogenetically selected sequence alignments; and (f) identifying HCV genomes of transmitted virus.
  • a "patient” or “subject” to be utilized by the disclosed methods can mean a human, chimpanzee or non-human primate. In certain embodiments the patient is a non-human mammal.
  • patient sample includes but is not limited to a blood, serum, plasma, or urine sample obtained from a patient.
  • the patient sample is plasma.
  • analyzing phylogentically represents the analysis of clinical viral isolates, regardless of the particular methodology employed. Comparative analysis of the genetic relatedness of any a collection of circulating viral isolates is used to select for nucleotide sequence of transmitted HCV genomes. This methodology is used to select for the actual HCV genomes present at the time of patient infection/viral transmission. This can be accomplished by any number of methods, including but not limited to: i) a novel mathematical model of random virus evolution disclosed herein; ii) star phylogeny; iii) Baysian analysis; or any other method of phylogenetic analysis known to one of skill in the art.
  • performing sequence alignment is meant to include aligning genetic sequences by any of a number of different procedures that produce a sufficient match between the corresponding residue in the sequences. Typically, Smith- Waterman or Needleman- Wunsch algorithms are used. However, other procedures such as BLAST, FASTA, PSI-BLAST can be used.
  • isolated viral RNA includes those methods well known in the art for extraction of viral RNA from a sample, cDNA synthesis from viral RNA template, optionally cloning of cDNA fragments, and/or amplification of polynucleotide sequence. Such methods are described in more detail, but such experimental procedures are provided in Sambrook and Russell, 2001, Molecular Cloning: A Laboratory Manual, Third Ed., Cold Spring Harbor Laboratory Press, Woodbury, N.Y.
  • nucleotide sequence amplification such as polymerase chain reaction (PCR) and modifications thereof (including for example reverse transcription (RT)-PCR, and stem-loop PCR), as well as reverse transcription and in vitro transcription.
  • PCR polymerase chain reaction
  • RT reverse transcription
  • these methods utilize one or a pair of oligonucleotide primers having sequence complimentary to sequences 5' and 3' to the sequence of interest, and in the use of these primers they are hybridized to a nucleotide sequence and extended during the practice of PCR amplification using DNA polymerase (preferably using a thermal-stable polymerase such as Taq polymerase).
  • RT-PCR may be performed on miRNA or mRNA with a specific 5' primer or random primers and appropriate reverse transcription enzymes such as avian (AMV-RT) or murine (MMLV-RT) reverse transcriptase enzymes.
  • AMV-RT avian
  • MMLV-RT murine reverse transcriptase enzymes.
  • over time represents a period of time between the collection of patient samples. For example, a patient sample is collected at time point A and then subsequently as a later date at time point B. The period between samplings is over time. In certain embodiments, samples are taken from the same patient. In alternative embodiments, samples may be taken from different patients. Performing the method of the invention at different time points followed by the differential analysis of HCV genomes at the different time points provides a means for assessing the evolution of HCV genomes both inter- and intra- patient. The identification of highly variable genomic regions is useful for the generation of effective therapeutics and vaccines to HCV.
  • transmitted refers to viral genotype and phenotype at the time of viral infection.
  • transmitted virus as used here is in reference to actual HCV virus that mediates patient infection.
  • transmitted viral sequence as used herein is meant to include nucleotide or amino acid sequence of HCV genomes of transmitted virus, an in a particular embodiment sequence corresponding to acute infection stage HCV.
  • circulating refers to HCV virus collected from patient post-infection. In general, patient samples comprise circulating virus because HCV infection has occurred at a prior time point. The half-life of plasma virus is less than 1 day, thus the collection of transmitted virus from a patient sample would be a rare event.
  • selected as used in the phrases “selected sequence alignments” or “selected HCV genome regions” is meant to include the identification and utilization of particular subsets of nucleotide sequence and/or genome regions for subsequent analysis. In certain embodiments, full-length HCV genomes are identified, however discrete portions of the genome are utilized or “selected” for further analysis.
  • a or “an” entity refers to one or more of that entity; for example, “a polynucleotide,” is understood to represent one or more polynucleotides.
  • the terms “a” (or “an”), “one or more,” and “at least one” can be used interchangeably herein.
  • the present invention also provides immunogenic compositions and methods for delivery of transmitted full-length HCV polynucleotide or polypeptide sequences to a vertebrate with optimal expression and safety conferred.
  • the transmitted full-length HCV genome comprises the polynucleotide of SEQ ID NO. 776.
  • These immunogenic compositions may be prepared and administered in such a manner that the encoded gene products are optimally expressed in the vertebrate of interest. As a result, these compositions and methods are useful in stimulating an immune response against HCV infection.
  • expression systems and delivery systems are also included in the invention.
  • the identified polynucleotides or polypeptides encoded by the polynucleotides of the invention may be in any form, and polypeptides are generated using techniques well known in the art. Examples include isolated HCV proteins produced recombinantly or proteins delivered in the form of an inactivated HCV vaccine, such as conventional vaccines.
  • an isolated HCV polynucleotide or polypeptide or fragment, variant or derivative thereof is administered in an immunologically effective amount.
  • the effective amount of conventional vaccines is determinable by one of ordinary skill in the art based upon several factors, including the antigen being expressed, the age and weight of the subject, and the precise condition requiring treatment and its severity, and route of administration.
  • the combination of conventional antigen vaccine compositions with optimized nucleic acid or polypeptide compositions provides for therapeutically beneficial effects at dose sparing concentrations.
  • immunological responses sufficient for a therapeutically beneficial effect in patients predetermined for an approved commercial product, such as for the conventional product described above, can be attained by using less of the approved commercial product when supplemented or enhanced with the appropriate amount of nucleic acid or polypeptide.
  • a desirable level of an immunological response afforded by a DNA based pharmaceutical alone may be attained with less DNA by including an aliquot of a conventional vaccine.
  • using a combination of conventional and DNA based pharmaceuticals may allow both materials to be used in lesser amounts while still affording the desired level of immune response arising from administration of either component alone in higher amounts (e.g. one may use less of either immunological product when they are used in combination). This may be manifest not only by using lower amounts of materials being delivered at any time, but also to reducing the number of administrations points in a vaccination regime (e.g. 2 versus 3 or 4 injections), and/or to reducing the kinetics of the immunological response (e.g. desired response levels are attained in 3 weeks instead of 6 after immunization).
  • Determining the precise amounts of DNA based pharmaceutical and conventional antigen is based on a number of factors as described above, and is readily determined by one of ordinary skill in the art.
  • an adjuvant to increase the immune response to an antigen is typically manifested by a significant increase in immune -mediated protection.
  • an increase in humoral immunity is typically manifested by a significant increase in the titer of antibodies raised to the antigen
  • an increase in T-cell activity is typically manifested in increased cell proliferation, or cellular cytotoxicity, or cytokine secretion.
  • Nucleic acid molecules and/or polynucleotides of the present invention may be solubilized in any of various buffers.
  • Suitable buffers include, for example, phosphate buffered saline (PBS), normal saline, Tris buffer, and sodium phosphate (e.g., 150 mM sodium phosphate).
  • PBS phosphate buffered saline
  • Tris buffer Tris buffer
  • sodium phosphate e.g. 150 mM sodium phosphate
  • Insoluble polynucleotides may be solubilized in a weak acid or weak base, and then diluted to the desired volume with a buffer. The pH of the buffer may be adjusted as appropriate.
  • a pharmaceutically acceptable additive can be used to provide an appropriate osmolarity.
  • compositions of the present invention can be formulated according to known methods. Suitable preparation methods are described, for example, in Remington's Pharmaceutical Sciences, 16th Edition, A. Osol, ed., Mack Publishing Co., Easton, Pa. (1980), and Remington's Pharmaceutical Sciences, 19th Edition, A. R. Gennaro, ed., Mack Publishing Co., Easton, Pa. (1995).
  • composition may be administered as an aqueous solution, it can also be formulated as an emulsion, gel, solution, suspension, lyophilized form, or any other form known in the art.
  • composition may contain pharmaceutically acceptable additives including, for example, diluents, binders, stabilizers, and preservatives.
  • Example 1 is illustrative of specific embodiments of the invention, and various uses thereof. They are set forth for explanatory purposes only, and are not to be taken as limiting the invention.
  • Example 1 is illustrative of specific embodiments of the invention, and various uses thereof. They are set forth for explanatory purposes only, and are not to be taken as limiting the invention.
  • Example 1 is illustrative of specific embodiments of the invention, and various uses thereof. They are set forth for explanatory purposes only, and are not to be taken as limiting the invention.
  • Plasma samples were obtained from 17 subjects with acute or very recent HCV infection representing subtypes la, lb, 2 or 3. These consisted of twice-weekly serial collections from source plasma donors who became HCV infected during the course of their plasma donations. The donors were untreated and asymptomatic throughout the collection period. Plasma samples from 14 subjects with chronic HCV infection from the U.S. served as controls. All patients were treatment-na ' ive. All subjects gave informed consent, and plasma collections were performed with institutional review board and other regulatory approvals. Plasma samples were tested for HCV RNA and viral specific antigen and antibodies by a battery of commercial tests (Abbott and Roche) (Fig. 1 and Table 2).
  • a total of 154 samples were tested including a median of 8 sequential specimens per acutely infected subject (range 5-11).
  • the initial samples from acute subjects were HCV vRNA and antibody negative, followed by a sharp rise in vRNA levels.
  • Five of 17 acutely infected subjects developed HCV antibodies by the last sampling time point.
  • RNA isolation and cDNA synthesis was performed as follows. Samples which contained approximately 100,000 viral RNA copies were extracted using the Qiagen BioRobot EZ1 Workstation with EZ1 Virus Mini Kit v2.0 (Qiagen, Valencia, CA). RNA was eluted in 60 ⁇ and immediately subjected to cDNA synthesis. Reverse transcription of RNA to single stranded cDNA was performed using Superscript III reverse transcriptase (Invitrogen Life Technologies, Carlsbad, CA).
  • a mixture of 3 ⁇ (0.5 mM) of 10 mM deoxynucleoside triphosphate, 1.5 ⁇ (0.25 ⁇ ) of 10 ⁇ anti-sense primer and 34.5 ⁇ of vRNA was incubated at 65°C for 5 min and then was placed on ice for at least 1 min. The contents of the tube were collected by brief centrifugation. Next, 12 ⁇ of 5X reaction buffer, 3 ⁇ of 0.1 M DTT (5 mM), 3 ⁇ of RNAseOUT Recombinant RNASE Inhibitor, (40 units/ ⁇ , Invitrogen) and 3 ⁇ of Superscript IIITM Reverse Transcriptase (2001 ⁇ / ⁇ 1) were added to the mixture.
  • the final reaction was incubated at 50°C for 60 min followed by an increase in temperature to 55°C for an additional 60 minutes.
  • the reaction was heat-inactivated at 70°C for 15 minutes and then treated with RNaseH at 37°C for 20 minutes.
  • SEQ ID NOS The antisense primers were designed specifically for different genotype.
  • Single genome amplification was performed from prepared viral cDNA.
  • cDNA was serially diluted and distributed among wells of replicate 96-well plates so as to identify a dilution where PCR positive wells constituted less than 30% of the total number of reactions. At this dilution, most wells contained amplicons derived from a single cDNA molecule. This was confirmed in every positive well by direct sequencing of the amplicon and inspection of the sequence for mixed bases (double peaks), which would be evidence of priming from more than one original template or the introduction of PCR error in early cycles. Any sequence with evidence of mixed bases was excluded from further analysis.
  • PCR amplification was carried out in the presence of 10 x Taq High Fidelity Platinum PCR buffer, 2 mM MgS0 4 , 0.2 mM of each deoxynucleoside triphosphate, 0.2 ⁇ of each primer, and 0.1 ⁇ (0.5 Unit) Platinum Taq High Fidelity polymerase in a 20 ⁇ reaction (Invitrogen, Carlsbad, CA).
  • genotype 1 1 st round sense primer l .core.Fl 5'- ATGAGCACGAATCCTAAACCTCAAAGA-3' (SEQ ID NO:761)(nt 342-368 H77) and 1 st round antisense primer 1.NS4A.R1 5'-GCACTCTTCCATCTCATCGAACTC-3' (SEQ ID NO:763) (nt 5451-5474 H77), 2 nd round sense primer l .core.F2 5 ' -TC AAAG AAAAAC C AAA CGTAACACCAACCG-3' (SEQ ID NO:764) (nt 362-391 H77) and 2 nd round antisense primer 1.NS3A4A.R2 5'-AGGTGCTCGTGACGACCTCCAGG-3' (SEQ ID NO:766) (nt 5297-5319 H77); (2) 5' quarter genome of
  • PCR was performed in MicroAmp 96-well reaction plates (Applied Biosystems, Foster City, CA) with the following PCR parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94 °C for 15 s, an annealing step of 58°C for 30 s, an extension step of 68 °C for 5 min, followed by a final extension of 68 °C for 10 min.
  • the product of the 1 st round PCR was subsequently used as a template in the 2 nd round PCR under same conditions but with a total of 45 cycles.
  • sequences of 2850 were unambiguous at every position. Sequences chromatograms from 150 amplicons contained one or more "double peaks" representing mixed bases. Because mixed bases generally represented only a minority of polymorphisms, it was inferred that such ambiguities had resulted from Taq polymerase errors in the initial PCR cycles and not from amplification from more than one initial target template; in such cases a correction of the assignment of the base was made. In rare instances where the only polymorphic site(s) present were represented by double peaks, possibly due to mixed initial templates, the sequences were discarded and not included in the analysis. This amounted to 30 sequences or less than 1% of total amplicons generated.
  • Example 2 Example 2
  • a total of 38 transmitted/founder lineages from 17 acutely infected subjects were used to calculate the mutation rate.
  • Each sequence within the lineage was compared with the T/F virus sequence of that lineage.
  • the insertion, deletion, transition and transversion frequencies were counted by self-developed computer program and by Highlighter tool.
  • the rates were calculated by taking the ratio of each frequency number and the total number of nucleotides of all the sequences within that lineage.
  • the SNAP program (www.HIV.lanl.gov) was applied to the codon-aligned sequences of each T/F lineage. Within each lineage, the accumulations of synonymous and nonsynonymous substitutions were counted by comparing to the transmitted/founder viral sequences. The Jukes-Cantor corrected accumulation rates of synonymous substitutions per potential synonymous site (ds) and nonsynonymous substitution per potential nonsynonymous site (dn) were compared to screen for positive selection.
  • Nucleotide polymorphisms in subject 10051 were essentially random, corresponding to a star-like phylogeny and a Poisson distribution of low frequency events.
  • Two sequences (2C3 and 2C2) contained a single common polymorphism at position -2200, and two others (2A2 and 2B34) contained a different common polymorphism at position -4375, indicating for each pair shared recent ancestry.
  • Figures 4C and 4D depict sequences from subjects having evidence of productive infection by more than one genetically distinct virus. Each of the subjects was sampled on 4 occasions and sequences again showed increasing diversity over time. The consensus sequences of each low diversity lineage in these subjects were interpreted to correspond to a unique T/F HCV genome. Thus, each of these subjects was productively infected by at least three viruses. The proportion of sequences represented in each lineage was approximately the same in subject 10012 (Fig 4C) but not in subject 10062 (Fig 4D). This could be due either to different replication rates of the different T/F viral viruses, differential effects of early innate or adaptive immune responses, different times of infection by different transmitted viruses, or early stochastic events in the infection process.
  • a mathematical model was developed to estimate the adequacy of sampling given the observed distribution of T/F lineage sequences and total number of sequences analyzed.
  • This model when applied to the 179 sequences analyzed for subject 10012, estimated that the actual number of T/F viruses was 3 at a 95 % confidence level.
  • the model when applied to the 164 sequences analyzed for subject 10062, the model estimated that the actual number of T/F viruses could be as high as 5, again at a 95 % confidence level.
  • Sequences from subject 10029 where low diversity sequence lineages corresponding to 9 T/F viruses could be unambiguously identified, but with even greater differences in sequence proportions.
  • the model estimated that the actual number of T/F viruses could be as high as 13 at a 95 % confidence level.
  • T/F viruses If an acutely-infected subject acquired multiple HCV genomes from a contact who himself was acutely infected by a single virus, then these T/F viruses would be expected to all be closely related to each other, since they must reflect the viral diversity in the donor. If that donor contact had been acutely infected by more than one genetically divergent virus and the acutely infected recipient acquired multiple progeny of these, then the expectation would be for T/F viral sequences to be represented by subsets of sequences, some closely related and some not. Evidence in five subjects of acute-to-acute transmission was found by these two scenarios.
  • Figures 13A-C depict ML trees and Highlighter plots from acutely infected subjects 10020, 10016 and 10003, where the maximum diversity of T/F viruses from a consensus was 0.18 %, 0.16 % and 0.12 %, respectively.
  • the phylogenetic patterns of these sequences were quite different from the single variant transmissions shown in Figs 3B and 4 A and 4B, which exhibited star-like phylogeny and conformed to a Poisson distribution of random mutations.
  • the sequences in 13A-C instead, violated the Poisson model. They were comprised of distinct subsets of highly related sequences that differed from each other and from a common consensus by 1-6 nucleotides.
  • the phylogenetic pattern of sequences from subject 106889 was far more complicated than that of any of the other 16 acutely infected subjects.
  • Subject 106889 exhibited a typical acute infection viral kinetic profile with four sequential plasma samples negative for HCV vRNA followed by rapid vR A ramp-up to nearly 10 7 vRNA IU/ml (Fig 14 insert). Eighty-six 5 ' half genome sequences were obtained from the initial plasma vRNA positive time point.
  • T/F viral sequences in 17 subjects provided a unique strategy for analyzing HCV sequence evolution in vivo in the critical period beginning at or near the moment of virus transmission and extending to the establishment of viral load setpoint and in some subjects HCV antibody seroconversion.
  • a summary of this analysis is presented in Table 1.
  • Maximum intra-lineage diversity for the 17 subjects ranged from 0.06 % to 0.22 %. Insertions (0.000001), deletions (0.00004) and stop codons (0.000004) were infrequent. Transitions outnumbered transversions by 8 to 1.
  • the overall mutation frequency including all sampled time points was low (0.000145) given the number of possible virus replication cycles between the first infected hepatocyte and setpoint viremia 6-10 weeks later when as many as 7-20 % of hepatocytes may be productively infected.
  • the dN/dS ratio was low at 0.33 as was its derivative pN/pS at 0.33 %.
  • the diversity of sequences within the T/F lineages of each acutely infected subject corresponded to a star-like pattern of random diversification that conformed to a Poisson distribution mutations.
  • Exceptions were of four types: (i) infrequent shared polymorphisms resulting from stochastic changes in the newly infected subjects (e.g., Fig 3B; 4A, 4B, 4C, 4D); (ii) transmission of multiple closely related variants (e.g., Figs 11, 13 A, 13B; 14); (iii) evidence of immune selection in later samples (Fig 12); and (iv) rare examples of short inverted repeats resulting from template switching between double-stranded RNA duplexed hairpin structures by the RNA dependent polymerase (Fig 13).
  • Non- random amino acid polymorphisms in multiple T/F sequences in subjects 10020 (Fig 13 A) and 10016 (Fig 13B) also corresponded to known or predicted CTL epitopes and thus are likely to represent CTL escape or reversion in the respective donors (not in subjects 10020 or 10016).
  • HCV-1 transmission for acutely infected subjects is believed to result from high virus loads, each of circulating neutralizing antibodies and restricted viral genetic diversity, an interpretation corroborated by acute infection studies of SIV in Indian rhesus macaques. All of these same factors could pertain to acute -to-acute HCV transmission.
  • the complicated HCV transmission scenario for subject 106889 illustrates the sensitivity and specificity with which the SGA-direct amplicon sequencing strategy can detect and discriminate T/F viral genomes and their progeny.
  • SGA-direct sequencing includes in vitro recombination artifacts, avoids Taq polymerase-mediated nucleotide substitution errors in finished sequences, precludes founder effects due to disproportionate target amplification, and avoids cloning bias. This allows for a proportional representation of target molecules in finished sequences. It also allows for genetic linkages to be preserved in finished sequences specific mutations, in this case NS3 DAA resistance mutations V36M and R155K to be precisely identified.
  • each T/F virus contained the V36M/R155K double mutation indicating high level NS3 protease resistance.
  • Subject 106889 as well as the transmitting partner, SGA-direct sequencing is ideally suited to identifying genetically-linked DAA resistant mutants within single or multiple genes (e.g., NS3, NS5A and NS5B) of transmitted viruses and characterizing their persistence or disappearance over time.
  • the transmitted/founder virus sequence of the full length genome of subject 10021 was obtained by analysis of five overlapped genome fragments: 5' UTR (nt 1 to 852), 5' half genome (core, El, E2, p7, NS2 andNS3, nt 391 to 5297), 3' half genome (NS4A, NS4B, NS5A andNS5B, nt 5168 to 9374), 3' U/C tract (nt 9082 to 9582) and 3' x-tail (nt 9563 to 9646).
  • the primers designed for cDNA synthesis and PCR amplification are listed below.
  • the nucleotide numbering system is based on reference sequence H77 (accession number NC 004102). The sequences of all five genome regions were analyzed and assembled.
  • the transmitted/founder sequence was inferred based on phylogenetic inference and mathematical modeling.
  • the cDNA synthesis and single genome amplification of '5 and '3 half genome was performed as described supra in Example 2.
  • the positive PCR reactions were subject to direct sequencing.
  • the cDNA synthesis of 3' U/C tract of subject 10021 was also performed using Superscript IIITM Reverse Transcriptase.
  • the mixture of 1 ⁇ (0.5 mM) of 10 mM deoxynucleoside triphosphate, 0.5 ⁇ (0.25 ⁇ ) of 10 ⁇ anti-sense primer 3UTR-R10 and 11.50 ⁇ of vRNA was heated at 65 °C for 5 min and then was placed on ice for at least 1 min.
  • PCR amplification was carried out in the presence of 2 ⁇ of 10 x Taq High Fidelity Platinum PCR buffer, 0.8 ⁇ of 50 mM MgS0 4 (2 mM), 0.4 ⁇ of 10 mM deoxynucleoside triphosphate (0.2 mM), 0.2 ⁇ of each primer (0.2 ⁇ ), and 0.1 ⁇ (0.5 Unit) Platinum Taq High Fidelity polymerase in a 20 ⁇ reaction.
  • the semi-nested PCR primers used are listed in Table 4.
  • PCR was performed in MicroAmp 96-well reaction plates with the following PCR parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94°C for 15 s, an annealing step of 50 °C for 30 s, an extension step of 68°C for 1.5 min, followed by a final extension of 68 °C for 10 min.
  • the product of the 1 st round PCR was subsequently used as a template in the 2 nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on precasted 1 % agarose E-gels 96 and 4 % Nusieve GTG Agarose gel (CAMBREX, Cat. No. 50080).
  • the positive PCR reactions were TA cloned into pGEM-T Easy Vector System (Promega, Cat. No. A3610). The inserts were identified and sequenced using Ml 3 forward/reverse primers.
  • the transmitted/founder virus sequence of 5' UTR of subject 10021 was obtained using FirstChoice RLM-RACE kit (Ambion, Cat. No. AM 1700), with slight modification of the methods recommended by the manufacturer.
  • TAP Tobacco Acid Pyrophosphatase
  • vRNA acquires the adapter sequence as its 5' end
  • a core gene specific RT-PCR amplified the complete 5' UTR genome.
  • the cDNA synthesis and PCR amplification were performed under same conditions as that described for 3' U/C tract except using 55 °C as the annealing temperature.
  • the positive PCR products were sequenced directly.
  • RNA oligonucleotide adaptor (Integrated DMA Technologies INC.) w r as iigated to the 3' end of the vRNA by using T4 RNA ligase.
  • the adaptor was 5' phosphorylated and 3' Dideoxy-C blocked to prevent intra-molecular ligation.
  • the 20 ul ligation reaction that contained 15 ⁇ of vRNA, 2 ⁇ of 10X ligation buffer, 1 ⁇ of 10 rnM ATP, 0.5 ⁇ of RNase Inhibitor (20 ⁇ / ⁇ 1), 1 ⁇ of 10 ⁇ oligo adaptor ( 0.5 ⁇ ) and 1 ⁇ of T4 RNA ligase (10 units) was incubated at 37 °C for 60 min.
  • the whole ligation reaction was subject to cDNA synthesis directly using the oligo adaptor specific primer.
  • the cDNA synthesis was carried at 50 °C for 60 min.
  • PCR was performed with the following parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94 °C for 15 s, an annealing step of 55 °C for 30 s, an extension step of 68 °C for 30 sec, followed by a final extension of 68 °C for 10 min.
  • the product of the 1 st round PCR was subsequently used as a template in the 2 nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on 4 % Nusieve GTG Agarose gel (CAMBREX). The positive PCR reactions were directly TA cloned into pGEM-T Easy Vector System (Promega, Cat. No. A3610). The inserts were identified and sequenced using Ml 3 forward/reverse primers.
  • the inferred full-length transmitted/founder sequences was chemically synthesized in four subgenomic fragments (Blue Heron Biotechnology). The fragments were overlapped at the unique restriction sites, BsrGI (nt 3640), SnaBI (nt 6644) and Sfil (nt 9404), respectively.
  • BsrGI nt 3640
  • SnaBI nt 6644
  • Sfil nt 9404
  • the noncutter Notl was attached to the 5' terminal of the first fragment and noncutter Xbal was attached to the 3' terminal of all the fragments during chemical synthesis.
  • the synthesized fragments were propagated, restriction enzyme digested and sequentially cloned into the MCS Notl-Xbal site of pBlueScriptKS(+). The final full-length clone was confirmed by sequence analysis. Table 1. Diversity and mutation analyses of HCV sequences in acute infection.
  • a 5h-5' half genome contains core, El, E2, p7, NS2, and NS3;
  • 5ql-5' quarter 1 genome contains core, El, E2, p7 and partial NS2;
  • 5q2-5' quarter 2 genome contains partial NS2 and NS3.
  • h Averages were calculated from total mutations in all transmitted/founder lineages from all subjects combined. Because of low numbers of sequences and mutations in some lineages, certain values (e.g. dN/dS for subjects 10062 v2 and 10004 v3) vary substantially from the mean.
  • a obtained by applying jackknife to the biasd estimator in this case is the clusters found.
  • the average cut-off method is the more conservative estimate of the minimum number of founder strains needed to explain the observed diversity.
  • the maximum cut-off distinguishes more lineages and separates them into clusters from distinct founders.
  • NS3.F2 (SEQ ID NO: 785) 5146-5168 +, 2nd round
  • 3UTR-R10 (SEQ ID NO: 791) 9582-9604 -, RT and 1st round

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Organic Chemistry (AREA)
  • Wood Science & Technology (AREA)
  • Engineering & Computer Science (AREA)
  • Zoology (AREA)
  • Genetics & Genomics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Immunology (AREA)
  • General Engineering & Computer Science (AREA)
  • Communicable Diseases (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Biotechnology (AREA)
  • Biochemistry (AREA)
  • Virology (AREA)
  • General Health & Medical Sciences (AREA)
  • Microbiology (AREA)
  • Physics & Mathematics (AREA)
  • Biomedical Technology (AREA)
  • Analytical Chemistry (AREA)
  • Biophysics (AREA)
  • Medicinal Chemistry (AREA)
  • Molecular Biology (AREA)
  • Medicines Containing Antibodies Or Antigens For Use As Internal Diagnostic Agents (AREA)
  • Peptides Or Proteins (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

The invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection, and methods for identifying nucleotide sequences corresponding to specific transmitted HCV genomes.

Description

FULL-LENGTH TRANSMITTED HEPATITIS C VIRUS (HCV) GENOMES
IDENTIFIED BY SINGLE GENOME AMPLIFICATION
CROSS-REFERENCES TO RELATED APPLICATIONS
This application claims the benefit of priority to United States Provisional Patent Application Serial No. 61/630,161, filed on December 5, 2011, which is incorporated by reference in its entirety.
GOVERNMENT RIGHTS
This invention was supported in part with funding provided by NIH Grant Nos. NIH Chavi Award (NIH U01 AI 067854) and the University of Pennsylvania Center for AIDS Research (NIH P30 AI 045008), each awarded by the National Institutes of Health. The government may have certain rights to this invention.
FIELD OF THE INVENTION
The invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection. The invention provides methods for identifying nucleotide sequences corresponding to specific transmitted HCV genomes, including structural gene nucleotide sequences of global HCV genotypes, subtypes, and drug resistant variants that actively mediate viral infection. The invention further provides methods of administering a vaccine comprising transmitted full-length hepatitis C virus (HCV) genomes or portions thereof.
BACKGROUND OF THE INVENTION
Hepatitis C Virus (HCV) is a positive strand, non-segmented, enveloped RNA virus of approximately 9.6 kb in length. The virus is classified in the genus Hepacivirus within the larger family of Flavivirus, which includes the human pathogens West Nile virus, yellow fever virus and dengue fever virus among others. A common feature among the Flaviviridae is their dependence on a virally-encoded RNA-dependent RNA polymerase (RdRp) for replication. RdRp is error-prone, and HCV is notable for its quasispecies complexity and broad genotypic diversity. Globally, there are seven major genotypes of HCV that differ by approximately 30- 35% in nucleotide sequence.
The extraordinary diversity of HCV has important implications for clinical management and for basic and translational research aimed at elucidating viral natural history, pathogenesis and susceptibility to novel therapeutics and vaccines. Clinically, the different HCV genotypes exhibit variable natural history, responsiveness to interferon and ribavirin (mainstays of current therapy), and sensitivity to the many recently approved or still investigational direct acting antiviral (DAA) drugs. The development of new drugs and drug combinations that effectively suppress HCV replication and prevent the emergence of DAA resistance is a major challenge given the high rates of virus replication and variation. HCV diversity poses similar challenges to the development of effective vaccines and to the elucidation of virus biology, immunopathogenesis, and gene structure-function relationships. It is of interest and practical relevance that the extraordinary diversity of HCV is mirrored by comparable diversity of HIV- 1 , and that a novel experimental strategy to identify transmitted/founder (T/F) HIV-1 genomes in acute infection has led to new insights into HIV-1 transmission, persistence, and evolution in the face of cellular and humoral immune responses and antiretroviral therapy.
Acute HCV infection, defined as the period between virus transmission and antibody seroconversion which occurs about 6 - 10 weeks later, sets in motion viral-host interactions that largely dictate the natural history of the infection. Depending on viral genotype and host immunogenetic factors, most importantly IL28B alleles, a proportion of newly infected individuals spontaneously control or eliminate the virus. A greater number of patients can be cured if the infection is treated with interferon and ribavirin alone or in combination with DAA drugs. Mechanistically, how this occurs is unknown, and how the emergence of DAA drug resistance in communities of chronically infected subjects or in acutely infected subjects who initiate early therapy will affect treatment responses is uncertain, but again there are parallels with HIV-1. From a vaccine perspective, the acute infection period is critical. Transmitted viruses are the targets of a vaccine and early stages of infection represent a period when the virus should be most vulnerable to elimination by vaccine-elicited immune responses. For these reasons, there is considerable interest in the molecular features of the initial population 'bottleneck' to HCV transmission and subsequent pathways of virus evolution leading to viral persistence.
Previous reports have described different experimental approaches to the analysis of the HCV transmission bottleneck. These can generally be divided into studies that employed a DNA heteroduplex gel shift analytical method; studies that used conventional polymerase chain reaction (PCR) methods to amplify, clone and sequence fragments of the HCV genome; and studies that used 454 pyrosequencing to analyze acute and early viral sequences more deeply. All of these studies documented a restriction in viral diversity associated with virus transmission. However, despite the use of increasingly sensitive methods, a precise quantitative and molecular description of HCV transmission and early diversification remain elusive. This is because previous studies employed methods that were unable to sufficiently resolve HCV diversity between transmitted genomes that can be as low as 0.01%; employed Tag polymerase amplification from heterogeneous target templates, thereby allowing for artifactual in vitro recombination that confounds phylogenetic interpretation; employed 454 pyrosequencing, which yields short sequences lacking genetic linkage across genes and genomes; or evaluated few subjects.
Thus, there is a need in the art to allow for precise and unambiguous molecular identification of the complete full-length nucleotide sequences of HCV genome, particularly acute infection stage HCV, in order to gain a more comprehensive molecular understanding of HCV transmission and for development of effective HCV vaccines and treatments.
SUMMARY OF THE INVENTION
The invention provides transmitted full-length hepatitis C virus (HCV) genomes that mediate viral transmission and clinical infection. One application of this invention is to use these transmitted full-length genomes for development of more effective vaccines and treatments. In one embodiment the transmitted full-length HCV genome comprises the polynucleotide of SEQ ID NO: 776. In certain embodiments, the HCV genome mediates viral transmission. In other embodiments, the polypeptide sequence comprises env and core genes. Nucleotide sequence of transmitted HCV genomes is set forth in Table 5.
In other aspects, the invention provides methods for identifying full-length transmitted HCV genomes, the methods comprising collecting a patient sample, isolating and preparing viral RNA for sequencing, sequencing viral RNA that includes HCV genomes of circulating virus, performing sequence alignment of selected HCV genome regions, analyzing phylogenetically selected sequence alignments; and identifying full-length HCV genomes of transmitted virus. In particular embodiments the HCV genome polypeptide sequence comprises SEQ ID NO. 776.
In other aspects, the invention provides an immunogenic composition comprising transmitted full-length HCV genome or portions thereof. In particular embodiments the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
In other aspects, the invention provides methods for administering vaccines comprising transmitted full-length HCV genomes or portions thereof, wherein an immune response is induced in the patient following vaccination. In particular embodiments the HCV genome comprises the polynucleotide of SEQ ID NO: 776. The advantages of administering the disclosed HCV sequences and/or polypeptides encoded by such HCV sequences is such vaccines would illicit an immune response that is specific to transmitted HCV genomes rather than circulating HCV genomes, thereby improving vaccine effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS
These and other objects and features of this invention will be better understood from the following detailed description taken in conjunction with the drawings, wherein:
Figure 1 illustrates HCV R A kinetics in 16 acute infection subjects. The shaded area represents that the HCV RNA is below the linear range of quantitation (43-69,000,000 IU/ml). Plus signs indicate the samples are Anti-HCV antibody positive. The circles represent the time points that were sequence analyzed from each subject.
Figure 2 is a Maximum-likelihood tree (ML) of 5' quarter genome {core, El and E2) sequences from 17 acutely and 14 chronically infected subjects. The sequences from acutely infected subjects are shown in red and the sequences from chronically infected subjects are shown in blue. The reference sequences of genotype 1 to 7 are in gray. Bootstrap values (>70%) are indicated for intra-subject clusters. The horizontal scale bar represents 3% genetic distance.
Figure 3 illustrates 5' half-genome sequence diversity in two subjects, one with chronic infection (WIMI4025 (Fig 3 A)) and one with acute infection (10051 Fig 3(B)). Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. Bootstrap values (>70%) are indicated for intra-linerage clusters. The horizontal scale bars represent 0.5% (24.5 nt) and 0.02% (1 nt) genetic distance for subject WIMI and 10051, respectively.
Figure 4 illustrates 5' quarter-genome {Core, El and E2) sequence diversity in four acutely infected subjects. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. Sequences from multiple time points are color coded with orange, green, blue and black in chronological order. Sequences from subject 10021 (Fig 4A) and 10025 (Fig 4B) each showed productive infection by a single virus. Sequences from subject 10012 (Fig 4C) and 10062 (Fig 4D) showed productive infection by at least three viruses, repectively. The horizotal scale bars represent 0.04% (1 nt) genetic distance for Figs. 4A, 4B and 4D and 0.4% (10 nt) for 4C. Bootstrap values (>70%) are indicated for within patient intra-lineage clusters in the ML tree.
Figure 5 illustrates 5' quarter-genome {Core, El and E2) sequence divsersity in an acutely infected subject. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' quarter genome {Core, El and E2) sequences generated from multiple time points showed productive infection by at least 9 transmitted/founder variants. The close genetic diversity between variant 1 and 2, variant 6 and 7 is indicative of transmission from an acutely infected donor. The horizotal scale bars represent 0.2% (5 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra- lineage clusters.
Figure 6 illustrates HCV diversity in acutely infected subject 10024. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' quarter genome sequences derived from multiple time points showed productive infection by at least 7 transmitted/founder variants. The horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 7 illustrates HCV diversity in acutely infected subject 6123. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 3 transmitted/founder variants. The horizotal scale bars represent 0.2% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 8 illustrates HCV diversity in acutely infected subject 6222. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 4 transmitted/founder variants. The horizotal scale bars represent 0.02% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 9 illustrates HCV diversity in acutely infected subject 10004. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of the 5' quarter genome sequences showed productive infection by at least 3 transmitted/founder variants. The horizotal scale bars represent 0.36% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 10 illustrates HCV diversity in acutely infected subject 10002. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of the 5' half genome sequences showed productive infection by at least 13 transmitted/founder variants. The horizotal scale bars represent 0.2% (10 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra- lineage clusters.
Figure 11 illustrates HCV diversity in acutely infected subject 10017. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' quarter genome sequences derived from multiple time points showed productive infection by at least 6 transmitted/founder variants. The horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 12 illustrates HCV diversity in acutely infected subject 9055. Neighbor-joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The ML tree and highlighter plot of 5' half genome sequences derived from multiple time points showed productive infection by a single transmitted/founder variants. The horizotal scale bars represent 0.04% (1 nt) genetic distance. Bootstrap values (>70%) are indicated for patient intra-lineage clusters.
Figure 13 illustrates 5' quarter-genome (Core, El and E2) sequence divsersity in acute- to-acute transmission. 5' quarter genome (Core, El and E2) sequences from subject 10020 (Fig 13 A) and 10016 (Fig 13B), and 5' half genome (core, El, E2, p7, NS2 and NS3A) sequences from subject 10003 (Fig 13C) are depicted by ML trees and by highlighter plots. Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The first sequence in red is the consensus sequence derived from the entire sequence dataset. The horizotal scale bars represent 0.04%> (1 nt), 0.036%) (1 nt) and 0.02%> (1 nt) genetic distance for Figs. 13 A, 13B and 13C, repectively. Bootstrap values (>70%) are indicated.
Figure 14 illustrates HCV RNA kinectics and diversity in subject 106889. The inset shows the HCV RNA kinetics. The time point that was sequence analyzed is circled. Neighbor- joining phylogenetic tree and a Highlighter plot are shown in the left and right parts of each figure, respectively. The 5' half genome (core, El, E2, p7, NS2 and NS3A) sequences formed 10 distinct lineages (labeled with letter L) with high statistical support in the ML tree. Each cluster is color coded. The T/F variants was identified within each lineage. A total of 33 T/F variants were found that were responsible for productive infection. The horizotal scale bars represent 0.04%) (1 nt) genetic distance. Bootstrap values (>70%>) are indicated.
Figure 15 illustrates strand transfers at stem loop structures of HCV RNA. Figures 17A- E illustrate four examples of stem loop structures in double-stranded HCV RNA leading to template switching by the HCV RNA polymerase. Fig 17A, SEQ ID NO: 800 First 5' to 3' horizontal sequence; SEQ ID NO: 801 Second 3' to 5' horizontal sequence; SEQ ID NO: 802 Third 5' to 3' horizontal sequence; SEQ ID NO: 800 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 802 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 801 Sixth 3' to 5' haripin sequence. Fig 17B, SEQ ID NO: 803 First 5' to 3' horizontal sequence; SEQ ID NO: 804 Second 3' to 5' horizontal sequence; SEQ ID NO: 805 Third 5' to 3' horizontal sequence; SEQ ID NO: 803 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 805 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 804 Sixth 3' to 5' hairpin sequence. Fig 17C, SEQ ID NO: 806 First 5' to 3' horizontal sequence; SEQ ID NO: 807 Second 3' to 5' horizontal sequence; SEQ ID NO: 808 Third 5' to 3' horizontal sequence; SEQ ID NO: 806 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 808 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 807 Sixth 3' to 5' hairpin sequence. Fig 17D, SEQ ID NO: 809 First 5' to 3' horizontal sequence; SEQ ID NO: 810 Second 3' to 5' horizontal sequence; SEQ ID NO: 811 Third 5' to 3' horizontal sequence; SEQ ID NO: 809 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 811 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 810 Sixth 3' to 5' hairpin sequence. Fig 17E, SEQ ID NO: 812 First 5' to 3' horizontal sequence; SEQ ID NO: 813 Second 3' to 5' horizontal sequence; SEQ ID NO: 814 Third 5' to 3' horizontal sequence; SEQ ID NO: 812 Fourth 5' to 3' hairpin sequence; SEQ ID NO: 814 Fifth 5' to 3' hairpin sequence; SEQ ID NO: 813 Sixth 3' to 5' hairpin sequence.
Figure 16 illsutrates the strategy for identification of full-length HCV genome by single gneome amplification (SGA). The complete HCV genome of subject 10021 was determined by amplifying five overlapping genome fragments: fragment I (nt 1 - 852) contains the complete 5' UTR and partial Core; fragment II (nt 391 - 5297) contains Core, El, E2, NS2 and NS3; fragment III (nt 5168 - 9374) contains NS4A, NS4B, NS5A and NS5B; fragment IV (nt 9082 - 9582) contains partial NS5B, variable region and the complete poly U/UC tract; and fragment V (nt 9563 - 9646) contains the complete x-tail. The nucleotide positions are based reference sequence H77.
Figure 17 illustrates the strategy for generating a full-length T/F HCV molcular clone from subject 10021.
Figure 18 illustrates HCV diversity in subject 10003. ML tree and Highlighter plot of 5' halfgenome sequences reveal many sets of closely related sequences distinguished by unique shared mutations.
Figure 19 illustrates HCV diversity analysis in subject 10016 suggests acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 5' quarter 1 genome sequences. Visualization of 15 potential T/F viral lineages distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveal 10 and 4 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
Figure 20 illustrates maximum diversity of discrete HCV sequence lineages from acute infection subjects versus maximum sequence diversity in chronic subjects. Primary data are derived from Tables 1 and 2. Mean (±95% CI) values are represented by horizontal lines. Differences between the two groups were highly significant (p<0.0001; unpaired T-test with Welch's correction), reflecting the recent and remote diversification histories of acute and chronic sequences, respectively.
Figure 21 shows HCV diversity in acute subject 10051. 5' quarter 1 genomesequences are color coded in orange, green, blue and black in chronological order to reflect sampling time points in Figure 1 and are represented in a ML tree and Highlighter plot. Sequences show evidence of productive clinical infection by a single virus. The horizontal scale bar indicates genetic distance.
Figure 22 illustrates nonsynonymous and synonymous mutations in HCV sequences from acute subjects 10003, 10020 and 10016. Highlighter plotsof 5' half or quarter 1 genomesequences are color coded to denote nonynonymous (red) and synonymous (green) mutations for subjects 10003 (panel A), 10020 (panel B) and 10016 (panel C).
Figure 23 illustrates amino acid alignment of the HCV Env coding region of acute subject 10003. The H77 reference sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
Figure 24 illustrates amino acid alignment of the HCV Env coding region of acute subject 10020. The H77 reference sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
Figure 25 illustrates amino acid alignment of the HCV Env coding region of acute subject 10016. The H77 sequence is shown at the top. Nonrandom concentrations of nonsynonymous mutations are evident.
Figure 26 shows HCV diversity analysis in subject 10020 to suggest acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 5' quarter 1 genome sequences. Visualization of 10 potential T/F viral sequences distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveals 5 and 3 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions.
Figure 27 shows HCV diversity analysis in subject 10003 to suggest acute-to-acute transmission. Highlighter plot and neighbor-joining tree of 59 half genome sequences. Visualization of 37 potential T/F viral sequences distinguished by unique shared mutations is indicated by lower case (blue) letters. Model estimates of T/F virus lineages using maximum (red) and average (green) cut-offs reveals 15 and 8 potential T/F virus lineages, respectively, based on increasingly stringent model assumptions. Figure 28 illustrates nonsynonymous and synonymous mutations in HCV sequences from acute subject 9055. A Highlighter plot (panel A) of 5' half genomesequences is color coded to denote nonynonymous (red) and synonymous (green) mutations. The boxed area reveals a temporal expansion of sequences with concentrated amino acid subtitutions in NS3. In panel B, amino acid selection is evident in a previously identified CTL epitope highlighted in red. The top-most sequence represents the genotype 3a consensus.
DETAILED DESCRIPTION
The invention provides full-length transmitted HCV genomes that mediate viral infection and transmission. Prior to the present invention the identification of these genomes was unavailable due to the following: (i) a significant time period from weeks to months between the moment of transmission and the first appearance of HCV in the blood (Bowen. D.G. and Walker, CM., 2005, Nature 436: 946-952; Moradpour, D. et al, 2007, Nature Rev Micro 5:453- 463); (ii) HCV is genetically highly variable in its nucleotide sequence due to its error-prone RNA-dependent RNA polymerase, and as a result, it exists in individuals as a complex mixture of sequences commonly referred to as a 'quasispecies' (Moradpour, D. et al., 2007, Nature Rev Micro 5:453-463); (iii) conventional experimental approaches to sequencing the HCV genome from clinical samples introduced addition variation into the sequences as a consequence of Taq polymerase induced recombination and nucleotide misincorporation errors (Salazar-Gonzalez, J.F. et al., 2008, J Virol 82:3952-3970); and (iv) identifying from these myriad of sequences which sequences corresponded to actual transmitted viruses, or even viruses that were replication-competent and responsible for ongoing virus replication and persistence in the infected subject was not achievable. The identification of the full-length transmitted viral sequence provides specific nucleotide and protein regions responsible for mediating viral infection. In certain embodiments the polypeptide sequence comprises env and core genes. Utilizing these specific regions in the preparation of vaccines and therapeutic provides for more effective HCV treatments.
The term "polynucleotide" is intended to encompass a singular nucleic acid or nucleic acid fragment as well as plural nucleic acids or nucleic acid fragments, and refers to an isolated molecule or construct, e.g., a virus genome (e.g., vRNA), messenger RNA (mRNA), plasmid DNA (pDNA), or derivatives of pDNA (e.g., minicircles as described in (Darquet, A-M et al., 1997, Gene Therapy, 4: 1341-1349) comprising a polynucleotide. A polynucleotide may comprise a conventional phosphodiester bond or a non-conventional bond (e.g., an amide bond, such as found in peptide nucleic acids (PNA)). The terms "nucleic acid" or "nucleic acid fragment" refer to any one or more nucleic acid segments, e.g., DNA or R A fragments, present in a polynucleotide or construct. A nucleic acid or fragment thereof may be provided in linear (e.g., mR A) or circular (e.g., plasmid) form as well as double-stranded or single-stranded forms. By "isolated" nucleic acid or polynucleotide is intended a nucleic acid molecule, DNA or RNA, which has been removed from its native environment. For example, a recombinant polynucleotide contained in a vector is considered isolated or "cloned" for the purposes of the present invention. Further examples of an isolated polynucleotide include recombinant polynucleotides maintained in heterologous host cells or purified (partially or substantially) polynucleotides in solution. Isolated RNA molecules include in vivo or in vitro RNA transcripts of the polynucleotides of the present invention. Isolated polynucleotides or nucleic acids according to the present invention further include such molecules produced synthetically.
The terms "fragment," "variant," and "derivative" when referring to HCV polypeptides of the present invention include any polypeptides which retain at least some of the immunogenicity or antigenicity of the corresponding native polypeptide. Fragments of HCV polypeptides of the present invention include proteolytic fragments, deletion fragments and in particular, fragments of HCV polypeptides which exhibit increased secretion from the cell or higher immunogenicity or reduced pathogenicity when delivered to an animal. Polypeptide fragments further include any portion of the polypeptide which comprises an antigenic or immunogenic epitope of the native polypeptide, including linear as well as three-dimensional epitopes. Variants of HCV polypeptides of the present invention include fragments, and also polypeptides with altered amino acid sequences due to amino acid substitutions, deletions, or insertions. Variants may occur naturally, such as an allelic variant. By an "allelic variant" is intended alternate forms of a gene occupying a given locus on a chromosome or genome of an organism or virus. Genes II, Lewin, B., ed., John Wiley & Sons, New York (1985), which is incorporated herein by reference. For example, as used herein, variations in a given gene product is a "variant". Naturally or non-naturally occurring variations such as amino acid deletions, insertions or substitutions may occur. Non-naturally occurring variants may be produced using art-known mutagenesis techniques. Variant polypeptides may comprise conservative or non-conservative amino acid substitutions, deletions or additions. Derivatives of HCV polypeptides of the present invention, are polypeptides which have been altered so as to exhibit additional features not found on the native polypeptide. Examples include fusion proteins. An analog is another form of an HCV polypeptide of the present invention. An example is a proprotein which can be activated by cleavage of the proprotein to produce an active mature polypeptide. In certain embodiments, the polynucleotide, nucleic acid, or nucleic acid fragment is DNA. In the case of DNA, a polynucleotide comprising a nucleic acid which encodes a polypeptide normally also comprises a promoter and/or other transcription or translation control elements operably associated with the polypeptide-encoding nucleic acid fragment. An operable association is when a nucleic acid fragment encoding a gene product, e.g., a polypeptide, is associated with one or more regulatory sequences in such a way as to place expression of the gene product under the influence or control of the regulatory sequence(s). Two DNA fragments (such as a polypeptide-encoding nucleic acid fragment and a promoter associated with the 5' end of the nucleic acid fragment) are "operably associated" if induction of promoter function results in the transcription of mRNA encoding the desired gene product and if the nature of the linkage between the two DNA fragments does not (1) result in the introduction of a frame-shift mutation, (2) interfere with the ability of the expression regulatory sequences to direct the expression of the gene product, or (3) interfere with the ability of the DNA template to be transcribed. Thus, a promoter region would be operably associated with a nucleic acid fragment encoding a polypeptide if the promoter was capable of effecting transcription of that nucleic acid fragment. The promoter may be a cell-specific promoter that directs substantial transcription of the DNA only in predetermined cells. Other transcription control elements, besides a promoter, for example enhancers, operators, repressors, and transcription termination signals, can be operably associated with the polynucleotide to direct cell-specific transcription. Suitable promoters and other transcription control regions are disclosed herein.
A variety of transcription control regions are known to those skilled in the art. These include, without limitation, transcription control regions which function in vertebrate cells, such as, but not limited to, promoter and enhancer segments from cytomegaloviruses (the immediate early promoter, in conjunction with intron-A), simian virus 40 (the early promoter), and retroviruses (such as Rous sarcoma virus). Other transcription control regions include those derived from vertebrate genes such as actin, heat shock protein, bovine growth hormone and rabbit β-globin, as well as other sequences capable of controlling gene expression in eukaryotic cells. Additional suitable transcription control regions include tissue-specific promoters and enhancers as well as lymphokine-inducible promoters (e.g., promoters inducible by interferons or interleukins).
Similarly, a variety of translation control elements are known to those of ordinary skill in the art. These include, but are not limited to ribosome binding sites, translation initiation and termination codons, elements from picornaviruses (particularly an internal ribosome entry site, or IRES, also referred to as a CITE sequence). A DNA polynucleotide of the present invention may be a circular or linearized plasmid or vector, or other linear DNA which may also be non-infectious and nonintegrating (i.e., does not integrate into the genome of vertebrate cells). A linearized plasmid is a plasmid that was previously circular but has been linearized, for example, by digestion with a restriction endonuclease. Linear DNA may be advantageous in certain situations as discussed, e.g., in Cherng, J. Y., et al., 1999, J Control. Release 60:343-53, and Chen, Z. Y., et al, 2001, Mol Ther 3:403-10. As used herein, the terms plasmid and vector can be used interchangeably. In other embodiments, a polynucleotide of the present invention is RNA, for example, in the form of messenger RNA (mRNA). Methods for introducing RNA sequences into vertebrate cells are described in U.S. Pat. No. 5,580,859.
Polynucleotides, nucleic acids, and nucleic acid fragments of the present invention may be associated with additional nucleic acids which encode secretory or signal peptides, which direct the secretion of a polypeptide encoded by a nucleic acid fragment or polynucleotide of the present invention. According to the signal hypothesis, proteins secreted by mammalian cells have a signal peptide or secretory leader sequence which is cleaved from the mature protein once export of the growing protein chain across the rough endoplasmic reticulum has been initiated. Those of ordinary skill in the art are aware that polypeptides secreted by vertebrate cells generally have a signal peptide fused to the N-terminus of the polypeptide, which is cleaved from the complete or "full length" polypeptide to produce a secreted, or "mature" form of the polypeptide. In certain embodiments, the native leader sequence is used, or a functional derivative of that sequence that retains the ability to direct the secretion of the polypeptide that is operably associated with it. Alternatively, a heterologous mammalian leader sequence, or a functional derivative thereof, may be used. For example, the wild-type leader sequence may be substituted with the leader sequence of human tissue plasminogen activator (TP A) or mouse beta-glucuronidase.
The term "expression" refers to the biological production of a product encoded by a coding sequence. In most cases a DNA sequence, including the coding sequence, is transcribed to form a messenger-RNA (mRNA). The messenger-RNA is then translated to form a polypeptide product which has a relevant biological activity. Also, the process of expression may involve further processing steps to the RNA product of transcription, such as splicing to remove introns, and/or post-translational processing of a polypeptide product. As used herein, the term "polypeptide" is intended to encompass a singular "polypeptide" as well as plural "polypeptides," and comprises any chain or chains of two or more amino acids. Thus, as used herein, terms including, but not limited to "peptide," "dipeptide," "tripeptide," "protein," "amino acid chain," or any other term used to refer to a chain or chains of two or more amino acids, are included in the definition of a "polypeptide," and the term "polypeptide" can be used instead of, or interchangeably with any of these terms. The term further includes polypeptides which have undergone post-translational modifications, for example, glycosylation, acetylation, phosphorylation, amidation, derivatization by known protecting/blocking groups, proteolytic cleavage, or modification by non-naturally occurring amino acids.
Also included as polypeptides of the present invention are fragments, derivatives, analogs, or variants of the foregoing polypeptides, and any combination thereof. Polypeptides, and fragments, derivatives, analogs, or variants thereof of the present invention can be antigenic and immunogenic polypeptides related to HCV polypeptides, which are used to prevent or treat, i.e., cure, ameliorate, lessen the severity of, or prevent or reduce contagion of infectious disease caused by the HCV.
As used herein, an "antigenic polypeptide" or an "immunogenic polypeptide" is a polypeptide which, when introduced into a vertebrate, reacts with the vertebrate's immune system molecules, i.e., is antigenic, and/or induces an immune response in the vertebrate, i.e., is immunogenic. It is quite likely that an immunogenic polypeptide will also be antigenic, but an antigenic polypeptide, because of its size or conformation, may not necessarily be immunogenic. Isolated antigenic and immunogenic polypeptides of the present invention in addition to those encoded by polynucleotides of the invention, may be provided as a recombinant protein, a purified subunit, a viral vector expressing the protein, or may be provided in the form of an inactivated HCV vaccine, e.g., a live-attenuated virus vaccine, a heat-killed virus vaccine, etc. By an "isolated" HCV polypeptide or a fragment, variant, or derivative thereof is intended an HCV polypeptide or protein that is not in its natural form. No particular level of purification is required. For example, an isolated HCV polypeptide can be removed from its native or natural environment. Recombinantly produced HCV polypeptides and proteins expressed in host cells are considered isolated for purposed of the invention, as are native or recombinant HCV polypeptides which have been separated, fractionated, or partially or substantially purified by any suitable technique, including the separation of HCV virions from culture cells in which they have been propagated. In addition, an isolated HCV polypeptide or protein can be provided as a live or inactivated viral vector expressing an isolated HCV polypeptide and can include those found in inactivated HCV vaccine compositions. Thus, isolated HCV polypeptides and proteins can be provided as, for example, recombinant HCV polypeptides, a purified subunit of HCV, a viral vector expressing an isolated HCV polypeptide, or in the form of an inactivated or attenuated HCV vaccine. The term "epitopes," as used herein, refers to portions of a polypeptide having antigenic or immunogenic activity in a vertebrate, for example a human. An "immunogenic epitope," as used herein, is defined as a portion of a protein that elicits an immune response in an animal, as determined by any method known in the art. The term "antigenic epitope," as used herein, is defined as a portion of a protein to which an antibody or T-cell receptor can immunospecifically bind as determined by any method well known in the art. Immunospecific binding excludes nonspecific binding but does not exclude cross-reactivity with other antigens. Where all immunogenic epitopes are antigenic, antigenic epitopes need not be immunogenic.
As to the selection of peptides or polypeptides bearing an antigenic epitope (e.g., that contain a region of a protein molecule to which an antibody or T cell receptor can bind), it is well known in that art that relatively short synthetic peptides that mimic part of a protein sequence are routinely capable of eliciting an antiserum that reacts with the partially mimicked protein. See, e.g., Sutcliffe, J. G., et al., 1983, Science 219:660-666.
The present invention also provides methods for identifying transmitted full-length HCV genomes that mediate viral infection and transmission. In one embodiment, the methods comprise: (a) collecting a patient sample; (b) isolating viral RNA from said sample; (c) sequencing said viral RNA, wherein viral RNA sequencing includes HCV genomes of circulating virus; (d) performing sequence alignment of selected HCV genome regions; (e) analyzing phylogenetically selected sequence alignments; and (f) identifying HCV genomes of transmitted virus.
As used herein, a "patient" or "subject" to be utilized by the disclosed methods can mean a human, chimpanzee or non-human primate. In certain embodiments the patient is a non-human mammal.
The term "patient sample" as used herein includes but is not limited to a blood, serum, plasma, or urine sample obtained from a patient. In particular embodiments, the patient sample is plasma.
The phrase "analyzing phylogentically" as used herein represents the analysis of clinical viral isolates, regardless of the particular methodology employed. Comparative analysis of the genetic relatedness of any a collection of circulating viral isolates is used to select for nucleotide sequence of transmitted HCV genomes. This methodology is used to select for the actual HCV genomes present at the time of patient infection/viral transmission. This can be accomplished by any number of methods, including but not limited to: i) a novel mathematical model of random virus evolution disclosed herein; ii) star phylogeny; iii) Baysian analysis; or any other method of phylogenetic analysis known to one of skill in the art. The phrase "performing sequence alignment" as used herein is meant to include aligning genetic sequences by any of a number of different procedures that produce a sufficient match between the corresponding residue in the sequences. Typically, Smith- Waterman or Needleman- Wunsch algorithms are used. However, other procedures such as BLAST, FASTA, PSI-BLAST can be used.
The phrase "isolating viral RNA" as used herein includes those methods well known in the art for extraction of viral RNA from a sample, cDNA synthesis from viral RNA template, optionally cloning of cDNA fragments, and/or amplification of polynucleotide sequence. Such methods are described in more detail, but such experimental procedures are provided in Sambrook and Russell, 2001, Molecular Cloning: A Laboratory Manual, Third Ed., Cold Spring Harbor Laboratory Press, Woodbury, N.Y.
The practice of this invention can involve procedures well known in the art, including for example nucleotide sequence amplification, such as polymerase chain reaction (PCR) and modifications thereof (including for example reverse transcription (RT)-PCR, and stem-loop PCR), as well as reverse transcription and in vitro transcription. Generally these methods utilize one or a pair of oligonucleotide primers having sequence complimentary to sequences 5' and 3' to the sequence of interest, and in the use of these primers they are hybridized to a nucleotide sequence and extended during the practice of PCR amplification using DNA polymerase (preferably using a thermal-stable polymerase such as Taq polymerase). RT-PCR may be performed on miRNA or mRNA with a specific 5' primer or random primers and appropriate reverse transcription enzymes such as avian (AMV-RT) or murine (MMLV-RT) reverse transcriptase enzymes.
The phrase "over time" as used herein represents a period of time between the collection of patient samples. For example, a patient sample is collected at time point A and then subsequently as a later date at time point B. The period between samplings is over time. In certain embodiments, samples are taken from the same patient. In alternative embodiments, samples may be taken from different patients. Performing the method of the invention at different time points followed by the differential analysis of HCV genomes at the different time points provides a means for assessing the evolution of HCV genomes both inter- and intra- patient. The identification of highly variable genomic regions is useful for the generation of effective therapeutics and vaccines to HCV.
The term "transmitted" as used herein refers to viral genotype and phenotype at the time of viral infection. The term "transmitted virus" as used here is in reference to actual HCV virus that mediates patient infection. The phrase "transmitted viral sequence" as used herein is meant to include nucleotide or amino acid sequence of HCV genomes of transmitted virus, an in a particular embodiment sequence corresponding to acute infection stage HCV. The term "circulating" refers to HCV virus collected from patient post-infection. In general, patient samples comprise circulating virus because HCV infection has occurred at a prior time point. The half-life of plasma virus is less than 1 day, thus the collection of transmitted virus from a patient sample would be a rare event.
The term "selected" as used in the phrases "selected sequence alignments" or "selected HCV genome regions" is meant to include the identification and utilization of particular subsets of nucleotide sequence and/or genome regions for subsequent analysis. In certain embodiments, full-length HCV genomes are identified, however discrete portions of the genome are utilized or "selected" for further analysis.
It is to be noted that the term "a" or "an" entity refers to one or more of that entity; for example, "a polynucleotide," is understood to represent one or more polynucleotides. As such, the terms "a" (or "an"), "one or more," and "at least one" can be used interchangeably herein.
The present invention also provides immunogenic compositions and methods for delivery of transmitted full-length HCV polynucleotide or polypeptide sequences to a vertebrate with optimal expression and safety conferred. In particular instances the transmitted full-length HCV genome comprises the polynucleotide of SEQ ID NO. 776. These immunogenic compositions may be prepared and administered in such a manner that the encoded gene products are optimally expressed in the vertebrate of interest. As a result, these compositions and methods are useful in stimulating an immune response against HCV infection. Also included in the invention are expression systems and delivery systems.
The identified polynucleotides or polypeptides encoded by the polynucleotides of the invention may be in any form, and polypeptides are generated using techniques well known in the art. Examples include isolated HCV proteins produced recombinantly or proteins delivered in the form of an inactivated HCV vaccine, such as conventional vaccines.
When utilized, an isolated HCV polynucleotide or polypeptide or fragment, variant or derivative thereof is administered in an immunologically effective amount. The effective amount of conventional vaccines is determinable by one of ordinary skill in the art based upon several factors, including the antigen being expressed, the age and weight of the subject, and the precise condition requiring treatment and its severity, and route of administration.
In the instant invention, the combination of conventional antigen vaccine compositions with optimized nucleic acid or polypeptide compositions provides for therapeutically beneficial effects at dose sparing concentrations. For example, immunological responses sufficient for a therapeutically beneficial effect in patients predetermined for an approved commercial product, such as for the conventional product described above, can be attained by using less of the approved commercial product when supplemented or enhanced with the appropriate amount of nucleic acid or polypeptide.
A desirable level of an immunological response afforded by a DNA based pharmaceutical alone may be attained with less DNA by including an aliquot of a conventional vaccine. Further, using a combination of conventional and DNA based pharmaceuticals may allow both materials to be used in lesser amounts while still affording the desired level of immune response arising from administration of either component alone in higher amounts (e.g. one may use less of either immunological product when they are used in combination). This may be manifest not only by using lower amounts of materials being delivered at any time, but also to reducing the number of administrations points in a vaccination regime (e.g. 2 versus 3 or 4 injections), and/or to reducing the kinetics of the immunological response (e.g. desired response levels are attained in 3 weeks instead of 6 after immunization).
Determining the precise amounts of DNA based pharmaceutical and conventional antigen is based on a number of factors as described above, and is readily determined by one of ordinary skill in the art.
The ability of an adjuvant to increase the immune response to an antigen is typically manifested by a significant increase in immune -mediated protection. For example, an increase in humoral immunity is typically manifested by a significant increase in the titer of antibodies raised to the antigen, and an increase in T-cell activity is typically manifested in increased cell proliferation, or cellular cytotoxicity, or cytokine secretion.
Nucleic acid molecules and/or polynucleotides of the present invention, e.g., plasmid DNA, mRNA, linear DNA or oligonucleotides, may be solubilized in any of various buffers. Suitable buffers include, for example, phosphate buffered saline (PBS), normal saline, Tris buffer, and sodium phosphate (e.g., 150 mM sodium phosphate). Insoluble polynucleotides may be solubilized in a weak acid or weak base, and then diluted to the desired volume with a buffer. The pH of the buffer may be adjusted as appropriate. In addition, a pharmaceutically acceptable additive can be used to provide an appropriate osmolarity. Such additives are within the purview of one skilled in the art. For aqueous compositions used in vivo, sterile pyrogen-free water can be used. Such formulations will contain an effective amount of a polynucleotide together with a suitable amount of an aqueous solution in order to prepare pharmaceutically acceptable compositions suitable for administration to a human. Compositions of the present invention can be formulated according to known methods. Suitable preparation methods are described, for example, in Remington's Pharmaceutical Sciences, 16th Edition, A. Osol, ed., Mack Publishing Co., Easton, Pa. (1980), and Remington's Pharmaceutical Sciences, 19th Edition, A. R. Gennaro, ed., Mack Publishing Co., Easton, Pa. (1995). Although the composition may be administered as an aqueous solution, it can also be formulated as an emulsion, gel, solution, suspension, lyophilized form, or any other form known in the art. In addition, the composition may contain pharmaceutically acceptable additives including, for example, diluents, binders, stabilizers, and preservatives.
The invention illustratively described herein suitably can be practiced in the absence of any element or elements, limitation or limitations that are not specifically disclosed herein. Thus, for example, in each instance herein any of the terms "comprising", "consisting essentially of, and "consisting of may be replaced with either of the other two terms, while retaining their ordinary meanings. The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention that in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention has been specifically disclosed by embodiments, optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the description and the appended claims.
This invention is more particularly described below and the Examples set forth herein are intended as illustrative only, as numerous modifications and variations therein will be apparent to those skilled in the art. As used in the description herein and throughout the claims that follow, the meaning of "a", "an", and "the" includes plural reference unless the context clearly dictates otherwise. The terms used in the specification generally have their ordinary meanings in the art, within the context of the invention, and in the specific context where each term is used. Some terms have been more specifically defined below to provide additional guidance to the practitioner regarding the description of the invention.
Examples
The Examples which follow are illustrative of specific embodiments of the invention, and various uses thereof. They are set forth for explanatory purposes only, and are not to be taken as limiting the invention. Example 1
HCV Patient Selection and Analysis
HCV patients were selected and samples collected as described. Plasma samples were obtained from 17 subjects with acute or very recent HCV infection representing subtypes la, lb, 2 or 3. These consisted of twice-weekly serial collections from source plasma donors who became HCV infected during the course of their plasma donations. The donors were untreated and asymptomatic throughout the collection period. Plasma samples from 14 subjects with chronic HCV infection from the U.S. served as controls. All patients were treatment-na'ive. All subjects gave informed consent, and plasma collections were performed with institutional review board and other regulatory approvals. Plasma samples were tested for HCV RNA and viral specific antigen and antibodies by a battery of commercial tests (Abbott and Roche) (Fig. 1 and Table 2). In particular, a total of 154 samples were tested including a median of 8 sequential specimens per acutely infected subject (range 5-11). The initial samples from acute subjects were HCV vRNA and antibody negative, followed by a sharp rise in vRNA levels. Five of 17 acutely infected subjects developed HCV antibodies by the last sampling time point. Chronic subjects had a median vRNA load of 1,975,569 IU/ml (range = 24,000 - 6,400,000 IU/ml) and all were HCV antibody positive. Viral load measurements were typical of acute and chronic HCV infection.
Viral RNA isolation and cDNA synthesis was performed as follows. Samples which contained approximately 100,000 viral RNA copies were extracted using the Qiagen BioRobot EZ1 Workstation with EZ1 Virus Mini Kit v2.0 (Qiagen, Valencia, CA). RNA was eluted in 60 μΐ and immediately subjected to cDNA synthesis. Reverse transcription of RNA to single stranded cDNA was performed using Superscript III reverse transcriptase (Invitrogen Life Technologies, Carlsbad, CA). For a 60 μΐ of reaction, a mixture of 3 μΐ (0.5 mM) of 10 mM deoxynucleoside triphosphate, 1.5 μΐ (0.25 μΜ) of 10 μΜ anti-sense primer and 34.5 μΐ of vRNA was incubated at 65°C for 5 min and then was placed on ice for at least 1 min. The contents of the tube were collected by brief centrifugation. Next, 12 μΐ of 5X reaction buffer, 3 μΐ of 0.1 M DTT (5 mM), 3 μΐ of RNAseOUT Recombinant RNASE Inhibitor, (40 units/μΐ, Invitrogen) and 3 μΐ of Superscript III™ Reverse Transcriptase (2001Ι/μ1) were added to the mixture. The final reaction was incubated at 50°C for 60 min followed by an increase in temperature to 55°C for an additional 60 minutes. The reaction was heat-inactivated at 70°C for 15 minutes and then treated with RNaseH at 37°C for 20 minutes. SEQ ID NOS The antisense primers were designed specifically for different genotype. 1.NS4A-R1 5'- GCACTCTTCC ATCTCATCGAACTC-3 ' (SEQ ID NO: 758) (nt 5451-5474 H77 (accession number NC 004102)) for genotype 1, 2NS2-R1 5'-CCCCAGACGATGACTTTCTTCTCCAT- 3' (SEQ ID NO: 777) (nt 5445-5467 H77) for genotype 2 and 3aNS3-R2V2 5'- TTACTTCC AGATCAGCTGAC A-3 ' (SEQ ID NO: 772) for genotype 3. The newly synthesized cDNA was used immediately or kept frozen at -80°C.
Single genome amplification was performed from prepared viral cDNA. cDNA was serially diluted and distributed among wells of replicate 96-well plates so as to identify a dilution where PCR positive wells constituted less than 30% of the total number of reactions. At this dilution, most wells contained amplicons derived from a single cDNA molecule. This was confirmed in every positive well by direct sequencing of the amplicon and inspection of the sequence for mixed bases (double peaks), which would be evidence of priming from more than one original template or the introduction of PCR error in early cycles. Any sequence with evidence of mixed bases was excluded from further analysis. PCR amplification was carried out in the presence of 10 x Taq High Fidelity Platinum PCR buffer, 2 mM MgS04, 0.2 mM of each deoxynucleoside triphosphate, 0.2 μΜ of each primer, and 0.1 μ(0.5 Unit) Platinum Taq High Fidelity polymerase in a 20 μΐ reaction (Invitrogen, Carlsbad, CA).
The nested or hemi-nested primers for generating 5' half and 5' quarter genome from different genotypes included: (1) genotype 1 : 1st round sense primer l .core.Fl 5'- ATGAGCACGAATCCTAAACCTCAAAGA-3' (SEQ ID NO:761)(nt 342-368 H77) and 1st round antisense primer 1.NS4A.R1 5'-GCACTCTTCCATCTCATCGAACTC-3' (SEQ ID NO:763) (nt 5451-5474 H77), 2nd round sense primer l .core.F2 5 ' -TC AAAG AAAAAC C AAA CGTAACACCAACCG-3' (SEQ ID NO:764) (nt 362-391 H77) and 2nd round antisense primer 1.NS3A4A.R2 5'-AGGTGCTCGTGACGACCTCCAGG-3' (SEQ ID NO:766) (nt 5297-5319 H77); (2) 5' quarter genome of genotype 2: 1st round sense primer 2.core.Fl 5'- ATGAGCACA AATCCTAAACCTCAAAGA-3' (SEQ ID NO: 767) (nt 342-368 H77) and 1st round antisense primer 2.NS2.R1 5'-CCCCACACAATGACCTTCTTCTCCATTG-3' (SEQ ID NO: 778) (nt 5445-5467 H77), 2nd round sense primer 2.core.F2 5'-AATCCTAAACCTCAAAGAAAAACC AAA-3' (SEQ ID NO: 769) (nt 351-377 H77) and 2nd round antisense primer 2.NS2.R2 5'-GG GGAGAGGTGGTCATAGATGTAA -3 '(SEQ ID NO 779); (3) 5' half genome of genotype 3 : 1st round sense primer 3a.core.Fl 5'-ATGAGCACACTTCCTAAACCTCAAAGA-3' (SEQ ID NO: 771) and 1st round antisense primer 3aNS3-R2V2 5 ' -TTACTTCCAGATCAGCTGACA- 3 '(SEQ ID NO: 760), 2nd round sense primer 3a.core.F2 5 ' -TC AAAG AAAAACC AAAAGAAA CACCATCCG-3' (SEQ ID NO: 773) and 2nd round antisense primer PCR 3a.NS3-R2V2 5'-TT ACTTCCAGATCAGCTGACA -3 '(SEQ ID NO. 774). Strain specific primers were used to amplify the quarter genomes from the early time point samples. The positive PCR reactions were subject to direct sequencing. PCR was performed in MicroAmp 96-well reaction plates (Applied Biosystems, Foster City, CA) with the following PCR parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94 °C for 15 s, an annealing step of 58°C for 30 s, an extension step of 68 °C for 5 min, followed by a final extension of 68 °C for 10 min. The product of the 1st round PCR was subsequently used as a template in the 2nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on precasted 1% agarose E- gels 96 (Invitrogen Life Technologies, Carlsbad, CA). All PCR procedures were carried out under PCR clean room conditions using procedural safeguards against sample contamination, including pre-aliquoting of all reagents, use of dedicated equipment, and physical separation of sample processing from pre- and post-PCR amplification steps.
For DNA sequencing, 5' half-genome amplicons were directly sequenced by cycle- sequencing using BigDye terminator chemistry and protocols recommended by the manufacturer (Applied Biosystems; Foster City, CA). Sequencing reaction products were analyzed with an ABI 3730x1 genetic analyzer (Applied Biosystems; Foster City, CA). Both DNA strands were sequenced using partially overlapping fragments. Individual sequence fragments for each amplicon were assembled and edited using the Sequencher program 4.7 (Gene Codes; Ann Arbor, MI). Inspection of individual chromatograms allowed for the identification of amplicons derived from single versus multiple templates. The absence of mixed bases at each nucleotide position throughout the entire 5 ' half-genome sequences was taken as evidence of single genome amplification from a single viral RNA/cDNA template. This quality control measure allowed the exclusion of amplicons that resulted from PCR-generated in vitro recombination events or Taq polymerase errors and to obtain multiple individual sequences that proportionately represented those circulating HCV virions. All the sequences alignments were initially made with ClustalW and then hand-checked using MacClade 4.08 to improve the alignments according to the codon translation.
Among the 3000 amplicons generated, the sequences of 2850 were unambiguous at every position. Sequences chromatograms from 150 amplicons contained one or more "double peaks" representing mixed bases. Because mixed bases generally represented only a minority of polymorphisms, it was inferred that such ambiguities had resulted from Taq polymerase errors in the initial PCR cycles and not from amplification from more than one initial target template; in such cases a correction of the assignment of the base was made. In rare instances where the only polymorphic site(s) present were represented by double peaks, possibly due to mixed initial templates, the sequences were discarded and not included in the analysis. This amounted to 30 sequences or less than 1% of total amplicons generated. Example 2
Phylogenetic Analysis: Identification and Enumeration of Transmitted Viruses
A total of 3000 Sequences corresponding to the core antigen genes from the initial time point from all 17 acutely-infected and 14 chronically-infected subjects were analyzed using neighbor-joining (NJ) phylogenetic tree methods together with a sequence visualization tool, Highlighter (www.HIV.lanl.gov). Phylogenetic trees were generated by the Maximum Likelihood (ML) method. For the subjects with multiple transmitted/founder viral lineages, the lineages that contained more than 5 identical or near identical sequences were included in the subsequent diversity analyses. The maximum sequence diversity within subject or lineage was calculated using Poisson Fitter program (www.HIV.lanl.gov).
A total of 38 transmitted/founder lineages from 17 acutely infected subjects were used to calculate the mutation rate. Each sequence within the lineage was compared with the T/F virus sequence of that lineage. The insertion, deletion, transition and transversion frequencies were counted by self-developed computer program and by Highlighter tool. The rates were calculated by taking the ratio of each frequency number and the total number of nucleotides of all the sequences within that lineage.
To determine the Synonymous and non-synonymous substitution rate, the SNAP program (www.HIV.lanl.gov) was applied to the codon-aligned sequences of each T/F lineage. Within each lineage, the accumulations of synonymous and nonsynonymous substitutions were counted by comparing to the transmitted/founder viral sequences. The Jukes-Cantor corrected accumulation rates of synonymous substitutions per potential synonymous site (ds) and nonsynonymous substitution per potential nonsynonymous site (dn) were compared to screen for positive selection.
To determine the likelihood of missing infrequent transmitted variants, a power study was performed to investigate the probability of sampling limitations. With a sample of at least n = 20 plasma vRNA sequences, a 95 % confidence level existed that a given missed variant comprised less than 15 % of the virus population. For samples for which n > 30, a 95 % confidence level existed that the results did not miss any variant that comprised at least 10 % of the total viral population.
Sequences from each subject formed monophyletic clades (bootstraps 99-100 %) corresponding to HCV genotypes la (n = 23), lb (n = 4), 2b (n = 2) and 3a (n = 2). In no case was a subject infected by more than one virus clade and there was no intermixing of sequences between subjects in the phylogenetic tree. Sequences from chronic subjects (blue shading) and acute subjects (red shading) showed variable degrees of within- subject diversity. Acute sequences were distinct, however, in exhibiting one or more discrete sublineages characterized by generally large numbers of sequences with extremely low diversity.
Sequences from chronic subject WIMI4025 showed broad genotypic heterogeneity with a maximum diversity of 3.97 % (median 1.24 %; range 0.12-3.97 %) (Fig 3A). Sequences from acute subject 10051 revealed a very different pattern of diversification (Fig 3B). These sequences, which were derived from the last sampled time point 21 days after the beginning of documented viremia, were very homogeneous with a maximum diversity of 0.18 % (median 0 %; range 0-0.18 %). The maximum diversities of sequence lineages from all 17 acutely infected subjects were similarly low, ranging from 0.04 % to 0.22 % (median = 0.1 1 %). This extremely limited diversity within sequence lineages from acutely infected subjects was significantly lower than the maximum viral diversity in chronic subjects (0.56 % to 3.97 %; median = 2.51 %; p<0.000X) (Tables 1 and 2). Nucleotide polymorphisms in subject 10051 were essentially random, corresponding to a star-like phylogeny and a Poisson distribution of low frequency events. Two sequences (2C3 and 2C2) contained a single common polymorphism at position -2200, and two others (2A2 and 2B34) contained a different common polymorphism at position -4375, indicating for each pair shared recent ancestry. Rare shared polymorphisms like these in newly infected subjects can be explained as having been generated subsequent to virus transmission as a consequence of RdRp errors very early in infection when the effective number (Ne) of productively infected cells is relatively low. The 60 sequences depicted in Fig. 3B coalesced to a single unambiguous consensus that was inferred, based on models of random virus evolution and empirical testing to represent the T/F virus in this subject. However, to obtain additional evidence that the sequences coalesced to a virus at or near the moment of transmission and not an intermediate time point, 243 additional sequences from three earlier time points were sampled. Whether sequences from each time point were considered separately or altogether, they coalesced to the same T/F genome (Table 1). Power calculations indicate that a sample size of 60 sequences as shown in Fig. 3B provide 95 % likelihood of detecting variants present at 5% in the population. A sample size of 300 sequences provides 95 % likelihood of detecting variants present at 1 % in the population. Thus, it was concluded that subject 10051 was productively infected by a single virus.
For subjects 10021 and 10025, at the initial time points, 19 of 24 (79 %) of 10021 sequences and 32 of 43 (74 %) of 10025 sequences were identical with the remainder differing from the respective consensus sequences by only 1 or 2 nucleotides. At the second and third time points 1 - 4 weeks later, an increasing proportion of sequences differed from the respective consensus sequences but only by 3 or 4 nucleotides. In each subject, sequences varied essentially randomly in a star-like fashion that fit a Poisson model of low frequency independent mutations. For each subject it was concluded that the consensus sequence corresponded to a single, unique T/F HCV genome. Altogether, of the 17 acute subjects in the present study, 4 had evidence of productive clinical infection by single viruses (Table 1). Figures 4C and 4D depict sequences from subjects having evidence of productive infection by more than one genetically distinct virus. Each of the subjects was sampled on 4 occasions and sequences again showed increasing diversity over time. The consensus sequences of each low diversity lineage in these subjects were interpreted to correspond to a unique T/F HCV genome. Thus, each of these subjects was productively infected by at least three viruses. The proportion of sequences represented in each lineage was approximately the same in subject 10012 (Fig 4C) but not in subject 10062 (Fig 4D). This could be due either to different replication rates of the different T/F viral viruses, differential effects of early innate or adaptive immune responses, different times of infection by different transmitted viruses, or early stochastic events in the infection process. A mathematical model was developed to estimate the adequacy of sampling given the observed distribution of T/F lineage sequences and total number of sequences analyzed. This model, when applied to the 179 sequences analyzed for subject 10012, estimated that the actual number of T/F viruses was 3 at a 95 % confidence level. In contrast, when applied to the 164 sequences analyzed for subject 10062, the model estimated that the actual number of T/F viruses could be as high as 5, again at a 95 % confidence level. Sequences from subject 10029 where low diversity sequence lineages corresponding to 9 T/F viruses could be unambiguously identified, but with even greater differences in sequence proportions. The model estimated that the actual number of T/F viruses could be as high as 13 at a 95 % confidence level. Comparable patterns of virus diversification following multivariant HCV transmission were observed in six additional subjects (Figs 6-12). Importantly, in contrast to HIV-1 where viral recombination in acute and early infection is exceedingly common and widespread, there was no evidence of viral recombination in any subject acutely infected by HCV. In addition, it was not observed in any of 13 subjects whose plasma virus was sequenced at multiple time points, the appearance of a distinct lineage of virus different from those sampled at the initial time points, indicating that there was no evidence of superinfection or substantial undersampling at the initial time points.
A mathematical approximation of early HCV evolution based on estimated parameters of RdRp error rate, virus generation time, infected cell lifespan, and viral reproductive ratio suggests that maximum viral diversity 10 weeks post-infection by a single virus is 0.4 % (95 % C.I. = 0.25 - 0.55 %), which is well within empirical determinations in these experiments (Table 1). With this as an upper bound, the diversity among T/F sequences from acutely infected subjects were evaluated to look for evidence of virus transmission from subjects who themselves were acutely infected. A similar approach has been successful in detecting acute-to-acute transmission of HIV- 1. If an acutely-infected subject acquired multiple HCV genomes from a contact who himself was acutely infected by a single virus, then these T/F viruses would be expected to all be closely related to each other, since they must reflect the viral diversity in the donor. If that donor contact had been acutely infected by more than one genetically divergent virus and the acutely infected recipient acquired multiple progeny of these, then the expectation would be for T/F viral sequences to be represented by subsets of sequences, some closely related and some not. Evidence in five subjects of acute-to-acute transmission was found by these two scenarios. Figures 13A-C depict ML trees and Highlighter plots from acutely infected subjects 10020, 10016 and 10003, where the maximum diversity of T/F viruses from a consensus was 0.18 %, 0.16 % and 0.12 %, respectively. The phylogenetic patterns of these sequences were quite different from the single variant transmissions shown in Figs 3B and 4 A and 4B, which exhibited star-like phylogeny and conformed to a Poisson distribution of random mutations. The sequences in 13A-C, instead, violated the Poisson model. They were comprised of distinct subsets of highly related sequences that differed from each other and from a common consensus by 1-6 nucleotides. If, however, the sequence subsets within each subject were considered independently, then their diversity conformed to the Poisson model. Thus, it was concluded that these three subjects were each productively infected by as many as 9-16 genetically distinct T/F viruses. The other scenario of acute -to-acute transmission is virus acquisition from a donor who is acutely infected by multiple viruses, where the expectation is for a multimodal distribution of diversity among T/F viruses in the recipient. This was observed in two subjects (Figs. 5 and 12).
The phylogenetic pattern of sequences from subject 106889 (Fig. 14) was far more complicated than that of any of the other 16 acutely infected subjects. Subject 106889 exhibited a typical acute infection viral kinetic profile with four sequential plasma samples negative for HCV vRNA followed by rapid vR A ramp-up to nearly 107 vRNA IU/ml (Fig 14 insert). Eighty-six 5 ' half genome sequences were obtained from the initial plasma vRNA positive time point. The recent evolutionary history of these sequences and the clinical circumstances of virus transmission could be inferred based on the ML tree, the Highlighter plot, and two additional lines of evidence: First, subject 106889 was a persistently HCV antibody-negative and HCV RNA-negative source plasma donor who had undergone regular biweekly plasmaphereses. Thus, when the subject became HCV viremic, it was clear that this was due to a new infection and not recrudescence of an old one. Second, 85 of the 86 sequences from subject 106889 depicted in Fig. 14 contained two signature mutations (V36M and R155R) in the NS3 protease that conferred high level resistance to the protease inhibitors Boceprevir and Telaprevir. A single sequence (02B11) contained only one of these mutations (V36M). It is clinically and biologically implausible for subject 106889 to have been diagnosed with acute HCV infection, to have been treated with an NS3 protease inhibitor (which was investigational at the time), and to have developed drug resistance, all in the span of one week. Instead, the transmitting partner to subject 106889 must have been chronically infected by HCV and treated with an NS3 protease inhibitor leading to narrowing of viral diversity and emergence of DAA drug resistance. Then, virus transmission to subject 106889 occurred, with the ML tree and Highlighter plot in Fig. 14 showing evidence of more than 30 T/F viruses. Mathematical modeling based on the Jackknife estimation tool provided a 95 % upper limit on the estimated number of T/F viruses in this subject at 48 (Table 3). The T/F viruses in subject 106889 are evident as discrete sublineages (designated as variants, v) within lineages designated Ll-10. Thus, in lineage 1 (LI), the progeny of as many as nine closely related T/F variants (vl-9) are evident. In lineage 2 (L2), the progeny of as many as six closely related T/F variants (vl-6) are evident (Fig. 14). This phylogenetic pattern of the progeny of multiple T/F variants within any one lineage is indistinguishable from T/F virus progeny in acutely infected subjects 10020, 10016 and 10003 (Figs. 13A-C). Lineage 5 (L5) in subject 106889 is comprised of sequences arising from a single T/F variant (vl), which is indistinguishable from examples of single variant transmission (subject 10051; Fig. 3B). Importantly, T/F lineages with sufficient numbers of sequences for model testing (e.g., L2-vl and L5-vl), conformed to a star-like phylogeny with a Poisson distribution mutations (Fig. 14; Table 1).
The identification of T/F viral sequences in 17 subjects provided a unique strategy for analyzing HCV sequence evolution in vivo in the critical period beginning at or near the moment of virus transmission and extending to the establishment of viral load setpoint and in some subjects HCV antibody seroconversion. A summary of this analysis is presented in Table 1. Maximum intra-lineage diversity for the 17 subjects ranged from 0.06 % to 0.22 %. Insertions (0.000001), deletions (0.00004) and stop codons (0.000004) were infrequent. Transitions outnumbered transversions by 8 to 1. The overall mutation frequency including all sampled time points was low (0.000145) given the number of possible virus replication cycles between the first infected hepatocyte and setpoint viremia 6-10 weeks later when as many as 7-20 % of hepatocytes may be productively infected. The dN/dS ratio was low at 0.33 as was its derivative pN/pS at 0.33 %.
The present study builds on a body of previous work aimed at characterizing the genetic bottleneck to HCV transmission and early and late patterns of virus diversification. Although previous studies demonstrated a transmission bottleneck, they could not identify with precision actual T/F viral genomes that were responsible for productive clinical infection, nor could they discriminate between T/F genomes that differed by as few as 1 to 3 nucleotides in 10,000 (0.01- 0.03% diversity). In addition, they could not evaluate virus diversity unencumbered by Taq polymerase-mediated nucleotide substitution errors or recombination artifacts that prevented mutational linkage to be analyzed across genes and genomes. The present study addressed these objectives by testing the hypothesis that SGA-sequencing could enable an unambiguous molecular identification and enumeration of T/F HCV genomes and their progeny. Although these goals had been previously reached for HIV-1 and simian immunodeficiency virus (SIV), it was unclear if a similar result could be achieved for HCV given the differences in viral replication strategies of the two viruses, the variable durations of the eclipse phases, differences in viral target cells (long-lived hepatocytes versus short-lived lymphocytes), the higher estimates of HCV RdRp error, and the possibility that higher multiplicities of infection by HCV in injection drug users could confound the identification of T/F viral lineages. Despite these challenges and uncertainties, the empirical findings of the present study make clear that even in complicated cases of multivariant HCV transmission in the setting of community-acquired infection, T/F viral genomes can be unambiguously and routinely identified.
With few exceptions, the diversity of sequences within the T/F lineages of each acutely infected subject corresponded to a star-like pattern of random diversification that conformed to a Poisson distribution mutations. Exceptions were of four types: (i) infrequent shared polymorphisms resulting from stochastic changes in the newly infected subjects (e.g., Fig 3B; 4A, 4B, 4C, 4D); (ii) transmission of multiple closely related variants (e.g., Figs 11, 13 A, 13B; 14); (iii) evidence of immune selection in later samples (Fig 12); and (iv) rare examples of short inverted repeats resulting from template switching between double-stranded RNA duplexed hairpin structures by the RNA dependent polymerase (Fig 13). The first three of these exceptions are also found in early HIV-1 and SIV diversification and can be explained by models of virus replication and diversification. The last exception, however, was unique to HCV. Among the 2670 amplicons sequenced, four sequences exhibited short concentrated stretches of 3 to 20 nucleotide substitutions compared with other sequences from the same subjects, including the T/F consensus sequences. In each of the four instances, the nucleotide mismatches represented perfect inverted repeats. This implies that double-stranded viral RNA served as the replication template for these sequences, although it cannot determine if they are the result of HCV RdRp-mediated replication in vivo or the product of cDNA synthesis from in vitro generated reverse transcription products of the MuLV reverse transcriptase. Finally, in any of the sequences analyzed no evidence of HCV recombination in vivo was observed, which would have been plainly evident in those subjects infected by multiple genetically diverse viral genomes (Fig 4C, D; 5; 6-11). Absence of viral recombination in vivo distinguishes HCV from HIV-1 and is consistent with epidemiological data showing that HCV recombination is rare.
Because early HCV sequences could be mapped precisely to T/F sequences, the genetic pathways of early HCV evolution were examined in a manner not previously possible, both at the level of nucleotide substitution and amino acid selection. For the proteome, one subject (9055) had evidence of a non-random accumulation of changes in a previously identified HLA- restricted cytotoxic T-cell (CTL) epitope. Since the HLA type of this subject was not known, it could not verified if this peptide sequence represented a CTL epitope, but it is likely based on parallel observation in HIV-1 that this region of HCV represents CTL escape or reversion. CTL responses and virus escape or reversion are generally believed to occur later in infection but early mutations could have been overlooked if sequences were not mapped to T/F proteomes. Again, this was a scenario observed in acute HIV-1 infection where identification of T/F genomes and proteomes enabled the detection of earliest CTL recognition and escape. Non- random amino acid polymorphisms in multiple T/F sequences in subjects 10020 (Fig 13 A) and 10016 (Fig 13B) also corresponded to known or predicted CTL epitopes and thus are likely to represent CTL escape or reversion in the respective donors (not in subjects 10020 or 10016). At the genome level, it was determined the frequency of infections, deletions, and termination codons were low (between 0.000001-0.000004) and the dN/dS ratio to be similarly low at 0.33, consistent with purifying or negative selective early in infection. Interestingly, the frequency of transitions to exceeded that of transversions by a factor of 8: 1, again consistent with a bias toward negative selection. Gotte et al have recently provided a biochemical explanation for this observation based on a strong bias for G:U/U:G mismatches by the HCV RdRp. The frequency of nucleotide substitutions in evolved sequences compared with T/F sequences was overall 0.000145. This value is derived from all sequences from all sampling time points and thus substantially overestimates the error rate per replication cycle of the HCV RdRp. In addition, it includes base substitutions resulting from the MuLV RT reverse transcription of HCV RNA.
An unexpected finding of the present study was evidence of acute-to-acute HCV transmission in a high proportion of subjects (5 of 17). A limitation of the current study is that virus from donor-recipient HCV transmission partners was not available. However, in cases of acute -to-acute HIV-1 transmission when this has been done, the phylogenetic patterns of closely related sequences are virtually indistinguishable from those found here for HCV (Fig 5; 11, 12, 13 A, 13B). Transmission of HIV-1 during the acute infection period is believed to contribute substantially to the epidemic in high prevalence and high incidence regions including Africa. Enhanced HCV-1 transmission for acutely infected subjects is believed to result from high virus loads, each of circulating neutralizing antibodies and restricted viral genetic diversity, an interpretation corroborated by acute infection studies of SIV in Indian rhesus macaques. All of these same factors could pertain to acute -to-acute HCV transmission.
The complicated HCV transmission scenario for subject 106889 (Fig 14) illustrates the sensitivity and specificity with which the SGA-direct amplicon sequencing strategy can detect and discriminate T/F viral genomes and their progeny. SGA-direct sequencing includes in vitro recombination artifacts, avoids Taq polymerase-mediated nucleotide substitution errors in finished sequences, precludes founder effects due to disproportionate target amplification, and avoids cloning bias. This allows for a proportional representation of target molecules in finished sequences. It also allows for genetic linkages to be preserved in finished sequences specific mutations, in this case NS3 DAA resistance mutations V36M and R155K to be precisely identified. In subject 106889, each T/F virus contained the V36M/R155K double mutation indicating high level NS3 protease resistance. Subject 106889 as well as the transmitting partner, SGA-direct sequencing is ideally suited to identifying genetically-linked DAA resistant mutants within single or multiple genes (e.g., NS3, NS5A and NS5B) of transmitted viruses and characterizing their persistence or disappearance over time.
Finally, this study demonstrates that that SGA-direct amplicon sequence allows for a precise and unambiguous molecular identification of the complete nucleotide sequences of T/F HCV genomes that are responsible for productive HCV infection of humans. These sequences, when synthesized, molecularly cloned, and expressed eukaryotically, recapitulate all of the features of naturally-occurring T/F HCV viruses. Because their nucleotide sequences match exactly those of viruses responsible for transmission and productive clinical infection of humans, the sequences of T/F viral genomes by definition contain all of the genetic elements necessary for a pathogenic infection of humans. These molecular HCV genomes should prove to be a valuable resource for in vitro and in vivo testing of therapeutic agents, drugs, and vaccines.
Example 3
Molecular Identification, Chemical Synthesis and Molecular Cloning of a Full-Length, Replication-Competent, Infectious HCV Genome
The transmitted/founder virus sequence of the full length genome of subject 10021 was obtained by analysis of five overlapped genome fragments: 5' UTR (nt 1 to 852), 5' half genome (core, El, E2, p7, NS2 andNS3, nt 391 to 5297), 3' half genome (NS4A, NS4B, NS5A andNS5B, nt 5168 to 9374), 3' U/C tract (nt 9082 to 9582) and 3' x-tail (nt 9563 to 9646). The primers designed for cDNA synthesis and PCR amplification are listed below. The nucleotide numbering system is based on reference sequence H77 (accession number NC 004102). The sequences of all five genome regions were analyzed and assembled. The transmitted/founder sequence was inferred based on phylogenetic inference and mathematical modeling.
For viral RNA extraction, the plasma sample contained approximately 100,000 viral RNA copies and was extracted using the Qiagen BioRobot EZ1 Workstation with EZ1 Vrius Mini Kit v2.0 (Qiagen, Valencia, CA). RNA was eluted and immediately subjected to cDNA synthesis or stored at -80°C. The cDNA synthesis and single genome amplification of '5 and '3 half genome was performed as described supra in Example 2. The positive PCR reactions were subject to direct sequencing.
The cDNA synthesis of 3' U/C tract of subject 10021 was also performed using Superscript III™ Reverse Transcriptase. For a 20 μΐ of reaction, the mixture of 1 μΐ (0.5 mM) of 10 mM deoxynucleoside triphosphate, 0.5 μΐ (0.25 μΜ) of 10 μΜ anti-sense primer 3UTR-R10 and 11.50 μΐ of vRNA was heated at 65 °C for 5 min and then was placed on ice for at least 1 min. Then 4 μΐ of 5X reaction buffer, Ιμΐ of 0.1 M DTT (5 mM final), 1 μΐ of RNAseOUT Recombinant RNASE Inhibitor and 1 μΐ of Superscript III™ Reverse Transcriptase (2001Ι/μ1) were added. The final reaction mix was incubated at 50 °C for 60 min followed by a 5 min heat inactivation at 85°C. PCR amplification was carried out in the presence of 2 μΐ of 10 x Taq High Fidelity Platinum PCR buffer, 0.8 μΐ of 50 mM MgS04 (2 mM), 0.4 μΐ of 10 mM deoxynucleoside triphosphate (0.2 mM), 0.2 μΐ of each primer (0.2 μΜ), and 0.1 μΐ (0.5 Unit) Platinum Taq High Fidelity polymerase in a 20 μΐ reaction. The semi-nested PCR primers used are listed in Table 4. PCR was performed in MicroAmp 96-well reaction plates with the following PCR parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94°C for 15 s, an annealing step of 50 °C for 30 s, an extension step of 68°C for 1.5 min, followed by a final extension of 68 °C for 10 min. The product of the 1st round PCR was subsequently used as a template in the 2nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on precasted 1 % agarose E-gels 96 and 4 % Nusieve GTG Agarose gel (CAMBREX, Cat. No. 50080). The positive PCR reactions were TA cloned into pGEM-T Easy Vector System (Promega, Cat. No. A3610). The inserts were identified and sequenced using Ml 3 forward/reverse primers.
The transmitted/founder virus sequence of 5' UTR of subject 10021 was obtained using FirstChoice RLM-RACE kit (Ambion, Cat. No. AM 1700), with slight modification of the methods recommended by the manufacturer. First, vRNA was treated with Tobacco Acid Pyrophosphatase (TAP) to remove two 5' P04 from 5' -triphosphate of the full-length vRNA, leaving one 5'-monophosphate. Then a 45 base RNA Adapter oligonucleotide was iigated to the R A population using T4 RNA ligase. During the ligation reaction, the full-length vRNA acquires the adapter sequence as its 5' end, A core gene specific RT-PCR amplified the complete 5' UTR genome. The cDNA synthesis and PCR amplification were performed under same conditions as that described for 3' U/C tract except using 55 °C as the annealing temperature. The positive PCR products were sequenced directly.
To determine the complete x tail sequences, a synthetic RNA oligonucleotide adaptor (Integrated DMA Technologies INC.) wras iigated to the 3' end of the vRNA by using T4 RNA ligase. The adaptor was 5' phosphorylated and 3' Dideoxy-C blocked to prevent intra-molecular ligation. The 20 ul ligation reaction that contained 15 μΐ of vRNA, 2 μΐ of 10X ligation buffer, 1 μΐ of 10 rnM ATP, 0.5 μΐ of RNase Inhibitor (20υ/μ1), 1 μΐ of 10 μΜ oligo adaptor ( 0.5 μΜ) and 1 μΐ of T4 RNA ligase (10 units) was incubated at 37 °C for 60 min. The whole ligation reaction was subject to cDNA synthesis directly using the oligo adaptor specific primer. The cDNA synthesis was carried at 50 °C for 60 min. PCR was performed with the following parameters: 1 cycle of 94 °C for 2 min; 35 cycles of a denaturing step of 94 °C for 15 s, an annealing step of 55 °C for 30 s, an extension step of 68 °C for 30 sec, followed by a final extension of 68 °C for 10 min. The product of the 1st round PCR was subsequently used as a template in the 2nd round PCR under same conditions but with a total of 45 cycles. Amplicons were inspected on 4 % Nusieve GTG Agarose gel (CAMBREX). The positive PCR reactions were directly TA cloned into pGEM-T Easy Vector System (Promega, Cat. No. A3610). The inserts were identified and sequenced using Ml 3 forward/reverse primers.
The inferred full-length transmitted/founder sequences was chemically synthesized in four subgenomic fragments (Blue Heron Biotechnology). The fragments were overlapped at the unique restriction sites, BsrGI (nt 3640), SnaBI (nt 6644) and Sfil (nt 9404), respectively. The noncutter Notl was attached to the 5' terminal of the first fragment and noncutter Xbal was attached to the 3' terminal of all the fragments during chemical synthesis. The synthesized fragments were propagated, restriction enzyme digested and sequentially cloned into the MCS Notl-Xbal site of pBlueScriptKS(+). The final full-length clone was confirmed by sequence analysis. Table 1. Diversity and mutation analyses of HCV sequences in acute infection.
Table 1. Continued
a 5h-5' half genome contains core, El, E2, p7, NS2, and NS3; 5ql-5' quarter 1 genome contains core, El, E2, p7 and partial NS2; 5q2-5' quarter 2 genome contains partial NS2 and NS3.
b rate calculations were derived from sequences from each discrete transmitted/founder lineage.
c sequences from 1st sampled time point with >4 sequences per lineage were analyzed. If 5' half genomes were not available, quarter genomes
1 and 2 were analyzed with each result shown.
d sequences from 2nd sampled time point were analyzed due to insufficient nos. of sequences or sequence diversity from 1st sampled time point.
e Insufficient numbers of sequences from each time point to calculate fit to Poisson or star-like phylogeny.
Transmitted/founder lineage identified by this sequence in respective ML tree and Highlighter plot.
g N/A, not applicable due to multiple transmitted/founder virus genomes.
h Averages were calculated from total mutations in all transmitted/founder lineages from all subjects combined. Because of low numbers of sequences and mutations in some lineages, certain values (e.g. dN/dS for subjects 10062 v2 and 10004 v3) vary substantially from the mean.
Table 2. Diversity analysis of 5' half HCV genome sequences from 14 chronically infected
*HCV infection only
Table 3. Jackknife estimation of number of T/F variants in acute infection
a obtained by applying jackknife to the biasd estimator, in this case is the clusters found.
b obtained by bootstrapping the fixed cluster membership.
c power calculation of the prevalence of the unseen variants based on the number of sequences obtained. Table 4. Estimates of numbers of T/F viruses in acute HCV infection using empirical and model based methods
a manual estimates were based on phylogenies and Highlighter plots of all time points combined.
b power calculation estimating an upper bound on the prevalence of unseen variants given the number of sequences analyzed. This estimate is based on the total number of sequences from all time points.
Table 5. Estimation of numbers of T/F viruses by time point using model-based methods with different cut-offs
Days Avg.
Maximum cut-off1'
Sampling No. of Seq. since cut-off"
Subject Genome
time pt. seq. length negative Total mutation Model- Model- sample cut-off based based
7 5* half 44 4992 6 5 1 1
9055 9 5* half 78 4993 15 5 1 1
11 5* half 35 4993 37 9 1 1
10021 8 5' quarter 1 31 2172 7 3 1 1 Days Avg.
Maximum cut-off1'
Sampling No. of Seq. since cut-off"
Subject Genome
time pt. seq. length negative Total mutation Model- Model- sample cut-off based based
8 5' quarter 2 24 2773 7 3 1 1
10 5* half 30 4960 14 5 1 1
14 5* half 66 4960 35 9 1 1
8 5' quarter 1 40 2318 3 3 1 1
8 5' quarter 2 43 2645 3 3 1 1
10025
9 5* half 38 4964 15 5 1 1
11 5* half 54 4964 31 8 1 1
9 5' quarter 1 46 2212 8 3 1 1
9 5' quarter 2 45 2733 8 3 1 1
10 5' quarter 1 54 2212 13 3 1 1
10051
10 5' quarter 2 33 2733 13 3 1
11 5* half 65 4948 15 5 1 1
14 5* half 60 4948 27 7 1 1
7 5* half 120 4995 7 5 19 9
10003 9 5* half 7 4995 14 5
12 5* half 6 4995 26 7
10 5' quarter 53 2851 4 3 9 3
10016
12 5' quarter 19 2851 28 4 5 3
6 5' quarter 58 2194 9 3 5 2
10020 8 5* half 24 4987 16 5 6 1
13 5* half 40 4987 41 10 1 1
6213 10 5* half 41 4905 29 7 3 3
6222 8 5* half 17 4985 38 10 4 4
4 5* half 5 4987 11 5
10002
7 5* half 26 4987 24 6 13 11
10004c 6 5' quarter 36 2849 12 3 3 3
6 5' quarter 1 52 2367 7 3 3 3
6 5' quarter 2 49 2596 7 3 3 3
10012c 8 5* half 36 4963 14 5 3 3
10 5* half 49 4963 21 5 3 3
13 5* half 44 4964 33 8 3 3
9 5' quarter 1 35 2288 11 3 3 3
9 5' quarter 2 22 2681 11 3 2 1
10 5* half 63 4970 16 5 3 3
10017
12 5* half 48 4969 24 6 2 2
14 5* half 54 4970 31 8 2 2
16 5* half 27 4970 42 11 1 1
6 5' quarter 1 40 2318 16 3 4 4
6 5' quarter 2 70 2700 16 3 3 3
10024
7 5* half 63 4964 20 5 3 3
8 5* half 49 4964 22 6 3 3
8 5' quarter 1 68 2358 6 3 7 7
10029 8 5' quarter 2 53 2599 6 3 7 6
9 5* half 68 4957 13 5 6 7 Days Avg.
Maximum cut-off1'
Sampling No. of Seq. since cut-off"
Subject Genome
time pt. seq. length negative Total mutation Model- Model- sample cut-off based based
11 5* half 75 4957 20 5 7 7
15 5* half 58 4957 34 9 6 5
3 5' quarter 1 24 2425 4 3 2 2
3 5' quarter 2 24 2514 4 3 2 2
10062 4 5* half 54 4938 7 5 3 3
5 5* half 59 4938 11 5 3 3
8 5* half 27 4938 42 11 2 2
106889 5 5* half 87 4984 11 5 28 16 a total mutation cut-off was based on the time of sampling relative to the last negative time point and the average diversity in a cluster and was used to define distinct founders.
b two methods were used to implement the automated clustering algorithm. The average cut-off method is the more conservative estimate of the minimum number of founder strains needed to explain the observed diversity. The maximum cut-off distinguishes more lineages and separates them into clusters from distinct founders.
c the model assumes absence of homoplasy. In these subjects infected by multiple divergent viruses, one mutation at one site could be explained as a homoplasy and hence the
corresponding column was removed from the alignment.
Table 6. Oligonucleotides for ligation, cDNA synthesis and PCR amplification of full genome
Genome Name Position Sense and usage
1.NS4A.R1 (SEQ ID NO: 758) 5451-5474 -, RT and 1st round l .core.Fl (SEQ ID NO: 761) 342-368 +, 1st round
5' half
1.NS3.4A.R2 (SEQ ID NO: 780) 5297-5319 -, 2nd round l .core.F2 (SEQ ID NO: 781) 362-391 +, 2nd round la.NS5Bend.Rla (SEQ ID NO: 782) 9374-9398 -, RT and 1st round
10021.NS3.F1 (SEQ ID NO: 783) 5039-5067 +, 1st round
3' half
la.NS5B.R2 (SEQ ID NO: 784) 9315-9341 -, 2nd round
10021. NS3.F2 (SEQ ID NO: 785) 5146-5168 +, 2nd round
5* RACE Adaptor (SEQ ID NO: 786)
l .Core.Rl (SEQ ID NO: 787) 852-874 -, RT and 1st round
5' UTR 5* RACE Outer (SEQ ID NO: 788) +, 1st round
l .Core.R2 (SEQ ID NO: 789) 822-844 -, 2nd round
5* RACE Inner (SEQ ID NO: 790) +, 2nd round
3UTR-R10 (SEQ ID NO: 791) 9582-9604 -, RT and 1st round
3' U/C laNS5B-F1.2 (SEQ ID NO: 792) 9028-9055 +, 1st round tract 3UTR-R10 (SEQ ID NO: 793) 9582-9604 -, 2nd round
laNS5B-F2 (SEQ ID NO: 794) 9060-9082 +, 2nd round
HL-3* adaptor (SEQ ID NO: 795)
3*-termi-Rl (SEQ ID NO: 796) -, RT and 1st round
3' xtail
10021 3*raceFl (SEQ ID NO: 797) 9537-9559 +, 1st round
3*-termi-R2 (SEQ ID NO: 798) -, 2nd round Genome Name Position Sense and usage
10021 3*raceF2 (SEQ ID NO: 799) 9544-9563 +, 2nd round
The examples given above are merely illustrative and are not meant to be an exhaustive list of all possible embodiments, applications or modifications of the invention. Thus, various modifications and variations of the described methods and systems of the invention will be apparent to those skilled in the art without departing from the scope and spirit of the invention. Although the invention has been described in connection with specific embodiments, it should be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention which are obvious to those skilled in molecular biology, immunology, chemistry, biochemistry or in the relevant fields are intended to be within the scope of the appended claims.

Claims

We claim:
1. A transmitted full-length HCV genome comprising the polynucleotide of SEQ ID NO: 776.
2. The HCV genome of claim 1, wherein the genome mediates viral transmission.
3. The HCV genome of claim 1, wherein the polynucleotide sequence comprises env and core genes.
4. An immunogenic composition comprising transmitted full-length HCV genome or portions thereof.
5. The immunogenic composition of claim 4, wherein the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
6. A method of administering a vaccine, the method comprising treating a patient with transmitted full-length HCV genome or portions thereof, wherein a patient immune response is induced.
7. The method of claim 6, wherein the HCV genome comprises the polynucleotide of SEQ ID NO: 776.
8. A method for identifying full-length transmitted HCV genome identified by the method comprising:
(a) collecting a patient sample;
(b) isolating viral RNA from said sample;
(c) sequencing said viral RNA, wherein viral RNA sequencing includes HCV genomes of circulating virus;
(d) performing sequence alignment of selected HCV genome regions;
(e) analyzing phylogenetically selected sequence alignments; and
(f) identifying full-length HCV genomes of transmitted virus.
9. The HCV genome as identified by the method of claim 8, wherein the polynucleotide sequence comprises SEQ ID NO. 776.
EP12854779.1A 2011-12-05 2012-12-04 Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification Withdrawn EP2788490A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US201161630161P 2011-12-05 2011-12-05
PCT/US2012/067731 WO2013085889A1 (en) 2011-12-05 2012-12-04 Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification

Publications (1)

Publication Number Publication Date
EP2788490A1 true EP2788490A1 (en) 2014-10-15

Family

ID=48574801

Family Applications (1)

Application Number Title Priority Date Filing Date
EP12854779.1A Withdrawn EP2788490A1 (en) 2011-12-05 2012-12-04 Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification

Country Status (2)

Country Link
EP (1) EP2788490A1 (en)
WO (1) WO2013085889A1 (en)

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
ATE526423T1 (en) * 2000-04-18 2011-10-15 Virco Bvba METHOD FOR MEASURING DRUG RESISTANCE TO HCV
US7659103B2 (en) * 2004-02-20 2010-02-09 Tokyo Metropolitan Organization For Medical Research Nucleic acid construct containing fulllength genome of human hepatitis C virus, recombinant fulllength virus genome-replicating cells having the nucleic acid construct transferred thereinto and method of producing hepatitis C virus Particle
US20140023683A1 (en) * 2010-09-08 2014-01-23 The Uab Research Foundation Identification of transmitted hepatitis c virus (hcv) genomes by single genome amplification

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
See references of WO2013085889A1 *

Also Published As

Publication number Publication date
WO2013085889A8 (en) 2013-11-28
WO2013085889A1 (en) 2013-06-13

Similar Documents

Publication Publication Date Title
Cecchinato et al. Avian metapneumovirus (AMPV) attachment protein involvement in probable virus evolution concurrent with mass live vaccine introduction
US7320857B2 (en) Characterization of the earliest stages of the severe acute respiratory syndrome (SARS) virus and uses thereof
WO2006086188A2 (en) Use of consensus sequence as vaccine antigen to enhance recognition of virulent viral variants
Costa-Mattioli et al. Evidence of recombination in natural populations of hepatitis A virus
Wu et al. Recombination of hepatitis D virus RNA sequences and its implications.
Leary et al. Three adjacent nucleotide changes spanning two residues in SARS-CoV-2 nucleoprotein: possible homologous recombination from the transcription-regulating sequence
AU2003218111A1 (en) Methods and compositions for identifying and characterizing hepatitis c
Han et al. Comparison of multiplex restriction fragment mass polymorphism and sequencing analyses for detecting entecavir resistance in chronic hepatitis B
Franco et al. Complete nucleotide sequence of genotype 4 hepatitis C viruses isolated from patients co-infected with human immunodeficiency virus type 1
US20140023683A1 (en) Identification of transmitted hepatitis c virus (hcv) genomes by single genome amplification
Peres-da-Silva et al. Genetic diversity of NS3 protease from Brazilian HCV isolates and possible implications for therapy with direct-acting antiviral drugs
Li et al. Molecular epidemiology of hepatitis C genotype 6a from patients with chronic hepatitis C from Hong Kong
Bracho et al. Complete genome of a European hepatitis C virus subtype 1g isolate: phylogenetic and genetic analyses
WO2020184730A1 (en) Dengue virus vaccine
EP2788490A1 (en) Identification and cloning of transmitted hepatitic c virus (hcv) genomes by single genome amplification
Harasawa et al. Evidence for Pestivirus Infection in Free‐Living Japanese Serows, Capricornis crispus
González-Horta et al. Analysis of hepatitis C virus core encoding sequences in chronically infected patients reveals mutability, predominance, genetic history and potential impact on therapy of Cuban genotype 1b isolates
KR102297300B1 (en) A live virus banked from an attenuated dengue virus strain, and a dengue vaccine using them as an antigen
JP2017500886A (en) HCV genotyping algorithm
Dinu et al. Screening of protease inhibitors resistance mutations in hepatitis c virus isolates infecting romanian patients unexposed to triple therapy
Mello et al. Conservation of hepatitis C virus nonstructural protein 3 amino acid sequence in viral isolates during liver transplantation
Larsson Hepatitis B virus replication and integration
US10815523B2 (en) Indexing based deep DNA sequencing to identify rare sequences
Ndjomou et al. Functional domains of the human immunodeficiency virus type 1 Nef protein are conserved among different clades in Cameroon
US20140271726A1 (en) Compositions and methods for predicting hcv susceptibility to antiviral agents

Legal Events

Date Code Title Description
PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

17P Request for examination filed

Effective date: 20140707

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC MK MT NL NO PL PT RO RS SE SI SK SM TR

DAX Request for extension of the european patent (deleted)
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE APPLICATION IS DEEMED TO BE WITHDRAWN

18D Application deemed to be withdrawn

Effective date: 20150701