EP4689159A1 - Method and kits - Google Patents
Method and kitsInfo
- Publication number
- EP4689159A1 EP4689159A1 EP24718330.4A EP24718330A EP4689159A1 EP 4689159 A1 EP4689159 A1 EP 4689159A1 EP 24718330 A EP24718330 A EP 24718330A EP 4689159 A1 EP4689159 A1 EP 4689159A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- cells
- adaptors
- barcode
- samples
- polynucleotide
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6806—Preparing nucleic acids for analysis, e.g. for polymerase chain reaction [PCR] assay
Definitions
- the invention relates to a method of uniquely labelling RNA molecules in a population of cells and kits for use in such methods.
- the methods and kits of the invention allow the study of the transcriptome in individual cells or populations of cells.
- the transcriptome is the set of all coding and non-coding RNA transcripts in a single cell or a population of cells. Data concerning the transcriptome can be used to study, amongst others, cellular differentiation, carcinogenesis, transcription regulation, and biomarker discovery. Transcriptome data can also be used to identify the source of the cell or cells, including its/their phylogeny.
- RNA-Seq uses NGS to measure the presence and amount of RNA molecules in cells and is capable of analysing the continuously changing cellular transcriptome.
- WO 2019/060771 describes methods of uniquely labelling or barcoding RNA molecules within each cell of a population of cells. This technique is known as "Split-Seq" and involves multiple rounds of separating (or splitting) the cells into separate samples, labelling the cells in each sample with different barcodes and then pooling the samples. However, all of the separating/splitting and pooling steps are conducted after reverse transcribing the RNA molecules in the population of cells into cDNA. The same method is described in Rosenberg et al., Science, 360 (6385): 176-182. This method may described as the RLL (RT + ligation + ligation) approach.
- Biological pores have great potential as direct, electrical biosensors for polymers and a variety of small molecules.
- nanopores have great potential DNA sequencing technology.
- an analyte such as a nucleotide
- Nanopore detection of the nucleotide gives a current change of known signature and duration.
- Strand sequencing can involve the use of a molecular brake to control the movement of the polynucleotide through the pore.
- the present inventors have identified a novel method of uniquely labelling RNA molecules in populations of cells.
- This approach involves multiple rounds of dividing (or splitting) the population of cells into separate samples, labelling the cells in each sample with different barcodes and then pooling the samples.
- the first round of dividing (or splitting) the cells, labelling and pooling is conducted before reverse transcription (RT).
- the second dividing (or splitting step) is also conducted before RT.
- RT reverse transcription
- the novel method steps differentiate the novel method of the invention from Split-Seq disclosed in WO 2019/060771 and Rosenberg et al., Science, 360 (6385): 176-182.
- the novel method may be described as the LRL (ligation + RT + ligation) approach.
- the novel method may be summarised as follows. After the first division of the population of cells, the RNA molecules in the first divided samples are labelled by first adaptors comprising primer sites for RT and the reverse complement of unique first barcodes for each first sample. The first divided samples are then pooled and divided for a second time before RT using RT primers comprising unique second barcodes for each second sample. The second divided samples and then pooled and divided a third time before labelling with second adaptors comprising unique third barcodes for each third sample. This provides triple barcoded constructs from all cells. The triple barcoded constructs comprise DNA sequences transcribed from the RNA molecules (during the second round).
- the barcode used for each sample i.e., for each first, second and third sample, is preferably different from any other barcode or all the other barcodes used in the method.
- the number of cells in the original population is preferably less than the number of first samples multiplied by the number of second samples multiplied the number of third samples. If the number of possible barcode combinations used is greater than the number of cells in the population, there is a high probability the RNA molecules from each cell will be labelled with a unique combination of three barcodes (/.e., the barcoded constructs produced from each cell will have a unique combination of three barcodes).
- the triple barcoded constructs may then be isolated, amplified, and characterised or sequenced. This allows the transcriptome of each individual cell in the population or multiple cells in the population to be measured and/or quantified.
- the method may be used in combination with nanopore sequencing but does not have to be.
- the method of the invention has several key advantages. For instance, by separating the first-round labelling and third-round labelling steps with a second-round RT step, the possibility of re-barcoding or crosstalk upon pooling is mitigated. This a common problem associated with RLL Split-Seq as disclosed in WO 2019/060771 and Rosenberg et al., Science, 360 (6385): 176-182 where two sequential rounds of ligation are used. Crosstalk in this instance would mean that fragments failed to be RT primed or adapter ligated in their single wells, but upon pooling are presented with surplus primer/adapter and reagents that permit further labelling of the fragments with barcodes that were not present for the previous round of labelling.
- RT steps it is also technically possible for an RT primer to prime internally within a transcript and then allow reverse transcription to initiate and strand displace previous RT products or create truncated RT products. This would be an example of re-barcoding, where the identity of a barcode that has been incorporated could be over-written with a new one. Again this would lead to transcripts presenting as if they originated from different cells when the combinatorial barcodes are considered.
- a first adaptor to the 3' ends of RNA transcripts in the first round enables RT of the full RNA transcript with a specific RT primer in the second round without the need of a poly(T)VN reverse transcription primer which is prone to off-target priming. This minimises intronic overlap and off-target amplification. It also allows for cDNA synthesis of full-length poly(A) tails in eukaryotic RNAs for end-to-end RT of RNA transcripts.
- Part of the first adaptor can be destroyed prior to cell pooling at the end of the first round and this mitigates the possibility of re-barcoding or crosstalk upon pooling.
- the invention provides a method of uniquely labelling RNA molecules in a population of cells, the method comprising:
- first adaptors to the 3' ends of the RNA molecules in the plurality of first samples, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode;
- the invention also provides a method of characterising the transcriptome of an individual cell or two or more cells in a population of cells, the method comprising:
- the invention also provides a kit for uniquely labelling RNA molecules in a population of cells, the kit comprising two or more first adaptors, wherein the two or more first adaptors comprise primer sites for RT, and wherein each of the two or more first adaptors comprises the reverse complement of a different first barcode.
- Figure 1 shows LRL first-round barcoding by ligation of the first adaptor.
- Figure 2 shows LRL second-round barcoding through RT.
- Figure 3 shows LRL third-round barcoding through ligation.
- Figure 4 shows LRL third-round adaptor blocking.
- Figure 5 shows how an SSC based rehydration buffer yields total RNA of a much higher integrity than a DPBS based rehydration buffer.
- the RIN (RNA integrity number) score - a measure of the ratio 28S to 18S rRNA peaks - for the SSC samples is 2-fold higher than those of the DPBS samples.
- RNA integrity throughout the cell fixation is crucial as the RNA transcripts must be retained for cDNA synthesis as part of the combinatorial barcoding method. This experiment demonstrates that the methanol fixation and rehydration does not impact the RNA integrity too greatly.
- Figure 6 shows how RNA integrity remains unchanged with incubation at 37°C compared to a non-incubated control, both with comparable RIN scores.
- RNA integrity throughout the CDRA ligation and USER digest is crucial as the RNA transcripts must be retained for cDNA synthesis as part of the combinatorial barcoding method.
- Figure 7 shows first round barcoding competition assay results. Sequenced reads aligned to BC01 or BC02 of the CDRA adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
- Figure 8 shows second round barcoding competition assay results. Sequenced reads aligned to BC03 or BC04 of the RT barcode primer against the primary Y axis. PCR yield is aligned against the secondary Y axis.
- Figure 9 shows third round barcoding competition assay results where no blocking oligo was utilised. Sequenced reads aligned to BC05 or BC06 of the third-round ligation adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
- Figure 10 shows third round barcoding competition assay results where the RLL approach standard blocking oligo was utilised in the LRL approach. Sequenced reads aligned to BC05 or BC06 of the third-round ligation adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
- Figure 11 shows third round barcoding competition assay results where the LRL approach alternative blocking oligo was utilised. Sequenced reads aligned to BC05 or BC06 of the third-round ligation adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
- Figure 12 shows third round barcoding competition assay results. Comparison of nonbarcoded primary third round ligations competed with BC06 with either no blocking oligo, RLL approach standard blocking oligo (std), or the LRL approach alternative blocking oligo (alt) was utilised. Sequenced reads aligned to BC05 or BC06 of the third-round ligation adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
- SEQ ID NOs: 1-18 are the oligonucleotides used in the Examples.
- “About” as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ⁇ 20 % or ⁇ 10 %, more preferably ⁇ 5 %, even more preferably ⁇ 1 %, and still more preferably ⁇ 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods.
- Nucleotide sequence refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA.
- nucleic acid as used herein, is a single or double stranded covalently linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds.
- the polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases.
- Nucleic acids may be manufactured synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-translational modification, for example 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and splicing.
- Nucleic acids may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA).
- Sizes of nucleic acids also referred to herein as "polynucleotides” are typically expressed as the number of base pairs (bp) or nucleotide pairs for double stranded polynucleotides, or in the case of single stranded polynucleotides as the number of nucleotides (nt).
- oligonucleotides typically called “oligonucleotides” and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR).
- PCR polymerase chain reaction
- amino acid in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH 2 ) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid.
- the amino acids typically refer to naturally occurring L o-amino acids or residues.
- amino acid further includes D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as P-amino acids.
- amino acid analogues naturally occurring amino acids that are not usually incorporated into proteins such as norleucine
- chemically synthesised compounds having properties known in the art to be characteristic of an amino acid such as P-amino acids.
- analogues or mimetics of phenylalanine or proline which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid.
- Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid.
- polypeptide and “peptide” are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers.
- Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like.
- a peptide can be made using recombinant techniques, e.g., through the expression of a recombinant or synthetic polynucleotide.
- a recombinantly produced peptide it typically substantially free of culture medium, e.g., culture medium represents less than about 20 %, more preferably less than about 10 %, and most preferably less than about 5 % of the volume of the protein preparation.
- the term "protein” is used to describe a folded polypeptide having a secondary or tertiary structure.
- the protein may be composed of a single polypeptide or may comprise multiple polypepties that are assembled to form a multimer.
- the multimer may be a homooligomer, or a heterooligmer.
- the protein may be a naturally occurring, or wild type protein, or a modified, or non-naturally, occurring protein.
- the protein may, for example, differ from a wild type protein by the addition, substitution, or deletion of one or more amino acids.
- a "variant" of a protein encompasses peptides, oligopeptides, polypeptides, proteins, and enzymes having amino acid substitutions, deletions and/or insertions relative to the unmodified or wild-type protein in question and having similar biological and functional activity as the unmodified protein from which they are derived.
- amino acid identity refers to the extent that sequences are identical on an amino acid- by-amino acid basis over a window of comparison.
- a "percentage of sequence identity” is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (/.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.
- the identical amino acid residue e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met
- a "variant" has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity can also be to a fragment or portion of the full-length polynucleotide or polypeptide. Hence, a sequence may have only 50 % overall sequence identity with a full-length reference sequence, but a sequence of a particular region, domain or subunit could share 80 %, 90 %, or as much as 99 % sequence identity with the reference sequence.
- wild-type refers to a gene or gene product isolated from a naturally occurring source.
- a wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the "normal” or “wild-type” form of the gene.
- modified refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications and/or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product.
- methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer.
- Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art.
- non-naturally occurring amino acids may be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, they may be introduced by expressing the mutant monomer in E.
- coli that are auxotrophic for specific amino acids in the presence of synthetic (/.e., non-naturally occurring) analogues of those specific amino acids. They may also be produced by naked ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side-chain volume.
- the amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acids they replace.
- the conservative substitution may introduce another amino acid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid.
- Conservative amino acid changes are well- known in the art and may be selected in accordance with the properties of the 20 main amino acids as defined in Table 1 below. Where amino acids have similar polarity, this can also be determined by reference to the hydropathy scale for amino acid side chains in Table 2.
- a mutant or modified protein, monomer or peptide can also be chemically modified in any way and at any site.
- a mutant or modified monomer or peptide is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art.
- the mutant of modified protein, monomer or peptide may be chemically modified by the attachment of any molecule.
- the mutant of modified protein, monomer or peptide may be chemically modified by attachment of a dye or a fluorophore.
- the invention provides a method of uniquely labelling RIMA molecules in a population of cells.
- the method achieves this by producing triple barcoded constructs transcribed from the RNA molecules in step (f).
- uniquely labelling RNA molecules is synonymous with producing uniquely labelled constructs transcribed from the RNA molecules.
- the method is preferably for producing uniquely labelled constructs transcribed from the RNA molecules in a population of cells.
- the uniquely labelled constructs transcribed from the RNA molecules are preferably uniquely labelled constructs comprising DNA sequences transcribed from the RNA molecules.
- the uniquely labelled constructs are triple barcoded constructs.
- the population may comprise any number of cells bearing in mind the discussion below concerning the number of possible barcode combinations.
- the population for use in the invention preferably comprises at least about 1.00 x 10 4 cells.
- the population more preferably comprises at least about 1.30 x 10 4 cells, at least about 2.00 x 10 4 cells, at least about 5.00 x 10 4 cells, at least about 1.00 x 10 5 cells, at least about 1.10 x 10 5 cells, at least about 2.00 x 10 5 cells, at least about 5.00 x 10 5 cells, at least about 6.00 x 10 5 cells, at least about 7.00 x 10 5 cells, at least about 8.00 x 10 5 cells, at least about 9.00 x
- 10 5 cells at least about 1.00 x 10 6 , at least about 2.00 x 10 6 , at least about 5.00 x 10 6 , at least about 1.00 x 10 7 , at least about 2.00 x 10 7 , at least about 5.00 x 10 7 , at least about 1.00 x 10 8 , at least about 2.00 x 10 8 cells, at least about 5.00 x 10 8 cells, at least about 1.00 x 10 9 cells, at least about 2.00 x 10 9 cells, at least about 3.00 x 10 9 cells, at least about 5.00 x 10 9 cells, at least about 1.00 x 10 10 cells, at least about 2.00 x 10 10 cells, at least about 4.00 x 10 10 cells or at least about 5.00 x 10 10 cells.
- the population may contain even more cells, such as at least about 10 11 cells or at least about 10 12 cells. A greater range of transcriptomes across individual cells or two or more cells will be seen with a larger number of starting cells. A greater number of possible barcode combinations is also needed
- the population preferably comprises about 5.00 x 10 10 or fewer cells.
- the population more preferably comprises such as about 4.12 x 10 10 or fewer cells, about 4.10 x 10 10 or fewer cells, about 4.00 x 10 10 or fewer, about 3.50 x 10 10 or fewer, about 3.00 x 10 10 or fewer, about 2.00 x 10 10 or fewer cells, about 1.00 x 10 10 or fewer cells, about 5.00 x 10 9 or fewer cells, about 3.62 x 10 9 or fewer cells, about 3.60 x 10 9 or fewer cells, about 3.50 x 10 9 or fewer cells, about 3.00 x 10 9 or fewer cells, about 2.00 x 10 9 or fewer cells, about 1.00 x 10 9 or fewer cells, about 5.00 x 10 8 or fewer cells, about 2.00 x 10 8 or fewer cells, about 1.00 x 10 8 or fewer cells, about 5.66 x 10 7 or fewer cells, about 5.60 x 10 7 or fewer cells, about 5.
- 10 6 or fewer cells about 1.00 x 10 6 or fewer cells, about 8.84 x 10 5 or fewer cells, about 8.80 x 10 5 or fewer cells, about 8.50 x 10 5 or fewer cells, about 8.00 x 10 5 or fewer cells, about 5.00 x 10 5 or fewer cells, about 2.00 x 10 5 or fewer cells, about 1.10 x 10 5 or fewer cells, about 1.00 x 10 5 or fewer cells, about 5.00 x 10 4 or fewer cells, about 2.00 x 10 4 or fewer cells, about 1.38 x 10 4 or fewer cells, about 1.30 x 10 4 or fewer cells or about 1.00 x 10 4 or fewer cells.
- the method may be conducted in a 24-well plate.
- the population preferably comprises about 1.38 x 10 4 or fewer cells, about 1.30 x 10 4 or fewer cells or about 1.00 x 10 4 or fewer cells.
- the method may be conducted in a 48-well plate.
- the population preferably comprises about 1.10 x 10 5 or fewer cells, about 1.00 x 10 5 or fewer cells, about 5.00 x 10 4 or fewer cells, or about 2.00 x 10 4 or fewer cells.
- the population may comprise any of the number of cells for a 24-well plate.
- the method may be conducted in a 96-well plate.
- the population preferably comprises about 8.84 x 10 5 or fewer cells, about 8.80 x 10 5 or fewer cells, about 8.50 x 10 5 or fewer cells, about 8.00 x 10 5 or fewer cells, about 5.00 x 10 5 or fewer cells, or about 2.00 x 10 5 or fewer cells.
- the population may comprise any of the number of cells for a 24-well plate or 48-well plate.
- the method may be conducted in a 384-well plate.
- the population preferably comprises about 5.66 x 10 7 or fewer, about 5.60 x 10 7 or fewer, about 5.50 x 10 7 or fewer, about 5.00 x 10 7 or fewer, about 2.00 x 10 7 cells or fewer, about 1.00 x 10 7 or fewer, about 5.00 x 10 6 or fewer, about 2.00 x 10 6 or fewer, or about 1.00 x 10 6 or fewer cells.
- the population may comprise any of the number of cells for a 24-well plate, 48-well plate, or 96-well plate.
- the method may be conducted in a 1536-well plate.
- the population preferably comprises about 3.62 x 10 9 or fewer, about 3.60 x 10 9 or fewer, about 3.50 x 10 9 or fewer, about 3.00 x 10 9 or fewer, about 2.00 x 10 9 or fewer, about 1.00 x 10 9 or fewer, about 5.00 x 10 8 or fewer, about 2.00 x 10 8 or fewer, or about 1.00 x 10 8 or fewer cells.
- the population may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate or 384- well plate.
- the method may be conducted in a 3456-well plate.
- the population preferably comprises about 4.12 x 10 10 or fewer, about 4.10 x 10 10 or fewer, about 4.00 x 10 10 or fewer, about 2.00 x 10 10 or fewer, about 1.00 x 10 10 or fewer, or about 5.00 x 10 9 or fewer cells.
- the population may comprise any of the number of cells for a 24-well plate, 48-well plate, 96- well plate, 384-well plate, or 1536-well plate.
- the cells may be any type of cells.
- the cells may be prokaryotic cells.
- the cells may be bacterial or archaeal cells.
- the cells are typically eukaryotic cells.
- the cells may be protozoan, algal, fungal, plant or animal cells.
- the cells may be mammalian cells, such as a human, dog, cat, primate, horse, murine, rat, rodent, bovine, murine, porcine, or ovine cells.
- the cells are preferably human cells.
- the cells may be plant cells, such as cereal, legume, fruit, or vegetable cells. Examples include, but are not limited to, wheat, barley, oat, canola, maize, soya, rice, banana, apple, tomato, potato, grape, tobacco, bean, lentil, sugar cane, cocoa, cotton, tea, or coffee cells.
- the animal cells may be derived from the ectoderm, endoderm, or mesoderm.
- the cells may be stem cells, such as embryonic stem cells, induced pluripotent stem cells or mesenchymal stem cells, bone cells, such as osteoclasts, osteoblasts or osteocytes, tendon cells, such as tenoblasts or tenocytes, chondrocytes, synovial cells, vascular cells, blood cells, such as red blood cells, immune cells, platelet, neutrophils or basophils, muscle cells, such as skeletal muscle cells, cardiac muscle cells or smooth muscle cells, reproductive cells, such as sperm, oocytes, duct cells or epididymal cells, secretory cells, adipocytes, liver lipocytes, epithelial cells, odontoblasts, cementoblasts, hormone-secreting cells, barrier cells, exocrine secretory epithelial cells, nerve cells, astrocytes, oligodendrocytes, or neurons.
- the immune cells may be neutrophil granulocyte and precursors, such as myeloblasts, promyelocytes, myelocytes, or metamyelocytes, eosinophil granulocyte and precursors, basophil granulocyte and precursors, mast cells, leukocytes, lymphocytes, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, B cells, macrophages, dendritic cells, plasma cells, neutrophils, or monocytes.
- neutrophil granulocyte and precursors such as myeloblasts, promyelocytes, myelocytes, or metamyelocytes, eosinophil granulocyte and precursors, basophil granulocyte and precursors, mast cells, leukocytes, lymphocytes, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, B cells, macrophages, dendritic cells, plasma cells, neutrophils, or monocytes.
- the cells may be wild-type or naturally occurring.
- the cells may be genetically modified or genetically engineered.
- the immune cells may be genetically engineered to express a recombinant chimeric antigen receptor (CAR) or T cell receptor (TCR).
- CAR chimeric antigen receptor
- TCR T cell receptor
- the cells may be genetically modified or genetically engineered using transduction or transfection or any other common techniques known to those skilled in the art.
- the cells may be healthy cells or obtained from a healthy donor or source.
- the cells may be diseased or damaged, associated with a disease or damage or obtained from a diseased or damaged donor or source.
- the cells in the population may be homologous.
- the cells may be the same type of cells or derived from a single donor or source.
- the cells in the population may be heterologous.
- the cells may be a mixture of different types of cells.
- the cells may be a mixture of different types of cells from the same donor or source.
- the cells in the population may be derived from different donors or sources.
- the method typically involves the loss and/or destruction of at least some of the cells in the population. This means not all of the cells in the population are uniquely labelled using the method.
- the number of cells at the end of the method is typically lower than the number of cells in the population at the beginning of the method.
- About 80% or fewer of the cells in the population are preferably present at the end of the method. Even fewer cells may be present at the end of the method.
- the method typically involves at least three rounds as explained in more detail below.
- the number of cells at the end of each round is typically lower than the number of cells at the beginning of each round.
- the number of cells at the end of the first round is typically lower than the number of cells at the beginning of the first round (/.e., than the number of cells in the population).
- the number of cells at the end of the second round is typically lower than the number of cells at the beginning of the second round.
- the number of cells at the end of the third round is typically lower than the number of cells at the beginning of the third round.
- About 95% or fewer of the cells at the beginning of a round are preferably present at the end of the round.
- the method uniquely labels RIMA molecules in a population of cells.
- This typically means the RNA molecules from at least about 80% of the cells at the end of the method are labelled with a unique combination of barcodes.
- the RNA molecules from at least about 80% of the cells at the end of the method are preferably labelled with a unique combination of at least three barcodes.
- the RNA molecules from at least about 80% of the cells at the end of the method are preferably labelled with a unique triple barcode.
- the RNA molecules from at least about 80% of the cells at the end of the method are labelled with a combination of barcodes which is not present for any other cell at the end of the method.
- RNA molecules from at least about 80% of the cells at the end of the method are preferably labelled with a combination of at least three barcodes which is not present for any other cell at the end of the method.
- the RNA molecules from at least about 80% of the cells at the end of the method are preferably labelled with a triple barcode which is not present for any other cell at the end of the method.
- RNA molecules from at least about 85%, at least 90%, at least about 95%, at least about 97%, at least about 98% or at least about 99% of the cells at the end of the method are preferably labelled in any of these ways.
- the RNA molecules from about 100% of the cells at the end of the method are preferably labelled in any of these ways.
- the method uniquely labels the constructs produced in step (f). This typically means the constructs from at least about 80% of the cells at the end of the method are labelled with a unique combination of barcodes.
- the constructs from at least about 80% of the cells at the end of the method are preferably labelled with a unique combination of at least three barcodes.
- the constructs from at least about 80% of the cells at the end of the method are preferably labelled with a unique triple barcode. In other words, the constructs from at least about 80% of the cells at the end of the method are labelled with a combination of barcodes which is not present for any other cell at the end of the method.
- the constructs from at least about 80% of the cells at the end of the method are preferably labelled with a combination of at least three barcodes which is not present for any other cell at the end of the method.
- the constructs from at least about 80% of the cells at the end of the method are preferably labelled with a triple barcode which is not present for any other cell at the end of the method.
- "which is not present for any other cell at the end of the method” is interchangeable with "which is not detectable for any other cell at the end of the method”.
- the constricts from at least about 85%, at least 90%, at least about 95%, at least about 97%, at least about 98% or at least about 99% of the cells at the end of the method are preferably labelled in any of these ways.
- the constructs in about 100% of the cells at the end of the method are preferably labelled in any of these ways.
- the method uniquely labels RIMA molecules in a population of cells.
- the RNA molecules from each cell at the end of the method are preferably labelled with a unique combination of at least three barcodes.
- the RNA molecules from each cell at the end of the method are preferably labelled with a unique triple barcode.
- the RNA molecules from every cell at the end of the method are labelled with a combination of barcodes which is not present for any other cell at the end of the method.
- the RNA molecules from every cell at the end of the method are preferably labelled with a combination of at least three barcodes which is not present for any other cell at the end of the method.
- RNA molecules from every cell at the end of the method are preferably labelled with a triple barcode which is not present for any other cell at the end of the method.
- “which is not present for any other cell at the end of the method” is interchangeable with “which is not detectable for any other cell at the end of the method”.
- the method uniquely labels the constructs produced in step (f). This typically means the constructs from each cell at the end of the method are labelled with a unique combination of barcodes.
- the constructs from each cell at the end of the method are preferably labelled with a unique combination of at least three barcodes.
- the constructs from each cell at the end of the method are preferably labelled with a unique triple barcode. In other words, the constructs from every cell at the end of the method are labelled with a combination of barcodes which is not present for any other cell at the end of the method.
- the constructs from every cell at the end of the method are preferably labelled with a combination of at least three barcodes which is not present for any other cell at the end of the method.
- constructs from every cell at the end of the method are preferably labelled with a triple barcode which is not present for any other cell at the end of the method.
- "which is not present for any other cell at the end of the method” is interchangeable with “which is not detectable for any other cell at the end of the method”.
- the triple barcode for at least 80% of the cells is typically unique.
- the triple barcode for at least about 80% of the cells at the end of the method is preferably not present in any other cell at the end of the method.
- the triple barcode for at least about 85%, at least 90%, at least about 95%, at least about 97%, at least about 98% or at least about 99% of the cells at the end of the method are preferably unique or not present in any other cell at the end of the method.
- the triple barcoded constructs from every cell at the end of the method typically comprise a unique triple barcode.
- the triple barcoded constructs from every cell at the end of the method preferably comprise a triple barcode which is not present for any other cell at the end of the method.
- “is not present for any other cell at the end of the method” is interchangeable with “is not detectable for any other cell at the end of the method”.
- the first round of the method of the invention comprises steps (a) and (b).
- Step (a) comprises dividing the population of cells into a plurality of first samples.
- the population may be divided into any number of first samples.
- the population is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about
- the population is preferably divided into about 24, about 48, about 96, about 384, about 1536 or about 3456 first samples.
- each first sample typically depends on the number of cells in the starting population as described above and the number of first samples.
- Each first sample typically comprises at least about 5.00 x 10 2 cells.
- Each first sample more preferably comprises at least about 5.70 x 10 2 cells, at least about 1.00 x 10 3 cells, at least about 2.00 x 10 3 cells, at least about 2.30 x 10 3 cells, at least about 5.00 x 10 3 cells, at least about 9.00 x 10 3 cells, at least about 9.20 x 10 3 cells, at least about 1.00 x 10 4 cells, at least about 2.00 x 10 4 , at least about 5.00 x 10 4 , at least about 1.00 x 10 5 , at least about 1.40 x 10 5 cells, at least about 2.00 x 10 5 , at least about 5.00 x 10 5 , at least about 1.00 x 10 6 , at least about 2.00 x 10 6 , at least about 2.30 x 10 6 cells, at least about 5.00 x 10 6 cells, at least about 1.00
- Each first sample preferably comprises about 5.00 x 10 7 or fewer cells.
- Each first sample more preferably comprises such as about 3.00 x 10 7 or fewer cells, about 2.00 x 10 7 or fewer cells, about 1.19 x 10 7 or fewer cells, about 1.15 x 10 7 or fewer cells, about 1.10 x 10 7 or fewer cells, about 1.00 x 10 7 or fewer cells, about 5.00 x 10 6 or fewer cells, about 2.35 x 10 6 or fewer cells, about 2.30 x 10 6 or fewer cells, about 2.00 x 10 6 or fewer cells, about 1.00 x 10 6 or fewer cells, about 5.00 x 10 5 or fewer cells, about 2.00 x 10 5 cells or fewer cells, about 1.50 x 10 5 or fewer cells, about 1.47 x 10 5 or fewer cells, about 1.45 x 10 5 or fewer cells, about 1.40 x 10 5 or fewer cells, 1.00 x 10 5 or fewer cells, about 5.00 x 10 4 or
- Each first sample preferably comprises about 576 or fewer cells, about 570 or fewer cells, about 550 or fewer or about 500 or fewer cells.
- Each first sample preferably comprises about 2304 or fewer cells, about 2300 or fewer cells or about 2000 or fewer cells.
- Each first sample may comprise any of the number of cells for a 24-well plate.
- Each first sample preferably comprises about 9216 or fewer cells, about 9200 or fewer cells, or about 9000 or fewer cells.
- Each first sample may comprise any of the number of cells for a 24-well plate or a 48-well plate.
- Each first sample preferably comprises about 147,456 or fewer, about 147,000 or fewer, about 1450,000 or fewer, or about 140,000 cells or fewer.
- Each first sample may comprise any of the number of cells for a 24- well plate, 48-well plate, or 96-well plate.
- Each first sample preferably comprises about 2,359,296 or fewer cells, about 2,300,000 or fewer cells, or about 2,000,000 or fewer cells.
- Each first sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, or 384-well plate.
- Each first sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells.
- Each first sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, 384-well plate, or 1536-well plate.
- the number of cells in each first sample is typically approximately or about the same.
- the skilled person will appreciate how the numbers of cells may differ between samples based on standard techniques for measuring numbers of cells and dividing cells into different samples.
- step (d) of the method of the invention which comprises reverse transcribing the RIMA molecules in the plurality of second samples.
- the method of the invention preferably does not comprise conducting reverse transcription (RT) before step (a).
- the method of the invention preferably comprises not conducting reverse transcription (RT) before step (a).
- the method of the invention is preferably carried out on RNA molecules which have not been reversed transcribed.
- the method of the invention is preferably for uniquely labelling RNA molecules which have not been reversed transcribed in a population of cells.
- Step (b) comprises ligating first adaptors to the 3' ends of the RNA molecules in the plurality of first samples.
- Any method of ligation may be used in accordance with the invention, including the method described in the Examples.
- the skilled person understands the ligases that may be used in the invention, such as T4 DNA ligase and other DNA ligases. Suitable conditions for ligation reactions are known in the art and described in the Examples.
- the first adaptors comprise primer sites for reverse transcription (RT). Any RT primer sites may be used in accordance with the invention, including those described in the Examples. RT primer sites are typically sequences in the first adaptors which specifically hybridise to a portion or region of the RT primers used in the second round.
- RT primer sites “specifically hybridise” to the portion or region of the RT primers when they hybridise with preferential or high affinity to the portion or region of the RT primers but do not substantially hybridise, does not hybridise, or hybridises with only low affinity to other polynucleotide sequences, especially other RNA molecules, adaptors or sequences used in the invention.
- Conditions that permit the hybridisation are well-known in the art (for example, Sambrook et al., 2001, Molecular Cloning: a laboratory manual, 3rd edition, Cold Spring Harbour Laboratory Press; and Current Protocols in Molecular Biology, Chapter 2, Ausubel et al., Eds., Greene Publishing and Wiley-lnterscience, New York (1995)).
- Hybridisation can be carried out under low stringency conditions, for example in the presence of a buffered solution of 30 to 35% formamide, 1 M NaCI and 1 % SDS (sodium dodecyl sulfate) at 37 °C followed by a 20 wash in from IX (0.1650 M Na+) to 2X (0.33 M Na+) SSC (standard sodium citrate) at 50 °C.
- Hybridisation can be carried out under moderate stringency conditions, for example in the presence of a buffer solution of 40 to 45% formamide, 1 M NaCI, and 1 % SDS at 37 °C, followed by a wash in from 0.5X (0.0825 M Na+) to IX (0.1650 M Na+) SSC at 55 °C.
- Hybridisation can be carried out under high stringency conditions, for example in the presence of a buffered solution of 50% formamide, 1 M NaCI, 1% SDS at 37 °C, followed by a wash in 0.1X (0.0165 M Na+) SSC at 60 °C.
- the RT primer sites "specifically hybridise” if they hybridise to the portion or region of the RT primers with a melting temperature (Tm) that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C or at least 10 °C, greater than its Tm for other polynucleotide sequences.
- Tm melting temperature
- the RT primer sites hybridises to the portion or region of the RT primers with a Tm that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C, greater than its Tm for other polynucleotide sequences.
- the RT primer sites hybridise to the portion or region of the RT primers with a Tm that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C, greater than its Tm for a polynucleotide which differs from the portion or region of the RT primers by one or more nucleotides, such as by 1, 2, 3, 4 or 5 or more nucleotides.
- the RT primer sites typically hybridise to the portion or region of the RT primers with a Tm of at least 90 °C, such as at least 92 °C or at least 95 °C.
- Tm can be measured experimentally using known techniques, including the use of DNA microarrays, or can be calculated using publicly available Tm calculators, such as those available over the internet.
- the RT primers sites typically comprise or consist of a sequence at least about 80% identical or homologous to the reverse complement of the portion or region of the RT primers.
- the RT primers sites preferably comprise or consist of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of the portion or region of the RT primers.
- the RT primers sites most preferably comprise or consist of a sequence which is the reverse complement of the portion or region of the RT primers.
- complementary it is meant the RT primer sites comprise or consist of a sequence 100% identical or homologous to the reverse complement of the portion or region of the RT primers. Complementarity is typically determined using canonical Watson-Crick base pairing.
- Standard methods in the art may be used to determine homology.
- the UWGCG Package provides the BESTFIT program which can be used to calculate homology, for example used on its default settings (Devereux et al (1984) Nucleic Acids Research 12, p387-395).
- the PILEUP and BLAST algorithms can be used to calculate homology or line up sequences (such as identifying equivalent residues or corresponding sequences (typically on their default settings)), for example as described in Altschul S. F. (1993) J Mol Evol 36:290- 300; Altschul, S.F et al (1990) J Mol Biol 215:403-10.
- Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http://www.ncbi.nlm.nih.gov/).
- the RT primer sites and/or the portion or region of the RT primers may be any length, such as at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 21 nucleotides, at least about 25 nucleotides or at least about 25 nucleotides in length.
- the RT primer sites and the portion or region of the RT primers are preferably the same length.
- the first adaptors used in each sample may use the same RT primers sites. This allows the method to use a standard set of first adaptors and RT primers in which only the first and second barcodes differ between samples.
- the first adaptors used in each first sample comprise the reverse complement of a different first barcode.
- the first adaptors comprise the reverse complement of different barcodes because the reverse complement is reverse transcribed in the second round to create the first different barcodes.
- the first adaptors used in each first sample comprise the reverse complement of a unique first barcode.
- the RNA molecules in each first sample are labelled with a different or unique first barcode.
- the first adaptors used in each first sample comprise the reverse complement of a first barcode that is different from the first barcode used in all the other first samples.
- the RNA molecules in each first sample are labelled with a first barcode which is different from the first barcode used to label the RNA molecules in all the other first samples.
- the first adaptors used in each first sample comprise the reverse complement of a first barcode that is different from any other barcode or all the other barcodes used in the method.
- the RNA molecules in each first sample are labelled with a first barcode which is different from any other barcode or all the other barcodes used in the method. Barcodes are discussed in more detail below.
- the first adaptors are typically polynucleotides.
- the first adaptors may be any type of polynucleotide.
- a polynucleotide such as a nucleic acid, is a macromolecule comprising two or more nucleotides.
- a polynucleotide can be single-stranded or double-stranded.
- a doublestranded polynucleotide is made of two single stranded polynucleotides hybridised together.
- the first adaptor can be a single-stranded polynucleotide or a double-stranded polynucleotide.
- a polynucleotide may comprise any combination of any nucleotides.
- the nucleotides can be naturally occurring or artificial.
- a nucleotide typically contains a nucleobase, a sugar and at least one phosphate group.
- the nucleobase and sugar form a nucleoside.
- the nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine (A), guanine (G), thymine (T), uracil (U) and cytosine (C).
- the sugar is typically a pentose sugar.
- Nucleotide sugars include, but are not limited to, ribose and deoxyribose.
- the sugar is preferably a deoxyribose.
- the polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and/or thymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC).
- the nucleotide is typically a ribonucleotide or deoxyribonucleotide.
- the nucleotide typically contains a monophosphate, diphosphate, or triphosphate.
- the nucleotide may comprise more than three phosphates, such as 4 or 5 phosphates. Phosphates may be attached on the 5' or 3' side of a nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5- hydroxy methylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate.
- a nucleotide may be abasic (/.e., lack a nucleobase).
- a nucleotide may also lack a nucleobase and a sugar (/.e., is a C3 spacer).
- the nucleotides in the polynucleotide may be attached to each other in any manner.
- the nucleotides are typically attached by their sugar and phosphate groups as in nucleic acids.
- the nucleotides may be connected via their nucleobases as in pyrimidine dimers.
- the polynucleotide can be a nucleic acid, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA).
- the polynucleotide can comprise one strand of RNA hybridized to one strand of DNA.
- the polynucleotide may be any synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), bridged nucleic acid (BNA) or other synthetic polymers with nucleotide side chains.
- the PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds.
- the GNA backbone is composed of repeating glycol units linked by phosphodiester bonds.
- the TNA backbone is composed of repeating threose sugars linked together by phosphodiester bonds.
- LNA is formed from ribonucleotides as discussed above having an extra bridge connecting the 2’ oxygen and 4’ carbon in the ribose moiety.
- the polynucleotide is preferably DNA, RNA or a DNA or RNA hybrid, most preferably DNA.
- a DNA/RNA hybrid may comprise DNA and RNA on the same strand.
- the DNA/RNA hybrid comprises one DNA strand hybridized to an RNA strand.
- the backbone of the polynucleotide can be altered to reduce the possibility of strand scission.
- DNA is known to be more stable than RNA under many conditions.
- the backbone of the polynucleotide strand can be modified to avoid damage caused by e.g., harsh chemicals such as free radicals.
- DNA or RNA that contains unnatural or modified bases can be produced by amplifying natural DNA or RNA polynucleotides in the presence of modified NTPs using an appropriate polymerase.
- the first adaptors are typically synthetic or semi-synthetic.
- DNA or RIMA may be purely synthetic, synthesised by conventional DNA synthesis methods such as phosphoramidite based chemistries.
- Synthetic polynucleotides subunits may be joined together by known means, such as ligation or chemical linkage, to produce longer strands.
- Internal self-forming structures e.g., hairpins, quadruplexes
- Synthetic polynucleotides can be copied and scaled up for production by means known in the art, including PCR, incorporation into bacterial factories, and the like.
- the first adaptors can be any length.
- the first adaptors can be at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 56 nucleotides, at least about 60, or at least 60 nucleotides or nucleotide pairs in length.
- the first adaptors are preferably double stranded.
- the first adaptors are preferably double stranded DNA.
- the first adaptors may be called barcoded cDNA reverse transcription adaptors (barcoded CDRAs).
- the first adaptors are preferably double stranded and comprise overhangs which are capable of hybridising or specifically hybridising to the 3' ends of the RNA molecules.
- the non-overhanging strands of such adaptors are typically ligated to the RNA molecules.
- the skilled person is capable of designing overhangs capable of hybridising or specifically hybridising to the 3' ends of the RNA molecules.
- the first adaptors used in each first sample are preferably double stranded and comprise overhangs comprising every possible combination of sequences based on nucleotides comprising adenosine (A), thymine (T), uracil (U), guanine (G) and cytosine (C).
- Step (a) preferably comprises hybridising or specifically hybridising overhangs of double stranded first adaptors to the 3' ends of the RNA molecules in the plurality of first samples, ligating the non-overhanging strands of the first adaptors to the 3' ends of the RNA molecules, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode.
- RT primer sites for reverse transcription
- Eukaryotic RNA molecules typically have poly(A) tails (see Figure 1).
- the first adaptors are preferably double stranded and comprise overhangs which are capable of hybridising or specifically hybridising to the poly(A) tails of the RNA molecules.
- the non-overhanging strands of such adaptors are typically ligated to the RNA molecules. An example of this is shown in Figure 1.
- Step (a) preferably comprises hybridising or specifically hybridising overhangs of double stranded first adaptors to poly(A) tails of the RNA molecules in the plurality of first samples, ligating the non-overhanging strands of the first adaptors to the poly(A) tails of the RIMA molecules, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode.
- RT reverse transcription
- Non-eukaryotic RNA molecules typically do not have poly(A) tails.
- Step (b) may further comprise polyadenylating the 3' ends of the RNA molecules and using double stranded first adaptors which comprise overhangs which are capable of hybridising or specifically hybridising to the polyadenylated 3' ends of the RNA molecules.
- the non-overhanging strands of such adaptors are typically ligated to the RNA molecules.
- Step (a) may comprise polyadenylating the 3' ends of the RNA molecules, hybridising or specifically hybridising overhangs of double stranded first adaptors to polyadenylated 3' ends of the RNA molecules in the plurality of first samples, ligating the non-overhanging strands of the first adaptors to the polyadenylated 3' ends of the RNA molecules, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode.
- the RNA molecules may be polyadenylated using any known method, such as using E. coli Poly(A) polymerase (EPAP).
- the overhangs can be any length.
- the overhangs can be at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 15, at least about 20, or at least 25 nucleotides in length.
- the overhangs are preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length. Hybridisation and specific hybridisation are discussed above and any of those embodiments apply to the overhangs to the 3' ends, the poly(A) tails or the polyadenylated 3' ends.
- the overhangs preferably comprise one or more of (i) thymine-containing nucleotides, (ii) uracil-containing nucleotides and (iii) universal nucleotides.
- Overhangs comprising (i) and (ii) are capable of specifically hybridising to the poly(A) tails or polyadenylated 3' ends as described above.
- Nucleotides are defined above.
- a universal nucleotide is one which will hybridise or bind to some degree to all of the nucleotides in a polynucleotide.
- a universal nucleotide is preferably one which will hybridise or bind to some degree to nucleotides comprising the nucleosides adenosine (A), thymine (T), uracil (U), guanine (G) and cytosine (C).
- the universal nucleotide may hybridise or bind more strongly to some nucleotides than to others.
- the universal nucleotide preferably comprises one of the following nucleobases: hypoxanthine, 4-nitroindole, 5-nitroindole, 6-nitroindole, formylindole, 3-nitropyrrole, nitroimidazole, 4-nitropyrazole, 4-nitrobenzimidazole, 5-nitroindazole, 4- aminobenzimidazole or phenyl (C6-aromatic ring).
- the universal nucleotide more preferably comprises one of the following nucleosides: 2'-deoxyinosine, inosine, 7-deaza-2'- deoxyinosine, 7-deaza-inosine, 2-aza-deoxyinosine, 2-aza-inosine, 2-O'-methylinosine, 4- nitroindole 2'-deoxyribonucleoside, 4-nitroindole ribonucleoside, 5-nitroindole 2'- deoxyribonucleoside, 5-nitroindole ribonucleoside, 6-nitroindole 2'-deoxyribonucleoside, 6- nitroindole ribonucleoside, 3-nitropyrrole 2'-deoxyribonucleoside, 3-nitropyrrole ribonucleoside, an acyclic sugar analogue of hypoxanthine, nitroimidazole 2'- deoxyribonucleoside, nitroimidazole ribonu
- the universal nucleotide more preferably comprises 2'-deoxyinosine.
- the universal nucleotide is more preferably IMP or dIMP.
- the universal nucleotide is most preferably dPMP (2'-Deoxy-P-nucleoside monophosphate) or dKMP (N6-methoxy-2, 6- diaminopurine monophosphate).
- the overhangs preferably comprise (in the 5' to 3' direction) one or more consecutive repeating units of UTT, such as 2 or more, 3 or more, 4 or more or 5 more consecutive repeating units of UTT.
- the first adaptors used in each sample may use the same overhangs. This allows the method to use a standard set of first adaptors in which only the first barcodes differ between first samples.
- the reverse complements of the different first barcodes are preferably present in the strands of the first adaptors which are ligated to the RIMA molecules.
- the reverse complements of the different first barcodes are preferably present in the opposite strands from the strands with the overhangs which hybridise or specifically hybridise to the 3' ends, the poly(A) tails or the polyadenylated 3' ends of the RNA molecules. If the reverse complements of the different first barcodes are ligated to the RNA molecules, the different first barcodes will be present in the cDNA product of the reverse transcription in the second round (see for example Figure 2).
- the nucleotides at the 3' ends of the first adaptors are preferably dideoxy cytosine (ddC) or inverted thymidine (inverted dT).
- the nucleotides at the 3' ends of one of both strands of the double stranded first adaptors are preferably ddC or inverted dT.
- the nucleotides at the 3' ends of the strand of the double stranded first adaptors are preferably ddC or inverted dT.
- the non-ligated strands of the first adaptors are preferably removed before step (d). If the first adaptors are double stranded and comprise overhangs which are capable of hybridising or specifically hybridising to the 3' ends, the poly(A) tails or the polyadenylated 3' ends of the RIMA molecules, the overhanging strands of the adaptors are preferably removed before step (d).
- the non-ligated/overhanging strands may be removed using any suitable method, including digestion.
- the strands are preferably removed using USER digestion. USER enzyme mix is a blend of Uracil DNA glycosylase (UDG) and the DNA glycosylase-lyase Endonuclease VIII. As explained above, this mitigates the possibility of re-barcoding or crosstalk upon pooling.
- the second round of the method of the invention comprises steps (c) and (d).
- Step (c) comprises pooling the plurality of first samples and dividing the pool of cells into a plurality of second samples.
- the pool may be divided into any number of second samples.
- the pool is preferably divided into any number of second samples discussed above with respect to the first samples.
- the pool is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 second samples.
- Each second sample is preferably divided into about 24, about 48, about 96, about 384, about 1536 or about 3456 second samples. The number of second samples is typically identical to the number of first samples.
- Each second sample typically comprises at least about 5.00 x 10 2 cells.
- Each second sample more preferably comprises at least about 5.70 x 10 2 cells, at least about 1.00 x 10 3 cells, at least about 2.00 x 10 3 cells, at least about 2.30 x 10 3 cells, at least about 5.00 x 10 3 cells, at least about 9.00 x 10 3 cells, at least about 9.20 x 10 3 cells, at least about 1.00 x 10 4 cells, at least about 2.00 x 10 4 , at least about 5.00 x 10 4 , at least about 1.00 x 10 5 , at least about 1.40 x 10 5 cells, at least about 2.00 x 10 5 , at least about 5.00 x 10 5 , at least about 1.00 x 10 6 , at least about 2.00 x 10 6 , at least about 2.30 x 10 6 cells, at least about 5.00 x 10 6 cells, at least about 1.00 x 10 7 cells, at least about 1.10 x 10 7 cells, at least about 2.00 x 10 7
- Each second sample preferably comprises about 5.00 x 10 7 or fewer cells.
- Each second sample more preferably comprises such as about 3.00 x 10 7 or fewer cells, about 2.00 x 10 7 or fewer cells, about 1.19 x 10 7 or fewer cells, about 1.15 x 10 7 or fewer cells, about 1.10 x 10 7 or fewer cells, about 1.00 x 10 7 or fewer cells, about 5.00 x 10 6 or fewer cells, about 2.35 x 10 6 or fewer cells, about 2.30 x 10 6 or fewer cells, about 2.00 x 10 6 or fewer cells, about 1.00 x 10 6 or fewer cells, about 5.00 x 10 5 or fewer cells, about 2.00 x 10 5 cells or fewer cells, about 1.50 x 10 5 or fewer cells, about 1.47 x 10 5 or fewer cells, about 1.45 x 10 5 or fewer cells, about 1.40 x 10 5 or fewer cells, 1.00 x 10 5 or fewer cells, about 5.00 x 10 4 or
- Each second sample preferably comprises about 576 or fewer cells, about 570 or fewer cells, about 550 or fewer or about 500 or fewer cells.
- Each second sample preferably comprises about 2304 or fewer cells, about 2300 or fewer cells or about 2000 or fewer cells.
- Each second sample may comprise any of the number of cells for a 24-well plate.
- Each second sample preferably comprises about 9216 or fewer cells, about 9200 or fewer cells, or about 9000 or fewer cells.
- Each second sample may comprise any of the number of cells for a 24-well plate or a 48-well plate.
- Each second sample preferably comprises about 147,456 or fewer, about 147,000 or fewer, about 1450,000 or fewer, or about 140,000 cells or fewer.
- Each second sample may comprise any of the number of cells for a 24-well plate, 48-well plate, or 96-well plate.
- the method may be conducted in a 1536-well plate.
- Each second sample preferably comprises about 2,359,296 or fewer cells, about 2,300,000 or fewer cells, or about 2,000,000 or fewer cells.
- Each second sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, or 384-well plate.
- the method may be conducted in a 3456-well plate.
- Each second sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells.
- Each second sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, 384-well plate, or 1536-well plate.
- the number of cells in each second sample is typically approximately or about the same.
- the number of cells is each second sample is typically approximately or about the same as the number of cells in each first sample.
- the skilled person will appreciate how the numbers of cells may differ between samples based on standard techniques for measuring numbers of cells and dividing cells into different samples.
- Step (d) comprises reverse transcribing the RIMA molecules in the plurality of second samples.
- Methods for conducting reverse transcription are well known in the art and any suitable conditions may be used. Preferred conditions are described in the Examples.
- Reverse transcription involves the use of the enzyme reverse transcriptase to convert RNA into cDNA Reverse transcriptases are commercially available (e.g. Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753), Superscript® II reverse transcriptase (Invitrogen) and Affinity script (Agilent)).
- the reverse transcriptase is preferably Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753).
- the doubles stranded constructs produced in step (b) typically comprise the RNA molecules hybridised to cDNA strands produced by reverse transcription.
- Step (d) uses the primer sites in the first adaptors and RT primers to form double stranded constructs.
- the skilled person is capable of designing suitable RT primers for use in the invention.
- the RT primers are typically polynucleotide RT primers and may be any of the polynucleotides discussed above with reference to the first adaptors.
- the RT primers are typically synthetic or semi-synthetic.
- the RT primers are preferably single stranded polynucleotide RT primers.
- the RT primers may be any length.
- the RT primers are preferably at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 32 nucleotides, at least about 35 nucleotides, or at least about 40 nucleotides in length.
- the RT primers are preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length.
- Step (d) preferably comprises hybridising the RT primers to the first adaptors.
- Step (d) preferably comprises hybridising the RT primers to the RT primers sites in the first adaptors.
- Step (d) preferably comprises hybridising RT primers to the first adaptors or to the RT primers sites in the first adaptors and reverse transcribing the RNA molecules in the plurality of second samples using the primer sites and RT primers to form double stranded constructs.
- Hybridisation is preferably specific hybridisation.
- the RT primer sites are typically sequences in the first adaptors which specifically hybridise to a portion or region of the RT primers used in the second round.
- RT primer sites and portions or regions of the RT primers are discussed above with reference to the first round.
- the RT primers may comprise any of the portions or regions discussed above.
- the RT primers may use the same portion or region which specifically hybridises to the RT primer sites. This allows the method to use a standard set of RT primers in which only the second barcodes differ between second samples.
- the RT primers used in each second sample comprise a different second barcode.
- the RT primers used in each second sample comprise a unique second barcode.
- the RNA molecules in each second sample are labelled with a different or unique second barcode.
- the RT primers used in each second sample comprise a second barcode that is different from the second barcode used in all the other second samples.
- the RNA molecules in each second sample are labelled with a second barcode which is different from the second barcode used to label the RNA molecules in all the other second samples.
- the RT primers used in each second sample comprise a second barcode that is different from any other barcode or all the other barcodes used in the method.
- the RNA molecules in each second sample are labelled with a second barcode which is different from any other barcode or all the other barcodes used in the method. Barcodes are discussed in more detail below.
- the RT primers preferably comprise sequences at or near their 5' ends which are capable of specifically hybridising to the second adaptors used in step (f).
- the sequences preferably specifically hybridise to overhangs in the second adaptors as discussed in more detail below. Specific hybridisation is defined above.
- the sequences may be any length.
- the sequences preferably at least about 5 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 12 nucleotides, at least about 15 nucleotides, or at least about 20 nucleotides in length.
- the sequences are preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length.
- the RT primers may use the same sequences at or near their 5' ends which are capable of specifically hybridising to the second adaptors used in step (f). This allows the method to use a standard set of RT primers in which only the second barcodes differ between second samples.
- RT primers are shown in the Examples.
- the RT in step (d) preferably produces strands comprising the different first barcodes, the different second barcodes and sequences complementary to the RNA molecules.
- the different first barcodes are transcribed from the reverse complements in the first adaptors. An example of this is shown in Figure 2.
- the components are preferably in the following order in the strands 5' to 3': different second barcodes, different first barcodes and complementary sequences. These strand constructs are double barcoded.
- Step (d) preferably further comprises (i) hybridising strand switching primers (SSPs) to overhangs at the 3' end of the strands in the double stranded constructs produced by reverse transcription of the RIMA molecules and (ii) reverse transcribing the SSPs to elongate the double stranded constructs.
- SSPs strand switching primers
- the SSPs are preferably specifically hybridised to the overhangs.
- the strands in the double constructs produced by reverse transcription are preferably cDNA strands. Overhangs, hybridisation, and specific hybridisation are discussed above and any of those embodiments apply to the SSP embodiments.
- the overhangs preferably comprise consecutive cytosine-containing nucleotides, such as consecutive deoxycytidine (dC)-containing nucleotides. Such overhangs can be created through the addition of non-template dC to the synthesised cDNA by reverse transcriptase when it reaches the end of an RNA template.
- the SSPs may be any of the types of polynucleotides and/or any of the lengths discussed above with reference to the RT primers used in the invention.
- the SSPs preferably comprise one or more unique molecular identifiers (UMIs) to facilitate characterisation.
- UMIs unique molecular identifiers
- the SSPs may comprise any number of UMIs, such as 2 or more, 3 or more, 4 or more, 5 or more or 10 or more.
- UMI is TTVVVVTT.
- This UMI structure is optimised for characterisation using nanopores by reducing homopolymers.
- the SSPs facilitate deduplication of PCR replicates. By including one or more UMIs in the SSPs, they can be removed from the second adapters used in the third round (unlike in the RLL approach) to save on production cost and complexity of the second adapters. Any method of reverse transcription, including any of those discussed above, may be used in step (ii).
- the third round of the method of the invention comprises steps (e) and (f).
- Step (e) comprises pooling the plurality of second samples and dividing the pool of cells into a plurality of third samples.
- the pool may be divided into any number of third samples.
- the pool is preferably divided into any number of third samples discussed above with respect to the first samples.
- the pool is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 third samples.
- Each third sample is preferably divided into about 24, about 48, about 96, about 384, about 1536 or about 3456 third samples.
- the number of third samples is typically identical to the number of first and/or second samples. The number of first samples, the number of second samples and the number of third samples is preferably the same.
- Each third sample typically comprises at least about 5.00 x 10 2 cells.
- Each third sample more preferably comprises at least about 5.70 x 10 2 cells, at least about 1.00 x 10 3 cells, at least about 2.00 x 10 3 cells, at least about 2.30 x 10 3 cells, at least about 5.00 x 10 3 cells, at least about 9.00 x 10 3 cells, at least about 9.20 x 10 3 cells, at least about 1.00 x
- Each third sample preferably comprises about 5.00 x 10 7 or fewer cells.
- Each third sample more preferably comprises such as about 3.00 x 10 7 or fewer cells, about 2.00 x 10 7 or fewer cells, about 1.19 x 10 7 or fewer cells, about 1.15 x 10 7 or fewer cells, about 1.10 x 10 7 or fewer cells, about 1.00 x 10 7 or fewer cells, about 5.00 x 10 6 or fewer cells, about 2.35 x 10 6 or fewer cells, about 2.30 x 10 6 or fewer cells, about 2.00 x 10 6 or fewer cells, about 1.00 x 10 6 or fewer cells, about 5.00 x 10 5 or fewer cells, about 2.00 x 10 5 cells or fewer cells, about 1.50 x 10 5 or fewer cells, about 1.47 x 10 5 or fewer cells, about 1.45 x
- 10 5 or fewer cells about 1.40 x 10 5 or fewer cells, 1.00 x 10 5 or fewer cells, about 5.00 x 10 4 or fewer cells, about 2.00 x 10 4 or fewer cells, about 1.00 x 10 4 or fewer cells, about 9.21 x 10 3 or fewer cells, about 9.20 x 10 3 or fewer cells, about 9.00 x 10 3 or fewer cells, about 5.00 x 10 3 or fewer cells, about 2.30 x 10 3 or fewer cells, about 2.00 x 10 3 or fewer cells, about 1.00 x 10 3 or fewer cells, about 5.76 x 10 2 or fewer cells, about 5.70 x 10 2 or fewer cells or about 5.00 x 10 2 or fewer cells.
- Each third sample preferably comprises about 576 or fewer cells, about 570 or fewer cells, about 550 or fewer or about 500 or fewer cells.
- the method may be conducted in a 48-well plate.
- Each third sample preferably comprises about 2304 or fewer cells, about 2300 or fewer cells or about 2000 or fewer cells.
- Each third sample may comprise any of the number of cells for a 24-well plate.
- the method may be conducted in a 96-well plate.
- Each third sample preferably comprises about 9216 or fewer cells, about 9200 or fewer cells, or about 9000 or fewer cells.
- Each third sample may comprise any of the number of cells for a 24-well plate or a 48-well plate.
- Each third sample preferably comprises about 147,456 or fewer, about 147,000 or fewer, about 1450,000 or fewer, or about 140,000 cells or fewer.
- Each third sample may comprise any of the number of cells for a 24- well plate, 48-well plate, or 96-well plate.
- Each third sample preferably comprises about 2,359,296 or fewer cells, about 2,300,000 or fewer cells, or about 2,000,000 or fewer cells.
- Each third sample may comprise any of the number of cells for a 24-well plate, 48- well plate, 96-well plate, or 384-well plate.
- Each third sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells.
- Each third sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, 384-well plate, or 1536-well plate.
- the number of cells in each third sample is typically approximately or about the same.
- the number of cells is each third sample is typically approximately about the same as the number of cells in each first sample and/or the same as the number of cells in each second sample.
- the number of cells is each first sample, in each second sample and in each third sample is typically approximately or about the same.
- the skilled person will appreciate how the numbers of cells may differ between samples based on standard techniques for measuring numbers of cells and dividing cells into different samples.
- Step (f) comprises ligating second adaptors to the double stranded constructs in the plurality of third samples to form triple barcoded constructs.
- the triple barcoded constructs comprise DNA sequences transcribed from the RIMA molecules. Ligation is discussed above and any of the methods above may be used. The skilled person is capable of designing suitable second adaptors for use in the invention.
- the second adaptors are typically polynucleotide adaptors and may be any of the polynucleotides discussed above with reference to the first adaptors.
- the second adaptors are typically synthetic or semisynthetic.
- the second adaptors are preferably double stranded polynucleotide adaptors.
- the second adaptors may be any length.
- the second adaptors are preferably at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 32, at least about 35, at least about 40, at least about 45, at least 47, or at least about 50 nucleotides or nucleotide pairs in length.
- the second adaptors are preferably ligated to the strands in the double stranded constructs comprising the different first barcodes, different second barcodes and complementary sequences. This allows the third different barcodes to be added to the double barcoded constructs produced in the second round.
- Step (f) preferably comprises hybridising, preferably specifically hybridising, the second adaptors to the double stranded constructs in the plurality of third samples and ligating the second adaptors to the double stranded constructs to form triple barcoded constructs.
- the triple barcoded constructs comprise DNA sequences transcribed from the RIMA molecules. Hybridisation, specific hybridisation, and ligation are discussed above and any of those approaches may be used with the second adaptors.
- the second adaptors preferably comprise overhangs capable of hybridising, preferably specifically hybridising, to the RT primers.
- the second adaptors preferably comprise overhangs capable of hybridising, preferably specifically hybridising, to sequences at or near the 5' ends of the RT primers.
- the overhangs may be any length.
- the overhangs are preferably at least about 5 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 12 nucleotides, at least about 15 nucleotides, or at least about 20 nucleotides in length.
- the overhangs are preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length.
- the overhangs are preferably the same length as the sequences at or near the 5' ends of the RT primers.
- the second adaptors may use the same overhangs. This allows the method to use a standard set of second adaptors in which only the third barcodes differ between third samples.
- the different third barcodes are preferably present in the non-overhanging strands of the second adaptors which are ligated to the double stranded constructs.
- An example of this is shown in Figure 3. This approach produces constructs which are triple barcoded.
- Step (f) is preferably carried out in the presence of 5' phosphorylated blockers which are complementary to the overhangs in the second adaptors.
- 5' phosphorylated blockers which are complementary to the overhangs in the second adaptors.
- These blockers may comprise the sequences at or near the 5' ends of the RT primers. An example of this is shown in Figure 3.
- the second adaptors preferably comprise additional sequences at their 5' ends which facilitate isolation, amplification and/or characterisation of the triple barcoded constructs.
- the additional sequences may comprise click chemistry.
- the additional sequences may comprise or facilitate the addition of sequencing adaptors. Such adaptors are discussed in more detail below.
- the second adaptors preferably comprise the same additional sequences. This allows the method to use a standard set of RT primers in which only the second barcodes differ between second samples.
- the second adaptors used in each third sample comprise a different third barcode.
- the second adaptors used in each third sample comprise a unique third barcode.
- the double stranded constructs in each third sample are labelled with a different or unique third barcode.
- the second adaptors used in each third sample comprise a third barcode that is different from the third barcode used in all the other third samples.
- the double stranded constructs in each third sample are labelled with a third barcode which is different from the third barcode used to label the double stranded constructs in all the other third samples.
- the second adaptors used in each third sample comprise a third barcode that is different from any other barcode or all the other barcodes used in the method.
- the double stranded constructs in each third sample are labelled with a third barcode which is different from any other barcode or all the other barcodes used in the method. Barcodes are discussed in more detail below.
- the method involves producing triple barcoded constructs.
- Step (b) comprises the use of first different barcodes.
- Step (d) involves the use of different second barcodes.
- Step (f) comprises the use of different third barcodes.
- the meaning of "different" or "unique” barcodes is discussed above.
- the barcodes are typically polynucleotide barcodes.
- the barcodes may be any of those discussed above.
- the first, second and third different polynucleotides all typically comprise the same type of polynucleotide.
- Polynucleotide barcodes are well-known in the art (Kozarewa, I. et al., (2011), Methods Mol. Biol. 733, p279-298).
- a barcode is a specific sequence of polynucleotide that can be characterised and identified.
- a barcode is preferably a specific sequence of polynucleotide that affects the current flowing through the pore in
- the barcode may comprise one or more different nucleotide species.
- T k-mers i.e. k-mers in which the central nucleotide is thymine-based, such as TTA, GTC, GTG and CTA
- TTA, GTC, GTG and CTA typically have the lowest current states.
- Modified versions of T nucleotides may be introduced into the modified polynucleotide to reduce the current states further and thereby increase the total current range seen when the barcode moves through the pore.
- G k-mers i.e. k-mers in which the central nucleotide is guanine-based, such as TGA, GGC, TGT and CGA
- modifying the G nucleotides in the modified polynucleotide may help them to have more independent current positions.
- Including three copies of the same nucleotide species instead of three different species may facilitate characterization because it is then only necessary to map, for example, 3- nucleotide k-mers in the modified polynucleotide. However, such modifications do reduce the information provided by the barcode.
- One or more abasic nucleotides may be included in the barcode. Using one or more abasic nucleotides results in characteristic current spikes. This allows the clear highlighting of the positions of the one or more nucleotide species in the barcode.
- the nucleotide species in the barcode may comprise a chemical atom or group such as a propynyl group, a thio group, an oxo group, a methyl group, a hydroxymethyl group, a formyl group, a carboxy group, a carbonyl group, a benzyl group, a propargyl group or a propargylamine group.
- the chemical group or atom may be or may comprise a fluorescent molecule, biotin, digoxigenin, DNP (dinitrophenol), a photo-labile group, an alkyne, DBCO, azide, free amino group, a redox dye, a mercury atom, or a selenium atom.
- the barcode may comprise a nucleotide species comprising a halogen atom.
- the halogen atom may be attached to any position on the different nucleotide species, such as the nucleobase and/or the sugar.
- the halogen atom is preferably fluorine (F), chlorine (Cl), bromine (Br) or iodine (I).
- the halogen atom is most preferably F or I.
- the barcodes may be any length.
- the barcodes are preferably at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 11 nucleotides, at least about 12 nucleotides, at least about 13 nucleotides, at least about 14 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, or at least about 50 nucleotides in length.
- the first, second and third different barcodes may the same length or different lengths.
- the invention is based on the use of unique or different barcodes for each sample. As explained in more detail below, the method often requires large numbers of unique or different barcodes. The longer the barcodes, the more different possible barcodes that can be generated and used in the method of the invention. The skilled person is capable of designing barcodes and sufficient barcode numbers for use in the method of the invention.
- the barcode used in each sample is preferably different from any other barcode or all the other barcodes used in the method.
- the barcode used in each first sample, each second sample and each third sample is preferably different from any other barcode or all the other barcodes used in the method. This means each sample is contacted with a unique or different barcode.
- the number of cells in the population is preferably less than the number of first samples multiplied by the number of second samples multiplied the number of third samples (/.e., first samples x second samples x third samples). If the number of possible barcode combinations used is greater than the number of cells in the population, there is a high probability the RIMA molecules from each cell or the constructs produced from each cell will be labelled with a unique combination of three barcodes. If the number of possible barcode combinations used is greater than the number of cells in the population, there is a high probability the RNA molecules from each cell or the constructs produced from each cell will be labelled with a unique combination of three barcodes. This allows RNA molecules to be associated with individual cells in the population and the transcriptome of each cell or two or more cells to be measured.
- the (i) plurality of first samples, (ii) plurality of second samples, and/or (ii) plurality of third samples preferably comprise(s) at least 24 samples.
- (i), (ii), (iii), (i) and (ii), (i) and (iii), (ii) and (iii), or (i), (ii) and (iii) preferably comprises at least about 24 samples.
- (i), (ii), (iii), (i) and (ii), (i) and (ii), (i) and (iii), (ii) and (iii), or (i), (ii) and (iii) preferably comprises at least 24 samples, at least 48 samples, at least 96 samples, at least 384 samples, at least 1536 samples or at least 3456 samples.
- the minimum number of unique or different barcodes is preferably at least about 13,824, at least about 110,592, at least about 884,736, at least about 56,623,104, at least about 3,623,878,656 or at least about 41,278,242,816 barcodes.
- the skilled person is capable of designing suitable barcode numbers and numbers of samples based on the number of cells in the starting population.
- the population of cells is preferably fixed and/or permeabilised before step (a). Suitable methods for doing this are known in the art (e.g., Chen et al. (2016). PBMC fixation and processing for Chromium single-cell RNA sequencing. Journal of Translational Medicine, 16(1).) Specific methods are also discussed in the Examples.
- the method preferably further comprises repeating steps (e) to (f) at least once comprising the formation of a plurality of fourth or more samples and using third or more adaptors which comprise a different fourth or more barcode for each fourth or more sample to produce quadruple or more barcoded constructs.
- This repetition may be conducted more than one to produce quintuple, sextuple, septuple or octuple or more barcoded constructs.
- the skilled person is capable of designing a method uniquely labelling RNA molecules or producing uniquely barcoded constructs comprising sequences transcribed from RNA molecules with any number of barcodes. Any of the embodiments discussed above with reference to the first, second and third rounds equally apply to any of the repeated steps.
- the method produces constructs having three or more barcodes.
- the method may produce constructs having any number of three or more barcodes, such as 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more or 20 or more barcodes.
- all of these constructs will be collectively called “uniquely labelled constructs" or “barcoded constructs”.
- the method preferably further comprises isolating the barcoded constructs from the cells. Any method of isolation may be used including any of the ones discussed below. Any of the adaptors or RT primers can comprise sequences and/or molecules which facilitate isolation.
- the method preferably further comprises amplifying the barcoded constructs.
- Any amplification method may be used, including polymerase chain reaction (PCR).
- PCR polymerase chain reaction
- Any of the adaptors or RT primers can comprise primer sites that allow amplification. This facilitates preparation of the barcoded constructs for amplification.
- the method preferably further comprises isolating the barcoded constructs from the cells and amplifying the barcoded constructs.
- the method preferably further comprises characterising or sequencing the barcoded constructs. Any method may be used for sequencing or characterising the barcoded constructs, including next generation sequencing.
- the barcoded constructs are preferably characterised or sequenced using a nanopore. This is discussed in more detail below.
- the invention provides a method of characterising the transcriptome of an individual cell or two or more cells in a population of cells, the method comprising a conducting a method of the invention on the population of cells and characterising the RNA molecules from the individual cell or two or more cells using their unique labelling.
- the RNA molecules may be characterised using their barcoding.
- the characterisation preferably comprises sequencing the RNA molecules. This can be achieved by sequencing the barcoded constructs produced from the RNA molecules.
- the invention also provides a method of characterising the transcriptome of an individual cell or two or more cells in a population of cells, the method comprising producing uniquely labelled constructs transcribed from the RNA molecules in the population of cells using a method of the invention and characterising the uniquely labelled constructs from the individual cell or two or more cells using their unique labelling.
- the constructs may be characterised using their barcode.
- the method of the invention can be conducted such there is a high probability the RIMA molecules from each cell or barcoded constructs produced from each cell will be labelled with a unique combination of three barcodes.
- the characterisation is preferably sequencing of the RNA molecules or barcoded constructs, preferably using a nanopore. Any embodiments discussed above, especially in relation to the first to third rounds, adaptors, or RT primers, equally apply to these embodiments.
- These methods may involve characterising the transcriptome of any number of two or more cells, such as 3 or more, 4 or more, 5 or more, 10 or more, 20 or more, 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more or 10,000 or more cells.
- the number of two or more cells may be any of the numbers discussed above with reference to the population of cells used in the method of the invention.
- the barcoded constructs produced using the method of the invention are preferably characterised or sequenced.
- the barcoded constructs are preferably modified using sequencing adaptors which facilitate characterisation or sequencing, especially using nanopores.
- the second adaptors may comprise additional sequences which comprise or facilitate the addition of sequencing adaptors.
- a sequencing adaptor typically comprises a polynucleotide strand capable of being attached to the end of a target polynucleotide.
- the target polynucleotide is typically intended for characterisation in accordance with methods disclosed herein, and includes the first adaptors, RT primers, second adaptors and barcoded constructs.
- a sequencing adaptor may be added to both ends of the target polynucleotide.
- different adaptors may be added to the two ends of the target polynucleotide.
- An adaptor may be added to just one end of the target polynucleotide.
- Methods of adding adaptors to polynucleotides are known in the art.
- Adaptors may be attached to polynucleotides, for example, by ligation, by click chemistry, by tagmentation, by topoisomerisation or by any other suitable method.
- An adaptor may be synthetic or artificial.
- an adaptor comprises a polymer as described herein.
- the adaptor preferably comprises a polynucleotide.
- An adaptor may comprise a single-stranded polynucleotide strand.
- An adaptor may comprise a doublestranded polynucleotide.
- a sequencing adaptor may comprise any of the polynucleotide discussed above with reference to the first adaptors and includes DNA, RNA, modified DNA (such as a basic DNA), RNA, PNA, LNA, BNA and/or PEG.
- the adaptor comprises single stranded and/or double stranded DNA or RNA.
- the sequencing adaptors may be Y adaptors.
- Y adaptors are typically double stranded and comprise (a) at one end, a region where the two strands are hybridised together and (b), at the other end, a region where the two strands are not complementary. The non- complementary parts of the strands form overhangs.
- the hybridised stem of the adaptors typically attaches to the 5' end of a first strand of a double-stranded polynucleotide and the 3' end of a second strand of a double-stranded polynucleotide; or to the 3' end of a first strand of a double-stranded polynucleotide and the 5' end of a second strand of a doublestranded polynucleotide.
- the presence of a non-complementary region in the Y adaptors gives them their Y shape since the two strands typically do not hybridise to each other unlike the double stranded portion.
- the hybridised stem end of the Y adaptors may also comprise a short overhang that allows them to specifically hybridise to and be attached to the first adaptors, RT primers, second adaptors and barcoded constructs.
- a polynucleotide binding protein may bind to an overhang of an adaptor such as a Y adaptor.
- a polynucleotide binding protein may bind to the double stranded region.
- a polynucleotide binding protein may bind to a single-stranded and/or a double-stranded region of the adaptor.
- a first polynucleotide binding protein may bind to the single-stranded region of such an adaptor and a second polynucleotide binding protein may bind to the double-stranded region of the adaptor.
- the sequencing adaptors preferably comprise a membrane anchor or a pore anchor.
- the anchors may be attached to polynucleotides that are complementary to and hence hybridised to the overhangs to which a polynucleotide binding protein is bound.
- One of the non-complementary strands of sequencing adaptors may comprise leader sequences, which when contacted with a nanopore are capable of threading into the nanopore.
- the leader sequences typically comprise a polymer such as a polynucleotide, for instance DNA or RNA, a modified polynucleotide (such as abasic DNA), PNA, LNA, polyethylene glycol (PEG) or a polypeptide.
- the leader sequences preferably comprise a single strand of DNA, such as a poly dT section.
- the leader sequences can be any length, but are typically 10 to 150 nucleotides in length, such as from 20 to 120, 30 to 100, 40 to 80 or 50 to 70 nucleotides in length.
- the sequencing adaptors may be hairpin loop adaptors.
- Hairpin loop adaptors are adaptors comprising a single polynucleotide strand, wherein the ends of the polynucleotide strand are capable of hybridising to each other, or are hybridized to each other, and wherein the middle section of the polynucleotide forms a loop.
- Suitable hairpin loop adaptors can be designed using methods known in the art.
- the 3' end of a hairpin loop adaptor attaches to the 5' end of a first strand of a double-stranded polynucleotide and the 5' end of the hairpin loop adaptor attaches to the 3' end of a second strand of a double-stranded polynucleotide; or the 5' end of a hairpin loop adaptor attaches to the 3' end of a first strand of a double-stranded polynucleotide and the 3' end of the hairpin loop adaptor attaches to the 5' end of a second strand of a double-stranded polynucleotide.
- sequencing adaptors can be attached to a target polynucleotide in order to characterise the target polynucleotide.
- the sequences of the adaptors are typically not determinative and can be controlled or chosen according to the polynucleotide binding protein and other experimental conditions such as any polynucleotides to be characterised. Exemplary sequences are provided solely by way of illustration in the examples.
- the adaptors may comprise a sequence such as one or more of SEQ ID NOs: 21-26 or 28-33 in WO 2021/255476 (incorporated herein by reference in its entirety) or polynucleotide sequences having at least 20%, such as at least 30%, e.g., at least 40% such as at least 50%, e.g., at least 60% such as at least 70%, e.g., at least 80%, for example at least 90% e.g., at least 95% sequence similarity or identity to one or more of SEQ ID NOs: 21-26 or 28-33 in WO 2021/255476 (incorporated herein by reference in its entirety).
- the sequences of the adaptors can typically be altered without negatively affecting the efficacy of the method of the invention.
- Sequencing adaptors may comprise a loading site for loading the polynucleotide binding protein.
- the loading site may be for instance a single-stranded region which can targeted by the polynucleotide binding protein.
- the loading site may be a region of the sequencing adaptor to which an exogenous polynucleotide strand comprising the polynucleotide binding protein can bind in order to transfer the polynucleotide binding protein to the polynucleotide to be assessed in the method of the invention.
- the polynucleotide binding protein if present may be provided on sequencing adaptors.
- WO 2015/110813 and WO 2020/234612 describe the loading of polynucleotide binding proteins onto a target polynucleotide such as an adaptor and are hereby incorporated by reference in their entireties.
- any of the polynucleotides described herein, including the first adaptors, RT primers, second adaptors and barcoded constructs, may comprise one or more spacers, e.g., from about one to about 10 spacers, e.g., from about 1 to about 5 spacers, e.g., about 1, 2, 3, 4 or 5 spacers.
- the spacer may comprise any suitable number of spacer units.
- a spacer typically provides an energy barrier which impedes movement of a polynucleotide binding protein.
- a spacer may impede movement of a polynucleotide binding protein by reducing the traction of the protein, e.g., using an abasic spacer.
- a spacer may physically block movement of the protein, for instance by introducing a bulky chemical group to physically impede the movement of the polynucleotide binding protein.
- One or more spacers are typically included in the polynucleotide or in a sequencing adaptor to provide a distinctive signal when they pass through or across a nanopore.
- One or more spacers may be used to define or separate one or more regions of a polynucleotide, e.g., to separate an adaptor from the target polynucleotide.
- a spacer may comprise a linear molecule, such as a polymer, e.g., a polypeptide or a polyethylene glycol (PEG).
- a spacer has a different structure from the target polynucleotide. For instance, if the target polynucleotide is DNA, the or each spacer typically does not comprise DNA.
- the or each spacer preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) or a synthetic polymer with nucleotide side chains.
- PNA peptide nucleic acid
- GNA glycerol nucleic acid
- TAA threose nucleic acid
- LNA locked nucleic acid
- a spacer may comprise one or more nitroindoles, one or more inosines, one or more acridines, one or more 2- aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dTs), one or more inverted dideoxy-thymidines (ddTs), one or more dideoxy-cytidines (ddCs), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-Methyl RNA bases, one or more Isodeoxycytidines (Iso-dCs), one or more Iso-deoxyguanosines (Iso-dGs), one or more C3 (OC 3 H 6 OPO 3 ) groups, one or more photo-cleavable (PC) [OC 3 H 6 -C(O)NHCH 2 -C 6 H
- a spacer may comprise any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9 and iSpl8 spacers are all available from IDT®. A spacer may comprise any number of the above groups as spacer units.
- a spacer may comprise one or more chemical groups, e.g., one or more pendant chemical groups.
- the one or more chemical groups may be attached to one or more nucleobases in a sequencing adaptor.
- the one or more chemical groups may be attached to the backbone of a sequencing adaptor. Any number of appropriate chemical groups may be present, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more.
- Suitable groups include, but are not limited to, fluorophores, streptavidin and/or biotin, cholesterol, methylene blue, dinitrophenols (DNPs), digoxigenin and/or anti-digoxigenin and di benzylcyclooctyne groups.
- a spacer may comprise one or more abasic nucleotides (/.e., nucleotides lacking a nucleobase), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more abasic nucleotides.
- the nucleobase can be replaced by -H (idSp) or -OH in the abasic nucleotide.
- Abasic spacers can be inserted into target polynucleotides by removing the nucleobases from one or more adjacent nucleotides.
- polynucleotides may be modified to include 3- methyladenine, 7-methylguanine, l,N6-ethenoadenine inosine or hypoxanthine and the nucleobases may be removed from these nucleotides using Human Alkyladenine DNA Glycosylase (hAAG).
- polynucleotides may be modified to include uracil and the nucleobases removed with Uracil-DNA Glycosylase (UDG).
- the one or more spacers preferably do not comprise any abasic nucleotides.
- Suitable spacers can be designed or selected depending on the nature of the polynucleotide or sequencing adaptor, the polynucleotide binding protein, and the conditions under which the method is to be carried out.
- any of the polynucleotides, including the first adaptors, RT primers, second adaptors, barcoded constructs and/or sequencing adaptors, used in the invention may comprise a tag or tether.
- a polynucleotide can bind to a tag on a nanopore, e.g., via its adaptor, and release at some point, e.g., during characterization of the polynucleotide by the nanopore.
- a strong non-covalent bond e.g., biotin/avidin is still reversible and can be useful in some embodiments of the methods described herein.
- the pair of pore tag and sequencing adaptor can be configured such that the binding strength or affinity of a binding site on the polynucleotide (e.g., a binding site provided by an anchor or a leader sequence of an adaptor or by a capture sequence within the duplex stem of an adaptor) to a tag on a nanopore is sufficient to maintain the coupling between the nanopore and polynucleotide until an applied force is placed on it to release the bound polynucleotide from the nanopore.
- a binding site on the polynucleotide e.g., a binding site provided by an anchor or a leader sequence of an adaptor or by a capture sequence within the duplex stem of an adaptor
- the tags or tethers are preferably uncharged. This can ensure that the tags or tethers are not drawn into the nanopore under the influence of a potential difference.
- One or more molecules that attract or bind the polynucleotide or adaptor may be linked to the detector (e.g., the pore). Any molecule that hybridizes to the adaptor and/or target polynucleotide may be used.
- the molecule attached to the pore may be selected from a PNA tag, a PEG linker, a short oligonucleotide, a positively charged amino acid and an aptamer. Pores having such molecules linked to them are known in the art. For example, pores having short oligonucleotides attached thereto are disclosed in Howarka et al (2001) Nature Biotech.
- oligonucleotide attached to the detector e.g., a nanopore
- which oligonucleotide comprises a sequence complementary to a sequence in the leader sequence or another single stranded sequence in the adaptor may be used to enhance capture of the target polynucleotide in the methods described herein.
- the tag or tether may comprise or be an oligonucleotide (e.g., DNA, RIMA, LNA, BNA, PNA, or morpholino).
- the oligonucleotide e.g., DNA, RNA, LNA, BNA, PNA, or morpholino
- the oligonucleotide for use in the tag or tether can have at least one end (e.g., 3'- or 5'-end) modified for conjugation to other modifications or to a solid substrate surface including, e.g., a bead.
- the end modifiers may add a reactive functional group which can be used for conjugation. Examples of functional groups that can be added include, but are not limited to amino, carboxyl, thiol, maleimide, aminooxy, and any combinations thereof.
- the functional groups can be combined with different length of spacers (e.g., C3, C9, C12, Spacer 9 and 18) to add physical distance of the functional group from the end of the oligonucleotide sequence.
- the tag or tether may comprise or be a morpholino oligonucleotide.
- the morpholino oligonucleotide can have about 10-30 nucleotides in length or about 10-20 nucleotides in length.
- the morpholino oligonucleotides can be modified or unmodified.
- the morpholino oligonucleotide can be modified on the 3' and/or 5' ends of the oligonucleotides.
- modifications on the 3' and/or 5' end of the morpholino oligonucleotides include, but are not limited to 3' affinity tag and functional groups for chemical linkage (including, e.g., 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyl dithio, and any combinations thereof); 5' end modifications (including, e.g., 5'-primary ammine, and/or 5'- dabcyl), modifications for click chemistry (including, e.g., 3'-azide, 3'-alkyne, 5'-azide, 5'- alkyne), and any combinations thereof.
- 3' affinity tag and functional groups for chemical linkage including, e.g., 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyl dithio, and any combinations thereof
- 5' end modifications including, e.g.,
- the tag or tether may further comprise a polymeric linker, e.g., to facilitate coupling to a detector e.g., a nanopore.
- a polymeric linker includes, but is not limited to, polyethylene glycol (PEG).
- the polymeric linker may have a molecular weight of about 500 Da to about 10 kDa (inclusive), or about 1 kDa to about 5 kDa (inclusive).
- the polymeric linker (e.g., PEG) can be functionalized with different functional groups including, e.g., but not limited to maleimide, NHS ester, dibenzocyclooctyne (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combinations thereof.
- the tag or tether may further comprise a 1 kDa PEG with a 5'-maleimide group and a 3'-DBCO group.
- the tag or tether may further comprise a 2 kDa PEG with a 5'-maleimide group and a 3'-DBCO group.
- the tag or tether may further comprise a 3 kDa PEG with a 5'-maleimide group and a 3'-DBCO group.
- the tag or tether may further comprise a 5 kDa PEG with a 5'-maleimide group and a 3'-DBCO group.
- Other examples of a tag or tether include, but are not limited to His tags, biotin or streptavidin, antibodies that bind to analytes, aptamers that bind to analytes, analyte binding domains such as DNA binding domains (including, e.g., peptide zippers such as leucine zippers, single-stranded DNA binding proteins (SSB)), and any combinations thereof.
- the tag or tether may be attached to the external surface of a nanopore, e.g., on the cis side of a membrane, using any methods known in the art.
- one or more tags or tethers can be attached to the nanopore via one or more cysteines (cysteine linkage), one or more primary amines such as lysines, one or more non-natural amino acids, one or more histidines (His tags), one or more biotin or streptavidin, one or more antibody-based tags, one or more enzyme modification of an epitope (including, e.g., acetyl transferase), and any combinations thereof. Suitable methods for carrying out such modifications are well-known in the art.
- Suitable non-natural amino acids include, but are not limited to, 4-azido-L- phenylalanine (Faz) and any one of the amino acids numbered 1-71 in Figure 1 of Liu C. C. and Schultz P. G., Annu. Rev. Biochem., 2010, 79, 413-444.
- the one or more cysteines can be introduced to one or more monomers that form the nanopore by substitution.
- the nanopore may be chemically modified by attachment of (i) Maleimides including diabromomaleimides such as: 4-phenylazomaleinanil, l.N-(2- Hydroxyethyl)maleimide, N-Cyclohexylmaleimide, 1.3-Maleimidopropionic Acid, 1.1-4- Aminophenyl-lH-pyrrole,2,5,dione, l.l-4-Hydroxyphenyl-lH-pyrrole,2,5,dione, N- Ethylmaleimide, N-Methoxycarbonylmaleimide, N-tert-Butylmaleimide, N-(2- Aminoethyl)maleimide , 3-Maleimido-PROXYL , N-(4-Chlor
- the tag or tether may be attached directly to a nanopore or via one or more linkers.
- the tag or tether may be attached to the nanopore using the hybridization linkers described in WO 2010/086602 (incorporated herein by reference in its entirety).
- peptide linkers may be used.
- Peptide linkers are amino acid sequences. The length, flexibility and hydrophilicity of the peptide linker are typically designed such that it does not to disturb the functions of the monomer and pore.
- Preferred flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10 or 16, serine and/or glycine amino acids.
- More preferred flexible linkers include (SG)i, (SG) 2 , (SG) 3 , (SG) 4 , (SG) 5 and (SG) 8 wherein S is serine and G is glycine.
- Preferred rigid linkers are stretches of 2 to 30, such as 4, 6, 8, 16 or 24, proline amino acids. More preferred rigid linkers include (P) i2 wherein P is proline.
- Suitable pore tags are also described in WO 2018/100370, which describes non-hairpin methods for characterising double-stranded polynucleotides and is herein incorporated by reference in its entirety.
- any of the polynucleotides, including the first adaptors, RT primers, second adaptors, barcoded constructs and/or sequencing adaptors, used in the invention may comprise a membrane anchor.
- the anchor typically assists in the characterisation of a target polynucleotide in accordance with the methods disclosed herein.
- a membrane anchor may promote localisation of the selected polynucleotides around a nanopore.
- the anchor may be a polypeptide anchor and/or a hydrophobic anchor that can be inserted into the membrane.
- the hydrophobic anchor is preferably a lipid, fatty acid, sterol, carbon nanotube, polypeptide, protein, or amino acid, for example cholesterol, palmitate, or tocopherol.
- the anchor may comprise thiol, biotin, or a surfactant.
- the anchor may be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or a fusion protein), Ni-NTA (for binding to poly-histidine or poly-histidine tagged proteins) or peptides (such as an antigen).
- the anchor preferably comprises a linker, or 2, 3, 4 or more linkers.
- Preferred linkers include, but are not limited to, polymers, such as polynucleotides, polyethylene glycols (PEGs), polysaccharides and polypeptides. These linkers may be linear, branched, or circular. For instance, the linker may be a circular polynucleotide. The adaptor may hybridise to a complementary sequence on a circular polynucleotide linker.
- the one or more anchors or one or more linkers may comprise a component that can be cut or broken down, such as a restriction site or a photolabile group.
- the linker may be functionalised with maleimide groups to attach to cysteine residues in proteins. Suitable linkers are described in WO 2010/086602 (incorporated herein by reference in its entirety).
- the anchor is preferably cholesterol or a fatty acyl chain.
- any fatty acyl chain having a length of from 6 to 30 carbon atom, such as hexadecanoic acid, may be used.
- the anchor may consist or comprise a hydrophobic modification to the polynucleotide or sequencing adaptor.
- the hydrophobic modification may comprise a modified phosphate group comprised within the polynucleotide or polynucleotide anchor.
- the hydrophobic modification may for example comprise a phosphorothioate such as a charge-neutralized alkyl-phosphorothioate (PPT) as described in Jones et al, J. Am. Chem. Soc. 2021, 143, 22, 8305, the entire contents of which are hereby incorporated by reference.
- PPT charge-neutralized alkyl-phosphorothioate
- Suitable alkyl groups include for example Ci-Cio alkyl groups such as C 2 -C 6 alkyl groups, e.g., methyl, ethyl, propyl, butyl, pentyl and hexyl groups. Incorporation of the charge-neutralized alkyl- phosphorothioate into a polynucleotide allows for the polynucleotide to anchor to a hydrophobic region such as a lipid bilayer.
- Ci-Cio alkyl groups such as C 2 -C 6 alkyl groups, e.g., methyl, ethyl, propyl, butyl, pentyl and hexyl groups.
- any of the polynucleotides preferably comprise biotin.
- the biotin may be used to isolate barcoded constructs. Suitable methods for biotin-based enrichment are known in the art. For instance, a surface, such as a bead, comprising avidin/or streptavidin may be used to barcoded constructs comprising biotin. Characterising methods
- the method of the invention preferably further comprises characterising or sequencing the barcoded constructs.
- This allows the RIMA molecules to be characterised or sequenced.
- the method of the invention preferably further comprises using sequencing adaptors to characterise or sequence the barcoded constructs.
- the invention also provides methods of characterising the transcriptome of an individual cell or two or more cells in a population of cells which comprises characterising the RNA molecules from the individual cell or characterising the uniquely labelled constructs from the individual cell or two or more cells using their unique labelling.
- the RNA molecules or the constructs can be characterised using their barcoding. These methods preferably comprising using sequencing adaptors.
- the method preferably uses next generation sequencing (NGS).
- NGS next generation sequencing
- the barcoded constructs are preferably moved with respect to a detector such as a nanopore.
- the detector may be selected from (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube; and (v) a nanopore.
- the detector is a nanopore.
- the barcoded constructs may be characterised in the method of the invention in any suitable manner.
- the barcoded constructs are preferably characterised by detecting an ionic current or optical signal as they move with respect to a nanopore. This is described in more detail herein.
- the method is amenable to these and other methods of characterising polynucleotides.
- the barcoded constructs are characterised by detecting the by-products of a polynucleotide-processing reaction, such as a sequencing by synthesis reaction.
- the method may thus involve detecting the product of the sequential addition of (poly)nucleotides by an enzyme such as a polymerase to barcoded constructs.
- the product may be a change in one or more properties of the enzyme such as in the conformation of the enzyme.
- Such methods may thus comprise subjecting an enzyme such as polymerase or a reverse transcriptase to the barcoded constructs as templates under conditions such that the template-dependent incorporation of nucleotide bases into a growing oligonucleotide strand causes conformational changes in the enzyme in response to sequentially encountering template nucleic acid bases and/or incorporating template-specified natural or analog bases (/.e., an incorporation event), detecting the conformational changes in the enzyme in response to such incorporation events, and thereby detecting the sequence of the templates.
- the barcoded constructs may be moved in accordance with the method of the invention.
- Such methods may involve detecting and/or measuring incorporation events using methods known to those skilled in the art, such as those described in US 2017/0044605.
- by-products may be labelled so that a phosphate labelled species is released upon the addition of a nucleotide to a synthesised nucleic acid strand that is complementary to the template barcoded constructs, and the phosphate labelled species is detected e.g., using a detector as described herein.
- the barcoded constructs being characterised in this way may be moved in accordance with the methods herein.
- Suitable labels may be optical labels that are detected using a nanopore, or a zero-mode wave guide, or by Raman spectroscopy, or other detectors.
- Suitable labels may be non-optical labels that are detected using a nanopore, or other detectors.
- nucleoside phosphates are not labelled and upon the addition of nucleotides to synthesised nucleic acid strands that are complementary to the barcoded constructs, natural by-product species are detected.
- Suitable detectors may be ion-sensitive field-effect transistors, or other detectors.
- Any suitable measurements can be taken using a detector as the barcoded constructs move with respect to the detector.
- the barcoded constructs are preferably characterised using a nanopore.
- the method preferably comprises (i) contacting the barcoded constructs with a nanopore such that the barcoded constructs move with respect to the nanopore and (ii) taking one or more measurements as the barcoded constructs move with respect to the nanopore wherein the measurements are indicative of one or more characteristics of the barcoded constructs and thereby characterising the barcoded constructs.
- the one or more characteristics are preferably selected from (i) the length of the barcoded constructs, (ii) the identity of the barcoded constructs, (iii) the sequence of the barcoded constructs, (iv) the secondary structure of the barcoded constructs and (v) whether or not the barcoded constructs is modified.
- the barcoded constructs may be modified by methylation, by oxidation, by damage, with one or more proteins or with one or more labels, tags, or spacers.
- the one or more characteristics of the barcoded constructs are preferably measured by electrical measurement and/or optical measurement.
- the electrical measurement is preferably a current measurement, an impedance measurement, a tunnelling measurement, or a field effect transistor (FET) measurement.
- the method more preferably comprises (i) contacting the barcoded constructs with a nanopore such that the barcoded constructs move through the nanopore and (ii) measuring the current moving through the nanopore as the barcoded constructs move through the nanopore wherein the current is indicative of one or more characteristics of the barcoded constructs and thereby characterising the barcoded constructs.
- the one or more characteristics may be any of those described above.
- the movement of the barcoded constructs with respect to the nanopore or through the nanopore is preferably controlled using a polynucleotide binding protein.
- a polynucleotide binding protein The use of such proteins in nanopore sequencing is known. Examples of suitable proteins are discussed in more detail below.
- the nanopore is preferably a transmembrane pore.
- a transmembrane pore is a structure that crosses the membrane to some degree. It permits hydrated ions driven by an applied potential to flow across or within the membrane.
- the transmembrane pore typically crosses the entire membrane so that hydrated ions may flow from one side of the membrane to the other side of the membrane.
- the transmembrane pore does not have to cross the membrane. It may be closed at one end.
- the pore may be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions may flow.
- the nanopore typically has a first opening and a second opening.
- the first opening is typically the cis opening and the second opening is typically the trans opening.
- the first opening may be the trans opening and the second opening may be the cis opening.
- Any polynucleotide binding protein used in the method of the invention is typically provided at the first opening of the nanopore and thus controls the movement of the target polynucleotide in the direction from the second opening of the nanopore towards the first opening of the nanopore.
- transmembrane pore may be used in the method of the invention.
- the pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores and solid-state pores.
- the pore may be a DNA origami pore (Langecker et al., Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013/083983.
- the nanopore is preferably a transmembrane protein pore.
- a transmembrane protein pore is a polypeptide or a collection of polypeptides that permits hydrated ions, such as polynucleotide, to flow from one side of a membrane to the other side of the membrane.
- the transmembrane protein pore is capable of forming a pore that permits hydrated ions driven by an applied potential to flow from one side of the membrane to the other.
- the transmembrane protein pore preferably permits polynucleotides to flow from one side of the membrane, such as a triblock copolymer membrane, to the other.
- the transmembrane protein pore allows a polynucleotide to be moved through the pore.
- the nanopore may be a transmembrane protein pore which is a monomer or an oligomer.
- the pore is preferably made up of several repeating subunits, such as at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, or at least about 16 subunits.
- the pore is preferably a hexameric, heptameric, octameric or nonameric pore.
- the pore may be a homo-oligomer or a hetero-oligomer.
- the transmembrane protein pore may comprise a barrel or channel through which the ions may flow.
- the subunits of the pore typically surround a central axis and contribute strands to a transmembrane [3-barrel or channel or a transmembrane a-helix bundle or channel.
- the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near a constriction of the barrel or channel.
- the transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate the interaction between the pore and nucleotides, polynucleotides, or nucleic acids.
- the nanopore may be a transmembrane protein pore derived from p-barrel pores or a-helix bundle pores, p-barrel pores comprise a barrel or channel that is formed from p-strands.
- Suitable p-barrel pores include, but are not limited to, p-toxins, such as a-hemolysin, anthrax toxin and leukocidins, and outer membrane proteins/porins of bacteria, such as Mycobacterium smegmatis porin (Msp), for example MspA, MspB, MspC or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A and Neisseria autotransporter lipoprotein (NalP) and other pores, such as lysenin.
- a-helix bundle pores comprise a barrel or channel that is formed from a-helices.
- the nanopore may be a transmembrane pore derived from or based on Msp, a-hemolysin (a-HL), lysenin, CsgG, ClyA, Spl or haemolytic protein fragaceatoxin C (FraC).
- the nanopore may be a transmembrane protein pore derived from CsgG, e.g., from CsgG from E. coli Str. K-12 substr. MC4100. Such a pore is oligomeric and typically comprises 7, 8, 9 or 10 monomers derived from CsgG.
- the pore may be a homo-oligomeric pore derived from CsgG comprising identical monomers.
- the pore may be a heterooligomeric pore derived from CsgG comprising at least one monomer that differs from the others.
- Suitable pores derived from CsgG are disclosed in WO 2016/034591, WO 2017/149316, WO 2017/149317, WO 2017/149318, and WO 2019/002893 (all of which are incorporated herein by reference in their entireties).
- the nanopore may be a transmembrane pore derived from lysenin.
- suitable pores derived from lysenin are disclosed in WO 2013/153359 (incorporated herein by reference in its entirety).
- the nanopore may be a transmembrane pore derived from or based on a-hemolysin (a-HL).
- a-HL a-hemolysin
- the wild type a-hemolysin pore is formed of 7 identical monomers or sub-units (/.e., it is heptameric).
- An a-hemolysin pore may be a-hemolysin-NN or a variant thereof.
- the variant preferably comprises N residues at positions Elll and K147.
- the nanopore may be a transmembrane protein pore derived from Msp, e.g., from MspA. Examples of suitable pores derived from MspA are disclosed in WO 2012/107778 (incorporated herein by reference in its entirety).
- the nanopore may be a transmembrane pore derived from or based on ClyA.
- the detector or nanopore is typically present in a membrane. Any suitable membrane may be used.
- the membrane is preferably an amphiphilic layer.
- An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties.
- the amphiphilic molecules may be synthetic or naturally occurring.
- Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450).
- Block copolymers are polymeric materials in which two or more monomer sub-units that are polymerized together to create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit.
- a block copolymer may have unique properties that polymers formed from the individual sub-units do not possess.
- Block copolymers can be engineered such that one of the monomer sub-units is hydrophobic (/.e., lipophilic), whilst the other sub-unit(s) are hydrophilic whilst in aqueous media.
- the block copolymer may possess amphiphilic properties and may form a structure that mimics a biological membrane.
- the block copolymer may be a diblock (consisting of two monomer sub-units) but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphipiles.
- the copolymer may be a triblock, tetrablock or pentablock copolymer.
- the membrane may be a triblock copolymer membrane.
- Archaebacterial bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipid forms a monolayer membrane. These lipids are generally found in extremophiles that survive in harsh biological environments, thermophiles, halophiles and acidophiles. Their stability is believed to derive from the fused nature of the final bilayer. It is straightforward to construct block copolymer materials that mimic these biological entities by creating a triblock polymer that has the general motif hydrophilic-hydrophobic- hydrophilic.
- This material may form monomeric membranes that behave similarly to lipid bilayers and encompass a range of phase behaviours from vesicles through to laminar membranes.
- Membranes formed from these triblock copolymers hold several advantages over biological lipid membranes. Because the triblock copolymer is synthesised, the exact construction can be carefully controlled to provide the correct chain lengths and properties required to form membranes and to interact with pores and other proteins.
- Block copolymers may also be constructed from sub-units that are not classed as lipid submaterials; for example, a hydrophobic polymer may be made from siloxane or other non- hydrocarbon-based monomers.
- the hydrophilic sub-section of block copolymer can also possess low protein binding properties, which allows the creation of a membrane that is highly resistant when exposed to raw biological samples.
- This head group unit may also be derived from non-classical lipid head-groups.
- Triblock copolymer membranes also have increased mechanical and environmental stability compared with biological lipid membranes, for example a much higher operational temperature or pH range.
- the synthetic nature of the block copolymers provides a platform to customise polymer-based membranes for a wide range of applications.
- the membrane may be one of the membranes disclosed in International Application No. WO2014/064443 or WO2014/064444 (both of which are incorporated herein by reference in their entireties).
- the amphiphilic molecules may be chemically modified or functionalised to facilitate coupling of the polynucleotide.
- the amphiphilic layer may be a monolayer or a bilayer.
- the amphiphilic layer is typically planar.
- the amphiphilic layer may be curved.
- the amphiphilic layer may be supported.
- Amphiphilic membranes are typically naturally mobile, essentially acting as two-dimensional fluids with lipid diffusion rates of approximately IO -8 cm s 4 . This means that the pore and coupled polynucleotide can typically move within an amphiphilic membrane.
- the membrane may be a lipid bilayer.
- Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies.
- lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording.
- lipid bilayers can be used as biosensors to detect the presence of a range of substances.
- the lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer, or a liposome.
- the lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008/102121, WO 2009/077734, and WO 2006/100484 (incorporated herein by reference in their entireties).
- Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566).
- a lipid bilayer may be formed as described in WO 2009/077734 (incorporated herein by reference in its entirety). In this method, the lipid bilayer is formed from dried lipids. A lipid bilayer may be formed across an opening as described in W02009/077734.
- the membrane may comprise a solid-state layer.
- Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as Si 3 N 4 , A1 2 O 3 , and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses.
- the solid-state layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009/035647 (incorporated herein by reference in its entirety).
- the pore is typically present in an amphiphilic membrane or layer contained within the solid-state layer, for instance within a hole, well, gap, channel, trench or slit within the solid-state layer.
- amphiphilic membrane or layer contained within the solid-state layer for instance within a hole, well, gap, channel, trench or slit within the solid-state layer.
- suitable solid state/amphiphilic hybrid systems are disclosed in WO 2009/020682 and WO 2012/005857 (incorporated herein by reference in their entireties). Any of the amphiphilic membranes or layers discussed above may be used.
- the methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell having a pore inserted therein.
- the methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer.
- the layer may comprise other transmembrane and/or intramembrane proteins as well as other molecules in addition to the pore. Suitable apparatus and conditions are discussed below.
- the method of the invention is typically carried out in vitro.
- polynucleotide binding protein any suitable polynucleotide binding protein can be used in the methods and products of the invention.
- the polynucleotide binding protein may be any protein that is capable of binding to a polynucleotide and controlling its movement with respect to a detector, e.g., a nanopore.
- polynucleotide binding proteins such as helicases can typically control the movement of DNA in at least two active modes of operation (when is provided with all the necessary components to facilitate movement e.g., ATP and Mg 2+ ) and one inactive mode of operation (when not provided with the necessary components to facilitate movement; or when the polynucleotide binding protein is modified in order to prevent the active mode).
- active modes of operation when is provided with all the necessary components to facilitate movement e.g., ATP and Mg 2+
- inactive mode of operation when not provided with the necessary components to facilitate movement; or when the polynucleotide binding protein is modified in order to prevent the active mode.
- a polynucleotide binding protein When provided with all the necessary components to facilitate movement, a polynucleotide binding protein may move along a polynucleotide such as DNA in either a 5'-3' direction or a 3'-5' direction. Many polynucleotide binding proteins process polynucleotides such as DNA in a 5'-3' direction. Polynucleotide binding proteins which control the movement of polynucleotides in this manner are typically suitable for use in the method of the invention.
- a polynucleotide binding protein when a polynucleotide binding protein is not provided with the necessary components to facilitate movement or is modified in order to prevent it from actively controlling the movement of the polynucleotide with respect to the nanopore, it can still passively control the movement of the polynucleotide with respect to the nanopore.
- the polynucleotide binding protein can bind to the polynucleotide and act as a brake slowing the movement of the polynucleotide when it is pulled into the pore by an applied field (e.g., by the first force in the method of the invention).
- the polynucleotide binding protein may still control the movement of the polynucleotide with respect to the nanopore e.g., by acting as a brake.
- the movement control of a polynucleotide by a polynucleotide binding protein can be described in a number of ways including ratcheting, sliding, and braking.
- the method of the invention do not comprise the use of a polynucleotide binding protein operating in the passive mode.
- a polynucleotide binding protein the polynucleotide binding protein is used, it may be a polynucleotide binding protein operating in the passive mode.
- Some methods of the invention may comprise use of a polynucleotide binding protein as a pausing moiety to impede the movement of the polynucleotide strand through the nanopore.
- the polynucleotide binding protein may be a protein which binds to polynucleotides but which does not have polynucleotide processing capacity, i.e., it is not a polynucleotide binding protein.
- a polynucleotide-handling enzyme is a polypeptide that is capable of interacting with a polynucleotide.
- the enzyme may modify the polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides.
- the enzyme may modify the polynucleotide by orienting it or moving it to a specific position.
- a polynucleotide binding protein as used herein may be, or may be derived from a polynucleotide handling enzyme.
- a polynucleotide binding protein may be, or may be derived from a polynucleotide- handling enzyme.
- the polynucleotide binding protein may be derived from a member of any of the Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30 and 3.1.31.
- EC Enzyme Classification
- the polynucleotide binding protein is a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.
- the polynucleotide binding protein may be modified to prevent the polynucleotide binding protein disengaging from the polynucleotide.
- the target polynucleotide preferably does not disengage from the polynucleotide binding protein.
- the term “disengaging” refers to the dissociation of the polynucleotide binding protein from the target polynucleotide.
- a polynucleotide binding protein may be modified to prevent it from dissociating from the target polynucleotide, e.g., into the reaction medium. It is important to distinguish potential “disengagement” of a polynucleotide binding protein from “unbinding" of a polynucleotide binding protein from a target polynucleotide.
- unbinding refers to the transient release of the target polynucleotide the active site of the polynucleotide binding protein (described in more detail herein) but does not imply disengagement.
- a polynucleotide binding protein may be modified to prevent the polynucleotide binding protein from disengaging from a polynucleotide, but without preventing the polynucleotide binding protein from unbinding from the polynucleotide. When unbound, the polynucleotide binding protein remains engaged with the target polynucleotide.
- the polynucleotide binding protein may remain engaged with the target polynucleotide (/.e., it may be prevented from disengaging from the target polynucleotide) because it is topologically closed around the target polynucleotide.
- the polynucleotide binding site may remain free to bind or unbind the target polynucleotide such that the polynucleotide binding protein may bind or unbind to the target polynucleotide, whilst the polynucleotide binding protein remains engaged with the target polynucleotide.
- the polynucleotide binding protein When the polynucleotide binding protein is unbound from the target polynucleotide it may be able to move on (e.g., along) the target polynucleotide under an applied force and may be capable of re-binding to the target polynucleotide. When engaged on the target polynucleotide but unbound from the target polynucleotide, the polynucleotide binding protein is not capable of dissociating from the target polynucleotide.
- the polynucleotide binding protein can be adapted to prevent disengagement in any suitable way.
- the polynucleotide binding protein can be loaded on the polynucleotide and then modified in order to prevent it from disengaging from the polynucleotide.
- the polynucleotide binding protein can be modified to prevent it from disengaging from the polynucleotide before it is loaded onto the polynucleotide.
- Modification of a polynucleotide binding protein and/or a polynucleotide binding protein in order to prevent it from disengaging from a polynucleotide can be achieved using methods known in the art, such as those discussed in WO 2014/013260, which is hereby incorporated by reference in its entirety, and with particular reference to passages describing the modification of polynucleotide binding proteins such as helicases in order to prevent them from disengaging with polynucleotide strands.
- a polynucleotide binding protein can be modified by treating with tetramethylazodicarboxamide (TMAD).
- TMAD tetramethylazodicarboxamide
- TMAD tetramethylazodicarboxamide
- Various other closing moieties are described in WO 2021/255476 (incorporated herein by reference in its entirety).
- a polynucleotide binding protein and/or a polynucleotide binding protein may have a polynucleotide-unbinding opening, e.g., a cavity, cleft or void through which a polynucleotide strand may pass when the polynucleotide binding protein disengages from the strand.
- the polynucleotide-unbinding opening may be the opening through which a polynucleotide may pass when the polynucleotide binding protein disengages from the polynucleotide.
- the polynucleotide-unbinding opening for a given polynucleotide binding protein can be determined by reference to its structure, e.g., by reference to its X-ray crystal structure.
- the X-ray crystal structure may be obtained in the presence and/or the absence of a polynucleotide substrate.
- the location of a polynucleotide-unbinding opening in a given polynucleotide binding protein may be deduced or confirmed by molecular modelling using standard packages known in the art.
- the polynucleotide-unbinding opening may be transiently produced by movement of one or more parts e.g., one or more domains of the polynucleotide binding protein.
- the polynucleotide binding protein may be modified by closing the polynucleotide-unbinding opening.
- the polynucleotide-unbinding opening may be closed with a closing moiety.
- Closing the polynucleotide-unbinding opening may therefore prevent the polynucleotide binding protein from disengaging from the polynucleotide.
- the polynucleotide binding protein may be modified by covalently closing the polynucleotide-unbinding opening.
- closing the polynucleotide-unbinding opening does not necessarily prevent the target polynucleotide from unbinding from the polynucleotide binding site of the polynucleotide binding protein.
- a preferred protein for addressing in this way is a helicase.
- the polynucleotide binding protein may be modified with a closing moiety for (i) topologically closing the polynucleotide binding site of the polynucleotide binding protein around the target polynucleotide and (ii) promoting unbinding of the target polynucleotide from the polynucleotide binding site of the polynucleotide binding protein and/or retarding re-binding of the target polynucleotide to the polynucleotide binding site of the polynucleotide binding protein.
- the polynucleotide binding protein may be modified in any suitable manner to facilitate attachment of such a closing moiety.
- a closing moiety may comprise a bifunctional cross-linking moiety.
- the closing moiety may comprise a bifunctional cross-linker.
- the bifunctional crosslinker may attach at two points on the polynucleotide binding protein and close the polynucleotide-unbinding opening of the polynucleotide binding protein thereby preventing disengagement of the polynucleotide from the polynucleotide binding protein whilst allowing unbinding of the polynucleotide from the polynucleotide-binding site of the polynucleotide binding protein.
- the closing moiety may attach at any suitable positions on the polynucleotide binding protein.
- the closing moiety may crosslink two amino acid residues of the polynucleotide binding protein.
- at least one amino acid crosslinked by the closing moiety is a cysteine or a non-natural amino acid.
- the cysteine or non-natural amino acid may be introduced into the polynucleotide binding protein by substitution or modification of a naturally occurring amino acid residue of the polynucleotide binding protein. Methods for introducing non-natural amino acids are well known in the art and include for example native chemical ligation with synthetic polypeptide strands comprising such non-natural amino acids.
- the closing moiety may have a length of from about 1 A to about 100 A.
- the length of the closing moiety may be calculated according to static bond lengths or more preferably using molecular dynamics simulations.
- the length may for example be from about 2 A to about 80 A, such as from about 5 A to about 50 A, e.g., from about 8 to about 30 A such as from about 10 to about 25 A or about 20 A, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 A.
- polynucleotide binding proteins suitable for being closed using a closing moiety as described above are discussed in more detail herein.
- the polynucleotide binding protein is preferably a helicase, e.g., a Dda helicase as described herein.
- the polynucleotide binding protein may be or may be derived from an exonuclease.
- Suitable enzymes include, but are not limited to, exonuclease I from E. coli, exonuclease III enzyme from E. coli, RecJ from T. thermophilus and bacteriophage lambda exonuclease, TatD exonuclease and variants thereof.
- the polynucleotide binding protein may be a polymerase.
- the polymerase may be PyroPhage® 3173 DNA Polymerase (which is commercially available from Lucigen® Corporation), SD Polymerase (commercially available from Bioron®), Klenow from NEB or variants thereof.
- the enzyme is Phi29 DNA polymerase or a variant thereof. Modified versions of Phi29 polymerase that may be used in the invention are disclosed in US Patent No. 5,576,204.
- the polynucleotide binding protein may be a topoisomerase.
- the topoisomerase is a member of any of the Moiety Classification (EC) groups 5.99.1.2 and 5.99.1.3.
- the topoisomerase may be a reverse transcriptase, which are enzymes capable of catalysing the formation of cDNA from a RNA template. They are commercially available from, for instance, New England Biolabs® and Invitrogen®.
- the polynucleotide binding protein is preferably a helicase.
- Any suitable helicase can be used in accordance with the method of the invention.
- the or each enzyme used in accordance with the present disclosure may be independently selected from a Hel308 helicase, a RecD helicase, a Tral helicase, a TrwC helicase, an XPD helicase, and a Dda helicase, or a variant thereof.
- Monomeric helicases may comprise several domains attached together. For instance, Tral helicases and Tral subgroup helicases may contain two RecD helicase domains, a relaxase domain and a C-terminal domain.
- the domains typically form a monomeric helicase that is capable of functioning without forming oligomers.
- suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pifl and Tral. These helicases typically work on single stranded DNA. Examples of helicases that can move along both strands of a double stranded DNA include FtfK and hexameric enzyme complexes, or multisubunit complexes such as RecBCD.
- the polynucleotide binding protein is preferably a Dda (DNA-dependent ATPase) helicase.
- Hel308 helicases are described in publications such as WO 2013/057495, the entire contents of which are incorporated by reference.
- RecD helicases are described in publications such as WO 2013/098562, the entire contents of which are incorporated by reference.
- XPD helicases are described in publications such as WO 2013/098561, the entire contents of which are incorporated by reference.
- Dda helicases are described in publications such as WO 2015/055981 and WO 2016/055777, the entire contents of each of which are incorporated by reference.
- the helicase may be Trwc Cba or a variant thereof, Hel308 Mbu or a variant thereof or Dda or a variant thereof. Variants may differ from the native sequences in any of the ways discussed herein.
- An example variant of Dda comprises E94C/A360C.
- a further example variant of Dda comprises E94C/A360C and then (AM1)G1G2 (/.e., deletion of Ml and then addition of G1 and G2).
- the method of the invention may be operated using any suitable detector, and as such any suitable apparatus for detecting polynucleotides can be used.
- the method of the invention may be carried out using any apparatus that is suitable for nanopore sensing.
- the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections.
- the barrier may have an aperture in which a membrane containing a transmembrane pore is formed. Transmembrane pores are described herein.
- the methods may be carried out using the apparatus described in WO 2008/102120, WO 2010/122293, or WO 00/28312 (incorporated herein by reference in their entireties).
- a molecule e.g., a target polynucleotide
- Variation in the open-channel ion flow can be measured using suitable measurement techniques by the change in electrical current.
- the degree of reduction in ion flow, as measured by the reduction in electrical current is related to the size of the obstruction within, or in the vicinity of, the pore.
- Binding of a molecule of interest e.g., the target polynucleotide
- a molecule of interest e.g., the target polynucleotide
- Binding of a molecule of interest e.g., the target polynucleotide
- Detecting the presence of biological molecules finds application in personalised drug development, medicine, diagnostics, life science research, environmental monitoring and in the security and/or the defence industry.
- the presence, absence or one or more characteristics of the target polynucleotide are determined.
- the methods may be for determining the presence, absence or one or more characteristics of at least one target polynucleotide.
- the methods may concern determining the presence, absence or one or more characteristics of two or more target polynucleotide.
- the methods may comprise determining the presence, absence or one or more characteristics of any number of target polynucleotides, such as 2, 5, 10, 15, 20, 30, 40, 50, 100 or more target polynucleotides. Any number of characteristics of the one or more target polynucleotides may be determined, such as 1, 2, 3, 4, 5, 10 or more characteristics.
- Characteristics amenable to being detected in the methods include the identity or sequence of the polynucleotide, the length, of the polynucleotide, whether or not the polynucleotide is modified, etc.
- the method of the invention are methods of sequencing the barcoded constructs.
- the sequences of the barcoded constructs may be determined in real-time by aligning real-time signal or basecalling to known references. Exemplary methods of determining a polynucleotide sequence are described in WO 2016/059427 (incorporated herein by reference in its entirety).
- the methods may involve measuring the ion current flow through the pore, typically by measurement of a current.
- the ion flow through the pore may be measured optically, such as disclosed by Heron et al: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore, the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore.
- the characterisation methods may be carried out using a patch clamp or a voltage clamp. The characterisation methods preferably involve the use of a voltage clamp.
- the methods may involve measuring an optical signal as described in Chen et al, Nature Communications (2018)9: 1733, the entire contents of which are hereby incorporated by reference.
- a nanopore such as an optically engineered nanopore structure (e.g., a plasmonic nanoslit) may be used to locally enable single-molecule surface enhanced Raman spectroscopy (SERS) to allow the characterisation of the polynucleotide through direct Raman spectroscopic detection.
- SERS surface enhanced Raman spectroscopy
- the methods may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.
- the methods may involve the measuring of a current flowing through the pore.
- the method is typically carried out with a voltage applied across the membrane and pore.
- the voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV.
- the voltage used is preferably in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20mV and 0 mV and an upper limit independently selected from + 10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV.
- the voltage used is more preferably in the range 100 mV to 240mV and most preferably in the range of 120 mV to 220 mV. It is possible to increase discrimination between different nucleotides by a pore by using an increased applied potential.
- the methods are typically carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salt.
- Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3-methyl imidazolium chloride.
- the salt is present in the aqueous solution in the chamber. Potassium chloride (KCI), sodium chloride (NaCI) or caesium chloride (CsCI) is typically used. KCI is preferred.
- the salt may be an alkaline earth metal salt such as calcium chloride (CaCI2).
- the salt concentration may be at saturation.
- the salt concentration may be 3M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M.
- the salt concentration is preferably from 150 mM to 1 M.
- the method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M.
- High salt concentrations provide a high signal to noise ratio and allow for currents indicative of binding/no binding to be identified against the background of normal current fluctuations.
- the methods are typically carried out in the presence of a buffer.
- the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used.
- the buffer is HEPES.
- Another suitable buffer is Tris-HCI buffer.
- the methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5.
- the pH used is preferably about 7.5.
- the methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C.
- the methods are typically carried out at room temperature.
- the methods are optionally carried out at a temperature that supports enzyme function, such as about 37 °C.
- any of the proteins described herein, such as the protein pores, may be made synthetically or by recombinant means.
- the pore may be synthesised by in vitro translation and transcription (IVTT).
- the amino acid sequence of the pore may be modified to include non-naturally occurring amino acids or to increase the stability of the protein.
- amino acids may be introduced during production.
- the pore may also be altered following either synthetic or recombinant production.
- any of the proteins described herein, such as the protein pores, can be produced using standard methods known in the art.
- Polynucleotide sequences encoding a pore or construct may be derived and replicated using standard methods in the art.
- Polynucleotide sequences encoding a pore or construct may be expressed in a bacterial host cell using standard techniques in the art.
- the pore may be produced in a cell by in situ expression of the polypeptide from a recombinant expression vector.
- the expression vector optionally carries an inducible promoter to control the expression of the polypeptide.
- the pore may be produced in large scale following purification by any protein liquid chromatography system from protein producing organisms or after recombinant expression.
- Typical protein liquid chromatography systems include FPLC, AKTA systems, the Bio-Cad system, the Bio-Rad BioLogic system, and the Gilson HPLC system.
- Kits and Systems The invention also provides a kit for uniquely labelling RIMA molecules in a population of cells.
- the kit is preferably for producing uniquely labelled constructs transcribed from the RNA molecules in a population of cells.
- the uniquely labelled constructs transcribed from the RNA molecules are preferably uniquely labelled constructs comprising DNA sequences transcribed from the RNA molecules.
- the kit comprises two or more first adaptors, wherein the two or more first adaptors comprise primer sites for RT, and wherein each of the two or more first adaptors comprises the reverse complement of a different first barcode.
- the kit preferably comprises at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 first adaptors.
- the kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536 or at least about 3456 first
- the kit may comprise any number of first adaptors.
- the two or more first adaptors are preferably double stranded and comprise overhangs which are capable of hybridising to the 3' ends, the poly(A) tails or the polyadenylated 3' ends of the RNA molecules.
- the nucleotides at the 3' ends of one of both strands of the double stranded first adaptors are preferably ddC or inverted dT.
- the kit preferably further comprises (a) two or more RT primers capable of hybridising to the two or more first adaptors and each comprising a different second barcode and/or (b) two or more second adaptors each comprising a different third barcode.
- the kit preferably comprises at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 first adaptors.
- the kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536 or at least about 3456 second adaptors and/or RT primers.
- the kit preferably comprises the same number of first adaptors, RT primers and second adaptors. Any of the embodiments discussed above with reference to the RT primers and second adaptors used in the method of the invention equally apply to the kits. Any of the embodiments discussed above with reference to the barcodes used in the method of the invention equally apply to the kits.
- Each barcode in the kit is preferably different from any other barcode or all the other barcodes in the kit.
- the kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536 or at least about 3456 first adaptors, second adaptors and RT primers.
- the invention also provides a system for conducting the method of the invention.
- the system is for uniquely labelling RNA molecules in a population of cells.
- the system is preferably for producing uniquely labelled constructs transcribed from the RNA molecules in a population of cells.
- the uniquely labelled constructs transcribed from the RNA molecules are preferably uniquely labelled constructs comprising DNA sequences transcribed from the RNA molecules.
- the system comprises (a) two or more first adaptors, wherein the two or more first adaptors comprise primer sites for RT, and wherein each of the two or more first adaptors comprises the reverse complement of a different first barcode, (b) two or more RT primers capable of hybridising to the two or more first adaptors and each comprising a different second barcode, (c) two or more second adaptors each comprising a different third barcode and (d) a nanopore.
- the system may comprise any of the numbers of first adaptors, RT primers and second adaptors discussed above with reference to the kits of the invention.
- the system preferably comprises at same number of first adaptors, RT primers and second adaptors.
- the nanopore is preferably present in a membrane. Suitable membranes are discussed above.
- the system may comprise any of the membranes disclosed above, such as an amphiphilic layer, a triblock copolymer membrane or a solid-state layer.
- the membrane is typically part of an array of membranes, wherein each membrane preferably comprises a nanopore.
- the array may be any of those described in WO 2018/060740 (incorporated herein by reference in its entirety).
- the system is preferably adapted to apply a voltage across the membrane and to take one or more electrical measurements. Suitable adaptations are discussed in WO 2018/060740 (incorporated herein by reference in its entirety).
- the kit or system preferably further comprises one or more sequencing adaptors.
- the one or more sequencing adaptors may be any of those discussed above with reference to the method of the invention.
- the kit or system may further comprise a polynucleotide binding protein.
- the kit or system may further comprise a microparticle. Any of the embodiments discussed above with reference to the method of the invention equally apply to the system of the invention.
- the kit or system may additionally comprise one or more other reagents or instruments which enable any of the embodiments mentioned above to be carried out.
- reagents or instruments include one or more of the following: suitable buffer(s) (aqueous solutions), means to obtain a sample from a subject (such as a vessel or an instrument comprising a needle), means to amplify and/or express polynucleotides, a membrane as defined above or voltage or patch clamp apparatus.
- Reagents may be present in the kit or system in a dry state such that a fluid sample is used to resuspend the reagents.
- the kit or system may also, optionally, comprise instructions to enable the kit or system to be used in the methods described herein or details regarding for which organism the method may be used.
- the kit or system may comprise a magnet or an electromagnet.
- the kit or system may, optionally, comprise nucleotides.
- This Example describes one exemplary embodiment of the method of the invention. This is called the ligation + RT + ligation (LRL) approach.
- the LRL approach applies the first-round barcode to separate sample of cells through ligation of a barcoded cDNA reverse transcription adaptor (CDRA) ( Figure 1).
- CDRA barcoded cDNA reverse transcription adaptor
- a strand switching primer may be used which contains a UMI sequence optimised to reduce homopolymers through a TTVVVVTT structure to facilitate detection through our nanopores ( Figure 2).
- SSP strand switching primer
- the fully barcoded molecule is then enriched through biotin-streptavidin pulldown and amplified with PCR primers.
- These primers may enable the use of click chemistry moieties for rapid attachment chemistry for expedited library prep.
- barcoded universal primers for a potential fourth-round of barcoding.
- Rehydration 14 Prepare rehydration buffer according to quantity of cells:
- CDRA ligation (first-round barcoding) 20. Prepare the CDRA ligation mix as below:
- RT barcode plate Place the RT barcode plate on a chilled rack and add the following to each well in order: in situ RT 10E3 cells 31. Seal the plate and then briefly centrifuge.
- Third round blocker 1E6 solution 10E3 cells 0.5E6 cells cells 43. Add 20
- Lysis buffer cells cells 1E6 cells 54. Incubate at 55°C for 1 hour.
- NEB monarch PCR clean up kit NEB, T1030S
- 5: 1 buffer:sample ratio as described by the manufacturer.
- a cell culture stock of GM12878 (Coriel Institute, GM12878) was fixed in methanol for testing.
- ⁇ 2E7 cells GM12878 were washed in 1 mL of ice-cold IX DPBS (Gibco, 14080089) and pelleted at 300 g. The supernatant was aspirated, pellet washed and pelleted once again.
- the supernatant was aspirated, resuspending the pellet to approximately 5E3 cells/pL in ice-cold IX DPBS.
- the cell suspension was then mixed with 4x volumes ice-cold 100% methanol to achieve a final concentration of 80% methanol.
- the suspension was chilled at -20°C for 30 minutes.
- the methanol fixed cell suspension was then split into 4 equal aliquots of 5E6 fixed cell, incubated on ice for 5 minutes, then pelleted at 1000 g.
- the supernatant was aspirated, and the cell pellet was resuspended in a rehydration buffer of either 3X SSC (Invitrogen, AM9770) or IX DPBS (Gibco, 14080089) supplemented with 1 mM DTT (Sigma, 43816), 200 ng/pL recombinant albumin (NEB, B9200s), and 0.2 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 1 mL final.
- 3X SSC Invitrogen, AM9770
- IX DPBS Gibco, 14080089
- Rehydrated cells were then pelleted at 1000 g, the supernatant was aspirated.
- RNA was then immediately extracted from the cell pellet following a TRIzol reagent (Invitrogen, 15596026) based extraction following the method described by Workman et al. (2018). Briefly, 400 pL of TRIzol reagent was added to each sample, incubating for 5 minutes at room temperature. 80 pL of chloroform was then added and the sample vortexed before incubating for 5 minutes at room temperature. The sample was vortexed then pelleted at 2000 g, 4°C. The supernatant was transferred to a 5 mL centrifuge tube and mixed with an equal volume of isopropanol.
- TRIzol reagent Invitrogen, 15596026
- the tube mixed with inversion and incubated at room temperature for 15 minutes before centrifugation at 2000 g, 4°C for 20 minutes.
- the supernatant was aspirated, and the pellet washed with 400 pL of ice-cold 70% ethanol before further centrifugation at 2000 g, 4°C for 5 minutes.
- the supernatant was aspirated, and the pellet resuspended in 10 pL of TE buffer.
- RNA integrity was then assessed using the Agilent 2100 Bioanalyzer (Agilent, G2939BA) with the RNA Nano 6000 kit and chip (Agilent, 5067-1511).
- RNA integrity would remain stable during the 37°C incubation steps required for the ODRA adapter ligation and USER digestion in the first round of combinatorial barcoding in the LRL approach.
- ⁇ 2E7 cells GM12878 were washed in 1 mL of ice-cold IX DPBS (Gibco, 14080089) and pelleted at 300 g. The supernatant was aspirated, pellet washed and pelleted once again. ⁇ The supernatant was aspirated, and then the pellet resuspended to approximately 5E3 cells/pL in ice-cold IX DPBS.
- the cell suspension was then mixed with 4x volumes ice-cold 100% methanol to achieve a final concentration of 80% methanol.
- the suspension was chilled at -20°C for 30 minutes.
- Rehydrated cells were then pelleted at 1000 g, the supernatant was aspirate.
- RNA was then immediately extracted from the cell pellet following a TRIzol reagent (Invitrogen, 15596026) based extraction following the method described by Workman et al. (2018). Briefly, 400 pL of TRIzol reagent was added to each sample, incubating for 5 minutes at room temperature. 80 pL of chloroform was then added and the sample vortexed before incubating for 5 minutes at room temperature. The sample was vortexed then pelleted at 2000 g, 4°C. The supernatant was transferred to a 5 mL centrifuge tube and mixed with an equal volume of isopropanol.
- TRIzol reagent Invitrogen, 15596026
- the tube mixed with inversion and incubated at room temperature for 15 minutes before centrifugation at 2000 g, 4°C for 20 minutes.
- the supernatant was aspirated, and the pellet washed with 400 pL of ice-cold 70% ethanol before further centrifugation at 2000 g, 4°C for 5 minutes.
- the supernatant was aspirated, and the pellet resuspended in 10 pL of TE buffer.
- RNA integrity was then assessed using the Agilent 2100 Bioanalyzer (Agilent, G2939BA) with the RNA Nano 6000 kit and chip (Agilent, 5067-1511).
- EXAMPLE 4 BARCODE CROSSTALK
- ENO2 RNA analyte derived from the enolase II gene of Saccharomyces cerevisiae
- the ENO2 gene was amplified from S. cerevisiae genomic DNA using primers with a gene specific portion complimentary to the first and last 22 nt of the coding sequence and a 5' flanking portion of 27 nt of exogenic sequence for downstream application (Table 1).
- the amplicon product was then further amplified using an upstream primer flanked with a 5' T7 RNA polymerase promoter sequence and a downstream primer flanked with a 5' poly(T) tail to yield an in vitro transcription template (see Table 2).
- RNA control strand test analyte SEQ ID NO: 18
- each reaction was supplemented with 0.25 U/pL Lambda exonuclease (NEB, M0262S) and 0.05 U/pL USER mix (NEB, M5505S) final in 22 pL and incubate for 15 minutes at 37°C, then 5 minutes on ice.
- NEB Lambda exonuclease
- RNA template CDRA ligation reactions were either mixed with an equal volume of reaction without competing CDRA adapter (no competition), or with an equal volume reaction with competing CDRA adapter (competition). The pooled reactions were then incubated on ice for 5 minutes.
- each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
- AMPure XP Reagent Beckman Coulter, A63882
- 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
- each reaction was supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) then subjected to a 1.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 10 pL nuclease-free water.
- each reaction was supplemented with 0.25 U/pL Lambda exonuclease (NEB, M0262S) and 0.05 U/pL USER mix (NEB, M5505S) final in 22 pL and incubate for 15 minutes at 37°C, then 5 minutes on ice.
- NEB Lambda exonuclease
- each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
- AMPure XP Reagent Beckman Coulter, A63882
- 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
- ⁇ Reactions were then supplemented with 25 pM RT primer BC03 or BC04 (SEQ ID NO: 4or SEQ ID NO: 5 respectively), 25 pM strand switching primer (SEQ ID NO: 6), IX RT buffer (ThermoFisher Scientific, EP0753), 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694), 500 pM dNTPs (NEB, N0447S), 10 U/pL Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753) in 20 pL final and incubated for 90 minutes at 42°C, then 5 minutes on ice. Meanwhile, separate competing barcoded RT primer reactions were assembled as above without RNA template and incubated on ice.
- RNA template RT reactions were either mixed with an equal volume of reaction without competing RT primer (no competition), or with an equal volume reaction with competing RT primer (competition). The pooled reactions were then incubated on ice for 5 minutes.
- each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 20 pL of IX NEB buffer r3.1 (NEB, B6003S).
- each reaction was supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) then subjected to a 1.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 10 pL nuclease-free water.
- each reaction was supplemented with 0.25 U/pL Lambda exonuclease (NEB, M0262S) and 0.05 U/pL USER mix (NEB, M5505S) final in 22 pL and incubate for 15 minutes at 37°C, then 5 minutes on ice.
- NEB Lambda exonuclease
- each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
- AMPure XP Reagent Beckman Coulter, A63882
- 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
- ⁇ Reactions were then supplemented with 25 pM RT primer BC03 (SEQ ID NO: 4), 25 pM strand switching primer (SEQ ID NO: 6), IX RT buffer (ThermoFisher Scientific, EP0753), 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694), 500 pM dNTPs (NEB, N0447S), 10 U/pL Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753) in 20 pL final and incubated for 90 minutes at 42°C, then 5 minutes on ice. Meanwhile, separate competing barcoded RT primer reactions were assembled as above without RNA template and incubated on ice.
- each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 20 pL of IX NEB buffer r3.1 (NEB, B6003S).
- cDNA template RT reactions were either mixed with an equal volume of reaction without competing 3 rd round barcode (no competition), or with an equal volume reaction with competing 3 rd round barcode (competition). The pooled reactions were then incubated on ice for 5 minutes.
- each reaction was supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) then subjected to a 1.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 10 pL nuclease-free water.
- Figure 12 demonstrates most clearly the effectiveness of using a blocking oligo prior to sample pooling. Both the standard and alternative blocking oligo resulted in decreased incorporation of the competing barcode with diminishing PCR yields, but the alternative blocking oligo resulted in the biggest decrease of competing barcode amplicons.
Landscapes
- Chemical & Material Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Organic Chemistry (AREA)
- Analytical Chemistry (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Molecular Biology (AREA)
- Biochemistry (AREA)
- Microbiology (AREA)
- Biotechnology (AREA)
- Biophysics (AREA)
- Physics & Mathematics (AREA)
- Chemical Kinetics & Catalysis (AREA)
- Immunology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- General Engineering & Computer Science (AREA)
- General Health & Medical Sciences (AREA)
- Genetics & Genomics (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
- Preparation Of Compounds By Using Micro-Organisms (AREA)
Abstract
The invention relates to a method of uniquely labelling RNA molecules in a population of cells and kits for use in such methods. The methods and kits of the invention allow the study of the transcriptome in individual cells or populations of cells.
Description
METHOD AND KITS
TECHNICAL FIELD
The invention relates to a method of uniquely labelling RNA molecules in a population of cells and kits for use in such methods. The methods and kits of the invention allow the study of the transcriptome in individual cells or populations of cells.
INTRODUCTION
The transcriptome is the set of all coding and non-coding RNA transcripts in a single cell or a population of cells. Data concerning the transcriptome can be used to study, amongst others, cellular differentiation, carcinogenesis, transcription regulation, and biomarker discovery. Transcriptome data can also be used to identify the source of the cell or cells, including its/their phylogeny.
Advances in next generation sequencing (NGS) have allowed the transcriptome to be studied effectively. For instance, RNA-Seq uses NGS to measure the presence and amount of RNA molecules in cells and is capable of analysing the continuously changing cellular transcriptome. In addition, WO 2019/060771 describes methods of uniquely labelling or barcoding RNA molecules within each cell of a population of cells. This technique is known as "Split-Seq" and involves multiple rounds of separating (or splitting) the cells into separate samples, labelling the cells in each sample with different barcodes and then pooling the samples. However, all of the separating/splitting and pooling steps are conducted after reverse transcribing the RNA molecules in the population of cells into cDNA. The same method is described in Rosenberg et al., Science, 360 (6385): 176-182. This method may described as the RLL (RT + ligation + ligation) approach.
Biological pores (and other nanopores) have great potential as direct, electrical biosensors for polymers and a variety of small molecules. In particular, recent focus has been given to nanopores as a potential DNA sequencing technology. When a potential is applied across a nanopore, there is a change in the current flow when an analyte, such as a nucleotide, resides transiently in the barrel for a certain period of time. Nanopore detection of the nucleotide gives a current change of known signature and duration. In the strand sequencing method, a single polynucleotide strand is passed through the pore and the identities of the nucleotides are derived. Strand sequencing can involve the use of a molecular brake to control the movement of the polynucleotide through the pore.
SUMMARY OF THE INVENTION
The present inventors have identified a novel method of uniquely labelling RNA molecules in populations of cells. This approach involves multiple rounds of dividing (or splitting) the population of cells into separate samples, labelling the cells in each sample with different barcodes and then pooling the samples. In this novel method, the first round of dividing (or
splitting) the cells, labelling and pooling is conducted before reverse transcription (RT). The second dividing (or splitting step) is also conducted before RT. These method steps differentiate the novel method of the invention from Split-Seq disclosed in WO 2019/060771 and Rosenberg et al., Science, 360 (6385): 176-182. The novel method may be described as the LRL (ligation + RT + ligation) approach.
The novel method may be summarised as follows. After the first division of the population of cells, the RNA molecules in the first divided samples are labelled by first adaptors comprising primer sites for RT and the reverse complement of unique first barcodes for each first sample. The first divided samples are then pooled and divided for a second time before RT using RT primers comprising unique second barcodes for each second sample. The second divided samples and then pooled and divided a third time before labelling with second adaptors comprising unique third barcodes for each third sample. This provides triple barcoded constructs from all cells. The triple barcoded constructs comprise DNA sequences transcribed from the RNA molecules (during the second round).
The barcode used for each sample, i.e., for each first, second and third sample, is preferably different from any other barcode or all the other barcodes used in the method. The number of cells in the original population is preferably less than the number of first samples multiplied by the number of second samples multiplied the number of third samples. If the number of possible barcode combinations used is greater than the number of cells in the population, there is a high probability the RNA molecules from each cell will be labelled with a unique combination of three barcodes (/.e., the barcoded constructs produced from each cell will have a unique combination of three barcodes).
The triple barcoded constructs may then be isolated, amplified, and characterised or sequenced. This allows the transcriptome of each individual cell in the population or multiple cells in the population to be measured and/or quantified. The method may be used in combination with nanopore sequencing but does not have to be.
The method of the invention has several key advantages. For instance, by separating the first-round labelling and third-round labelling steps with a second-round RT step, the possibility of re-barcoding or crosstalk upon pooling is mitigated. This a common problem associated with RLL Split-Seq as disclosed in WO 2019/060771 and Rosenberg et al., Science, 360 (6385): 176-182 where two sequential rounds of ligation are used. Crosstalk in this instance would mean that fragments failed to be RT primed or adapter ligated in their single wells, but upon pooling are presented with surplus primer/adapter and reagents that permit further labelling of the fragments with barcodes that were not present for the previous round of labelling. This can lead to transcripts within a cell being exposed to a mixed population of barcodes, and therefore presenting as transcripts from different cells when the combinatorial barcodes are considered. With RT steps, it is also technically
possible for an RT primer to prime internally within a transcript and then allow reverse transcription to initiate and strand displace previous RT products or create truncated RT products. This would be an example of re-barcoding, where the identity of a barcode that has been incorporated could be over-written with a new one. Again this would lead to transcripts presenting as if they originated from different cells when the combinatorial barcodes are considered. These disadvantages of the RLL approach are avoided by the method of the invention.
The ligation of a first adaptor to the 3' ends of RNA transcripts in the first round enables RT of the full RNA transcript with a specific RT primer in the second round without the need of a poly(T)VN reverse transcription primer which is prone to off-target priming. This minimises intronic overlap and off-target amplification. It also allows for cDNA synthesis of full-length poly(A) tails in eukaryotic RNAs for end-to-end RT of RNA transcripts.
Part of the first adaptor can be destroyed prior to cell pooling at the end of the first round and this mitigates the possibility of re-barcoding or crosstalk upon pooling.
Additional advantages of the method of the invention are discussed below with reference to specific embodiments.
The invention provides a method of uniquely labelling RNA molecules in a population of cells, the method comprising:
(a) dividing the population of cells into a plurality of first samples;
(b) ligating first adaptors to the 3' ends of the RNA molecules in the plurality of first samples, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode;
(c) pooling the plurality of first samples and dividing the pool into a plurality of second samples;
(d) reverse transcribing the RNA molecules in the plurality of second samples using the primer sites and RT primers to form double stranded constructs, wherein the RT primers used in each second sample comprise a different second barcode;
(e) pooling the plurality of second samples and dividing the pool into a plurality of third samples; and
(f) ligating second adaptors to the double stranded constructs in the plurality of third samples to form triple barcoded constructs, wherein the second adaptors used in each third sample comprise a different third barcode.
The invention also provides a method of characterising the transcriptome of an individual cell or two or more cells in a population of cells, the method comprising:
(a) conducting a method of the invention on the population of cells; and
(b) characterising the RIMA molecules from the individual cell or two or more cells using their unique labelling.
The invention also provides a kit for uniquely labelling RNA molecules in a population of cells, the kit comprising two or more first adaptors, wherein the two or more first adaptors comprise primer sites for RT, and wherein each of the two or more first adaptors comprises the reverse complement of a different first barcode.
DESCRIPTION OF THE FIGURES
It is to be understood that Figures are for the purpose of illustrating particular embodiments of the invention only and are not intended to be limiting.
Figure 1 shows LRL first-round barcoding by ligation of the first adaptor.
Figure 2 shows LRL second-round barcoding through RT.
Figure 3 shows LRL third-round barcoding through ligation.
Figure 4 shows LRL third-round adaptor blocking.
Figure 5 shows how an SSC based rehydration buffer yields total RNA of a much higher integrity than a DPBS based rehydration buffer. The RIN (RNA integrity number) score - a measure of the ratio 28S to 18S rRNA peaks - for the SSC samples is 2-fold higher than those of the DPBS samples. RNA integrity throughout the cell fixation is crucial as the RNA transcripts must be retained for cDNA synthesis as part of the combinatorial barcoding method. This experiment demonstrates that the methanol fixation and rehydration does not impact the RNA integrity too greatly.
Figure 6 shows how RNA integrity remains unchanged with incubation at 37°C compared to a non-incubated control, both with comparable RIN scores. RNA integrity throughout the CDRA ligation and USER digest is crucial as the RNA transcripts must be retained for cDNA synthesis as part of the combinatorial barcoding method.
Figure 7 shows first round barcoding competition assay results. Sequenced reads aligned to BC01 or BC02 of the CDRA adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
Figure 8 shows second round barcoding competition assay results. Sequenced reads aligned to BC03 or BC04 of the RT barcode primer against the primary Y axis. PCR yield is aligned against the secondary Y axis.
Figure 9 shows third round barcoding competition assay results where no blocking oligo was utilised. Sequenced reads aligned to BC05 or BC06 of the third-round ligation adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
Figure 10 shows third round barcoding competition assay results where the RLL approach standard blocking oligo was utilised in the LRL approach. Sequenced reads aligned to BC05 or BC06 of the third-round ligation adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
Figure 11 shows third round barcoding competition assay results where the LRL approach alternative blocking oligo was utilised. Sequenced reads aligned to BC05 or BC06 of the third-round ligation adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
Figure 12 shows third round barcoding competition assay results. Comparison of nonbarcoded primary third round ligations competed with BC06 with either no blocking oligo, RLL approach standard blocking oligo (std), or the LRL approach alternative blocking oligo (alt) was utilised. Sequenced reads aligned to BC05 or BC06 of the third-round ligation adapter against the primary Y axis. PCR yield is aligned against the secondary Y axis.
DESCRIPTION OF THE SEQUENCE LISTING
SEQ ID NOs: 1-18 are the oligonucleotides used in the Examples.
DETAILED DESCRIPTION
It is to be understood that different applications of the disclosed products and methods may be tailored to the specific needs in the art. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting.
All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and/or take precedence over any such contradictory material.
Definitions
Where an indefinite or definite article is used when referring to a singular noun e.g., "a" or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.
"About" as used herein when referring to a measurable value such as an amount, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods.
"Nucleotide sequence", "DNA sequence" or "nucleic acid molecule(s)" as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. The term "nucleic acid" as used herein, is a single or double stranded covalently linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids may be manufactured synthetically in vitro or isolated from natural sources. Nucleic acids may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post-translational modification, for example 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and splicing. Nucleic acids may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA) and peptide nucleic acid (PNA). Sizes of nucleic acids, also referred to herein as "polynucleotides" are typically expressed as the number of base pairs (bp) or nucleotide pairs for double stranded polynucleotides, or in
the case of single stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotides in length are typically called "oligonucleotides" and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR).
The term "amino acid" in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid. The amino acids typically refer to naturally occurring L o-amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala; C=Cys; D=Asp; E=Glu; F=Phe; G=Gly; H=His; I=Ile; K=Lys; L=Leu; M = Met; N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as P-amino acids. For example, analogues or mimetics of phenylalanine or proline, which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid. Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.
The terms "polypeptide", and "peptide" are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like. A peptide can be made using recombinant techniques, e.g., through the expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide it typically substantially free of culture medium, e.g., culture medium represents less than about 20 %, more preferably less than about 10 %, and most preferably less than about 5 % of the volume of the protein preparation.
The term "protein" is used to describe a folded polypeptide having a secondary or tertiary structure. The protein may be composed of a single polypeptide or may comprise multiple
polypepties that are assembled to form a multimer. The multimer may be a homooligomer, or a heterooligmer. The protein may be a naturally occurring, or wild type protein, or a modified, or non-naturally, occurring protein. The protein may, for example, differ from a wild type protein by the addition, substitution, or deletion of one or more amino acids.
A "variant" of a protein encompasses peptides, oligopeptides, polypeptides, proteins, and enzymes having amino acid substitutions, deletions and/or insertions relative to the unmodified or wild-type protein in question and having similar biological and functional activity as the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid- by-amino acid basis over a window of comparison. Thus, a "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison (/.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.
For all aspects and embodiments of the invention, a "variant" has at least 50%, 60%, 70%, 80%, 90%, 95% or 99% complete sequence identity to the amino acid sequence of the corresponding wild-type protein. Sequence identity can also be to a fragment or portion of the full-length polynucleotide or polypeptide. Hence, a sequence may have only 50 % overall sequence identity with a full-length reference sequence, but a sequence of a particular region, domain or subunit could share 80 %, 90 %, or as much as 99 % sequence identity with the reference sequence.
The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the "normal" or "wild-type" form of the gene. In contrast, the term "modified", "mutant" or "variant" refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications and/or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For instance, methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For instance, non-naturally occurring amino acids may be
introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, they may be introduced by expressing the mutant monomer in E. coli that are auxotrophic for specific amino acids in the presence of synthetic (/.e., non-naturally occurring) analogues of those specific amino acids. They may also be produced by naked ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side-chain volume. The amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acids they replace. Alternatively, the conservative substitution may introduce another amino acid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well- known in the art and may be selected in accordance with the properties of the 20 main amino acids as defined in Table 1 below. Where amino acids have similar polarity, this can also be determined by reference to the hydropathy scale for amino acid side chains in Table 2.
Table 1 - Chemical properties of amino acids
Table 2 - Hydropathy scale
Side Chain Hydropathy
He 4.5
Vai 4.2
Leu 3.8
Phe 2.8
Cys 2.5 Met 1.9 Ala 1.8 Gly -0.4 Thr -0.7 Ser -0.8 Trp -0.9 Tyr -1.3 Pro -1.6 His -3.2
Glu -3.5 Gin -3.5 Asp -3.5 Asn -3.5 Lys -3.9 Arg -4.5
A mutant or modified protein, monomer or peptide can also be chemically modified in any way and at any site. A mutant or modified monomer or peptide is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. The mutant of modified protein, monomer or peptide may be chemically modified by the attachment of any molecule. For instance, the mutant of modified protein, monomer or peptide may be chemically modified by attachment of a dye or a fluorophore.
Methods of the invention
The invention provides a method of uniquely labelling RIMA molecules in a population of cells. The method achieves this by producing triple barcoded constructs transcribed from the RNA molecules in step (f).
In the context of the invention, uniquely labelling RNA molecules is synonymous with producing uniquely labelled constructs transcribed from the RNA molecules. The method is preferably for producing uniquely labelled constructs transcribed from the RNA molecules in a population of cells. The uniquely labelled constructs transcribed from the RNA molecules are preferably uniquely labelled constructs comprising DNA sequences transcribed from the RNA molecules. The uniquely labelled constructs are triple barcoded constructs.
The population may comprise any number of cells bearing in mind the discussion below concerning the number of possible barcode combinations. The population for use in the invention preferably comprises at least about 1.00 x 104 cells. The population more preferably comprises at least about 1.30 x 104 cells, at least about 2.00 x 104 cells, at least about 5.00 x 104 cells, at least about 1.00 x 105 cells, at least about 1.10 x 105 cells, at least about 2.00 x 105 cells, at least about 5.00 x 105 cells, at least about 6.00 x 105 cells, at least about 7.00 x 105 cells, at least about 8.00 x 105 cells, at least about 9.00 x
105 cells, at least about 1.00 x 106, at least about 2.00 x 106, at least about 5.00 x 106, at least about 1.00 x 107, at least about 2.00 x 107, at least about 5.00 x 107, at least about 1.00 x 108, at least about 2.00 x 108 cells, at least about 5.00 x 108 cells, at least about 1.00 x 109 cells, at least about 2.00 x 109 cells, at least about 3.00 x 109 cells, at least about 5.00 x 109 cells, at least about 1.00 x 1010 cells, at least about 2.00 x 1010 cells, at least about 4.00 x 1010 cells or at least about 5.00 x 1010 cells. The population may contain even more cells, such as at least about 1011 cells or at least about 1012 cells. A greater range of transcriptomes across individual cells or two or more cells will be seen with a larger number of starting cells. A greater number of possible barcode combinations is also needed.
The population preferably comprises about 5.00 x 1010 or fewer cells. The population more preferably comprises such as about 4.12 x 1010 or fewer cells, about 4.10 x 1010 or fewer cells, about 4.00 x 1010 or fewer, about 3.50 x 1010 or fewer, about 3.00 x 1010 or fewer, about 2.00 x 1010 or fewer cells, about 1.00 x 1010 or fewer cells, about 5.00 x 109 or fewer cells, about 3.62 x 109 or fewer cells, about 3.60 x 109 or fewer cells, about 3.50 x 109 or fewer cells, about 3.00 x 109 or fewer cells, about 2.00 x 109 or fewer cells, about 1.00 x 109 or fewer cells, about 5.00 x 108 or fewer cells, about 2.00 x 108 or fewer cells, about 1.00 x 108 or fewer cells, about 5.66 x 107 or fewer cells, about 5.60 x 107 or fewer cells, about 5.50 x 107 or fewer cells, about 5.00 x 107 or fewer cells, about 2.00 x 107 cells or fewer cells, about 1.00 x 107 or fewer cells, about 5.00 x 106 or fewer cells, about 2.00 x
106 or fewer cells, about 1.00 x 106 or fewer cells, about 8.84 x 105 or fewer cells, about 8.80 x 105 or fewer cells, about 8.50 x 105 or fewer cells, about 8.00 x 105 or fewer cells, about 5.00 x 105 or fewer cells, about 2.00 x 105 or fewer cells, about 1.10 x 105 or fewer cells, about 1.00 x 105 or fewer cells, about 5.00 x 104 or fewer cells, about 2.00 x 104 or fewer cells, about 1.38 x 104 or fewer cells, about 1.30 x 104 or fewer cells or about 1.00 x 104 or fewer cells.
The method may be conducted in a 24-well plate. The population preferably comprises about 1.38 x 104 or fewer cells, about 1.30 x 104 or fewer cells or about 1.00 x 104 or fewer cells.
The method may be conducted in a 48-well plate. The population preferably comprises about 1.10 x 105 or fewer cells, about 1.00 x 105 or fewer cells, about 5.00 x 104 or fewer
cells, or about 2.00 x 104 or fewer cells. The population may comprise any of the number of cells for a 24-well plate.
The method may be conducted in a 96-well plate. The population preferably comprises about 8.84 x 105 or fewer cells, about 8.80 x 105 or fewer cells, about 8.50 x 105 or fewer cells, about 8.00 x 105 or fewer cells, about 5.00 x 105 or fewer cells, or about 2.00 x 105 or fewer cells. The population may comprise any of the number of cells for a 24-well plate or 48-well plate.
The method may be conducted in a 384-well plate. The population preferably comprises about 5.66 x 107 or fewer, about 5.60 x 107 or fewer, about 5.50 x 107 or fewer, about 5.00 x 107 or fewer, about 2.00 x 107 cells or fewer, about 1.00 x 107 or fewer, about 5.00 x 106 or fewer, about 2.00 x 106 or fewer, or about 1.00 x 106 or fewer cells. The population may comprise any of the number of cells for a 24-well plate, 48-well plate, or 96-well plate.
The method may be conducted in a 1536-well plate. The population preferably comprises about 3.62 x 109 or fewer, about 3.60 x 109 or fewer, about 3.50 x 109 or fewer, about 3.00 x 109 or fewer, about 2.00 x 109 or fewer, about 1.00 x 109 or fewer, about 5.00 x 108 or fewer, about 2.00 x 108 or fewer, or about 1.00 x 108 or fewer cells. The population may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate or 384- well plate.
The method may be conducted in a 3456-well plate. The population preferably comprises about 4.12 x 1010 or fewer, about 4.10 x 1010 or fewer, about 4.00 x 1010 or fewer, about 2.00 x 1010 or fewer, about 1.00 x 1010 or fewer, or about 5.00 x 109 or fewer cells. The population may comprise any of the number of cells for a 24-well plate, 48-well plate, 96- well plate, 384-well plate, or 1536-well plate.
The cells may be any type of cells. The cells may be prokaryotic cells. The cells may be bacterial or archaeal cells. The cells are typically eukaryotic cells. The cells may be protozoan, algal, fungal, plant or animal cells. For example, the cells may be mammalian cells, such as a human, dog, cat, primate, horse, murine, rat, rodent, bovine, murine, porcine, or ovine cells. The cells are preferably human cells. The cells may be plant cells, such as cereal, legume, fruit, or vegetable cells. Examples include, but are not limited to, wheat, barley, oat, canola, maize, soya, rice, banana, apple, tomato, potato, grape, tobacco, bean, lentil, sugar cane, cocoa, cotton, tea, or coffee cells.
The animal cells may be derived from the ectoderm, endoderm, or mesoderm. The cells may be stem cells, such as embryonic stem cells, induced pluripotent stem cells or mesenchymal stem cells, bone cells, such as osteoclasts, osteoblasts or osteocytes, tendon cells, such as tenoblasts or tenocytes, chondrocytes, synovial cells, vascular cells, blood cells, such as red blood cells, immune cells, platelet, neutrophils or basophils, muscle cells,
such as skeletal muscle cells, cardiac muscle cells or smooth muscle cells, reproductive cells, such as sperm, oocytes, duct cells or epididymal cells, secretory cells, adipocytes, liver lipocytes, epithelial cells, odontoblasts, cementoblasts, hormone-secreting cells, barrier cells, exocrine secretory epithelial cells, nerve cells, astrocytes, oligodendrocytes, or neurons.
The immune cells may be neutrophil granulocyte and precursors, such as myeloblasts, promyelocytes, myelocytes, or metamyelocytes, eosinophil granulocyte and precursors, basophil granulocyte and precursors, mast cells, leukocytes, lymphocytes, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, B cells, macrophages, dendritic cells, plasma cells, neutrophils, or monocytes.
The cells may be wild-type or naturally occurring. The cells may be genetically modified or genetically engineered. For instance, the immune cells may be genetically engineered to express a recombinant chimeric antigen receptor (CAR) or T cell receptor (TCR). The cells may be genetically modified or genetically engineered using transduction or transfection or any other common techniques known to those skilled in the art.
The cells may be healthy cells or obtained from a healthy donor or source. The cells may be diseased or damaged, associated with a disease or damage or obtained from a diseased or damaged donor or source.
The cells in the population may be homologous. The cells may be the same type of cells or derived from a single donor or source. The cells in the population may be heterologous. The cells may be a mixture of different types of cells. The cells may be a mixture of different types of cells from the same donor or source. The cells in the population may be derived from different donors or sources.
Numbers of cells for use in the method of the invention and in each round are discussed in more detail below. The method typically involves the loss and/or destruction of at least some of the cells in the population. This means not all of the cells in the population are uniquely labelled using the method. The number of cells at the end of the method is typically lower than the number of cells in the population at the beginning of the method. About 80% or fewer of the cells in the population, such as about 75% or fewer, about 70% or fewer, about 65% or fewer, about 60% or fewer, about 55% or fewer, about 50% or fewer, about 45% or fewer, about 40% or fewer, about 35% or fewer, about 30% or fewer, about 25% or fewer, or about 20% or fewer of the cells in the population, are preferably present at the end of the method. Even fewer cells may be present at the end of the method.
The method typically involves at least three rounds as explained in more detail below. The number of cells at the end of each round is typically lower than the number of cells at the
beginning of each round. The number of cells at the end of the first round is typically lower than the number of cells at the beginning of the first round (/.e., than the number of cells in the population). The number of cells at the end of the second round is typically lower than the number of cells at the beginning of the second round. The number of cells at the end of the third round is typically lower than the number of cells at the beginning of the third round. About 95% or fewer of the cells at the beginning of a round, such as about 90% or fewer, about 85% or fewer, about 80% or fewer, about 75% or fewer, about 70% or fewer, about 65% or fewer or about 60% of fewer of the cells at the beginning of a round, are preferably present at the end of the round.
The method uniquely labels RIMA molecules in a population of cells. This typically means the RNA molecules from at least about 80% of the cells at the end of the method are labelled with a unique combination of barcodes. The RNA molecules from at least about 80% of the cells at the end of the method are preferably labelled with a unique combination of at least three barcodes. The RNA molecules from at least about 80% of the cells at the end of the method are preferably labelled with a unique triple barcode. In other words, the RNA molecules from at least about 80% of the cells at the end of the method are labelled with a combination of barcodes which is not present for any other cell at the end of the method. The RNA molecules from at least about 80% of the cells at the end of the method are preferably labelled with a combination of at least three barcodes which is not present for any other cell at the end of the method. The RNA molecules from at least about 80% of the cells at the end of the method are preferably labelled with a triple barcode which is not present for any other cell at the end of the method. In this paragraph, "which is not present from any other cell at the end of the method" is interchangeable with "which is not detectable from any other cell at the end of the method". The RNA molecules from at least about 85%, at least 90%, at least about 95%, at least about 97%, at least about 98% or at least about 99% of the cells at the end of the method are preferably labelled in any of these ways. The RNA molecules from about 100% of the cells at the end of the method are preferably labelled in any of these ways.
The method uniquely labels the constructs produced in step (f). This typically means the constructs from at least about 80% of the cells at the end of the method are labelled with a unique combination of barcodes. The constructs from at least about 80% of the cells at the end of the method are preferably labelled with a unique combination of at least three barcodes. The constructs from at least about 80% of the cells at the end of the method are preferably labelled with a unique triple barcode. In other words, the constructs from at least about 80% of the cells at the end of the method are labelled with a combination of barcodes which is not present for any other cell at the end of the method. The constructs from at least about 80% of the cells at the end of the method are preferably labelled with a combination of at least three barcodes which is not present for any other cell at the end of
the method. The constructs from at least about 80% of the cells at the end of the method are preferably labelled with a triple barcode which is not present for any other cell at the end of the method. In this paragraph, "which is not present for any other cell at the end of the method" is interchangeable with "which is not detectable for any other cell at the end of the method". The constricts from at least about 85%, at least 90%, at least about 95%, at least about 97%, at least about 98% or at least about 99% of the cells at the end of the method are preferably labelled in any of these ways. The constructs in about 100% of the cells at the end of the method are preferably labelled in any of these ways.
The method uniquely labels RIMA molecules in a population of cells. This typically means the RNA molecules from each cell at the end of the method are labelled with a unique combination of barcodes. The RNA molecules from each cell at the end of the method are preferably labelled with a unique combination of at least three barcodes. The RNA molecules from each cell at the end of the method are preferably labelled with a unique triple barcode. In other words, the RNA molecules from every cell at the end of the method are labelled with a combination of barcodes which is not present for any other cell at the end of the method. The RNA molecules from every cell at the end of the method are preferably labelled with a combination of at least three barcodes which is not present for any other cell at the end of the method. The RNA molecules from every cell at the end of the method are preferably labelled with a triple barcode which is not present for any other cell at the end of the method. In this paragraph, "which is not present for any other cell at the end of the method" is interchangeable with "which is not detectable for any other cell at the end of the method".
The method uniquely labels the constructs produced in step (f). This typically means the constructs from each cell at the end of the method are labelled with a unique combination of barcodes. The constructs from each cell at the end of the method are preferably labelled with a unique combination of at least three barcodes. The constructs from each cell at the end of the method are preferably labelled with a unique triple barcode. In other words, the constructs from every cell at the end of the method are labelled with a combination of barcodes which is not present for any other cell at the end of the method. The constructs from every cell at the end of the method are preferably labelled with a combination of at least three barcodes which is not present for any other cell at the end of the method. The constructs from every cell at the end of the method are preferably labelled with a triple barcode which is not present for any other cell at the end of the method. In this paragraph, "which is not present for any other cell at the end of the method" is interchangeable with "which is not detectable for any other cell at the end of the method".
At the end of step (f), the triple barcode for at least 80% of the cells is typically unique. The triple barcode for at least about 80% of the cells at the end of the method is preferably not present in any other cell at the end of the method. The triple barcode for at least about
85%, at least 90%, at least about 95%, at least about 97%, at least about 98% or at least about 99% of the cells at the end of the method are preferably unique or not present in any other cell at the end of the method. At the end of step (f), the triple barcoded constructs from every cell at the end of the method typically comprise a unique triple barcode. The triple barcoded constructs from every cell at the end of the method preferably comprise a triple barcode which is not present for any other cell at the end of the method. In this paragraph, "is not present for any other cell at the end of the method" is interchangeable with "is not detectable for any other cell at the end of the method".
Methods for achieving the uniqueness of barcode combinations for cells are discussed in more detail below.
First round
The first round of the method of the invention comprises steps (a) and (b). Step (a) comprises dividing the population of cells into a plurality of first samples. The population may be divided into any number of first samples. The population is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about
300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 first samples. The population is preferably divided into about 24, about 48, about 96, about 384, about 1536 or about 3456 first samples.
The number of cells in each first sample typically depends on the number of cells in the starting population as described above and the number of first samples. Each first sample typically comprises at least about 5.00 x 102 cells. Each first sample more preferably comprises at least about 5.70 x 102 cells, at least about 1.00 x 103 cells, at least about 2.00 x 103 cells, at least about 2.30 x 103 cells, at least about 5.00 x 103 cells, at least about 9.00 x 103 cells, at least about 9.20 x 103 cells, at least about 1.00 x 104 cells, at least about 2.00 x 104, at least about 5.00 x 104, at least about 1.00 x 105, at least about 1.40 x 105 cells, at least about 2.00 x 105, at least about 5.00 x 105, at least about 1.00 x 106, at least about 2.00 x 106, at least about 2.30 x 106 cells, at least about 5.00 x 106 cells, at least about 1.00 x 107 cells, at least about 1.10 x 107 cells, at least about 2.00 x 107 cells, at least about 3.00 x 107 cells, or at least about 5.00 x 107 cells. Each first sample may contain even more cells, such as at least about 108 cells or at least about 109 cells.
Each first sample preferably comprises about 5.00 x 107 or fewer cells. Each first sample more preferably comprises such as about 3.00 x 107 or fewer cells, about 2.00 x 107 or
fewer cells, about 1.19 x 107 or fewer cells, about 1.15 x 107 or fewer cells, about 1.10 x 107 or fewer cells, about 1.00 x 107 or fewer cells, about 5.00 x 106 or fewer cells, about 2.35 x 106 or fewer cells, about 2.30 x 106 or fewer cells, about 2.00 x 106 or fewer cells, about 1.00 x 106 or fewer cells, about 5.00 x 105 or fewer cells, about 2.00 x 105 cells or fewer cells, about 1.50 x 105 or fewer cells, about 1.47 x 105 or fewer cells, about 1.45 x 105 or fewer cells, about 1.40 x 105 or fewer cells, 1.00 x 105 or fewer cells, about 5.00 x 104 or fewer cells, about 2.00 x 104 or fewer cells, about 1.00 x 104 or fewer cells, about 9.21 x 103 or fewer cells, about 9.20 x 103 or fewer cells, about 9.00 x 103 or fewer cells, about 5.00 x 103 or fewer cells, about 2.30 x 103 or fewer cells, about 2.00 x 103 or fewer cells, about 1.00 x 103 or fewer cells, about 5.76 x 102 or fewer cells, about 5.70 x 102 or fewer cells or about 5.00 x 102 or fewer cells.
The method may be conducted in a 24-well plate. Each first sample preferably comprises about 576 or fewer cells, about 570 or fewer cells, about 550 or fewer or about 500 or fewer cells.
The method may be conducted in a 48-well plate. Each first sample preferably comprises about 2304 or fewer cells, about 2300 or fewer cells or about 2000 or fewer cells. Each first sample may comprise any of the number of cells for a 24-well plate.
The method may be conducted in a 96-well plate. Each first sample preferably comprises about 9216 or fewer cells, about 9200 or fewer cells, or about 9000 or fewer cells. Each first sample may comprise any of the number of cells for a 24-well plate or a 48-well plate.
The method may be conducted in a 384-well plate. Each first sample preferably comprises about 147,456 or fewer, about 147,000 or fewer, about 1450,000 or fewer, or about 140,000 cells or fewer. Each first sample may comprise any of the number of cells for a 24- well plate, 48-well plate, or 96-well plate.
The method may be conducted in a 1536-well plate. Each first sample preferably comprises about 2,359,296 or fewer cells, about 2,300,000 or fewer cells, or about 2,000,000 or fewer cells. Each first sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, or 384-well plate.
The method may be conducted in a 3456-well plate. Each first sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells. Each first sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, 384-well plate, or 1536-well plate.
The number of cells in each first sample is typically approximately or about the same. The skilled person will appreciate how the numbers of cells may differ between samples based
on standard techniques for measuring numbers of cells and dividing cells into different samples.
It is step (d) of the method of the invention which comprises reverse transcribing the RIMA molecules in the plurality of second samples. This is the "R" in the LRL approach of the invention. The method of the invention preferably does not comprise conducting reverse transcription (RT) before step (a). The method of the invention preferably comprises not conducting reverse transcription (RT) before step (a). The method of the invention is preferably carried out on RNA molecules which have not been reversed transcribed. The method of the invention is preferably for uniquely labelling RNA molecules which have not been reversed transcribed in a population of cells.
Step (b) comprises ligating first adaptors to the 3' ends of the RNA molecules in the plurality of first samples. Any method of ligation may be used in accordance with the invention, including the method described in the Examples. The skilled person understands the ligases that may be used in the invention, such as T4 DNA ligase and other DNA ligases. Suitable conditions for ligation reactions are known in the art and described in the Examples.
The first adaptors comprise primer sites for reverse transcription (RT). Any RT primer sites may be used in accordance with the invention, including those described in the Examples. RT primer sites are typically sequences in the first adaptors which specifically hybridise to a portion or region of the RT primers used in the second round.
The RT primer sites "specifically hybridise" to the portion or region of the RT primers when they hybridise with preferential or high affinity to the portion or region of the RT primers but do not substantially hybridise, does not hybridise, or hybridises with only low affinity to other polynucleotide sequences, especially other RNA molecules, adaptors or sequences used in the invention. Conditions that permit the hybridisation are well-known in the art (for example, Sambrook et al., 2001, Molecular Cloning: a laboratory manual, 3rd edition, Cold Spring Harbour Laboratory Press; and Current Protocols in Molecular Biology, Chapter 2, Ausubel et al., Eds., Greene Publishing and Wiley-lnterscience, New York (1995)). Hybridisation can be carried out under low stringency conditions, for example in the presence of a buffered solution of 30 to 35% formamide, 1 M NaCI and 1 % SDS (sodium dodecyl sulfate) at 37 °C followed by a 20 wash in from IX (0.1650 M Na+) to 2X (0.33 M Na+) SSC (standard sodium citrate) at 50 °C. Hybridisation can be carried out under moderate stringency conditions, for example in the presence of a buffer solution of 40 to 45% formamide, 1 M NaCI, and 1 % SDS at 37 °C, followed by a wash in from 0.5X (0.0825 M Na+) to IX (0.1650 M Na+) SSC at 55 °C. Hybridisation can be carried out under high stringency conditions, for example in the presence of a buffered solution of 50% formamide, 1 M NaCI, 1% SDS at 37 °C, followed by a wash in 0.1X (0.0165 M Na+) SSC at 60 °C. The
RT primer sites "specifically hybridise" if they hybridise to the portion or region of the RT primers with a melting temperature (Tm) that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C or at least 10 °C, greater than its Tm for other polynucleotide sequences. More preferably, the RT primer sites hybridises to the portion or region of the RT primers with a Tm that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C, greater than its Tm for other polynucleotide sequences. Preferably, the RT primer sites hybridise to the portion or region of the RT primers with a Tm that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C, greater than its Tm for a polynucleotide which differs from the portion or region of the RT primers by one or more nucleotides, such as by 1, 2, 3, 4 or 5 or more nucleotides. The RT primer sites typically hybridise to the portion or region of the RT primers with a Tm of at least 90 °C, such as at least 92 °C or at least 95 °C. Tm can be measured experimentally using known techniques, including the use of DNA microarrays, or can be calculated using publicly available Tm calculators, such as those available over the internet.
The RT primers sites typically comprise or consist of a sequence at least about 80% identical or homologous to the reverse complement of the portion or region of the RT primers. The RT primers sites preferably comprise or consist of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of the portion or region of the RT primers. The RT primers sites most preferably comprise or consist of a sequence which is the reverse complement of the portion or region of the RT primers. By complementary, it is meant the RT primer sites comprise or consist of a sequence 100% identical or homologous to the reverse complement of the portion or region of the RT primers. Complementarity is typically determined using canonical Watson-Crick base pairing.
Standard methods in the art may be used to determine homology. For example, the UWGCG Package provides the BESTFIT program which can be used to calculate homology, for example used on its default settings (Devereux et al (1984) Nucleic Acids Research 12, p387-395). The PILEUP and BLAST algorithms can be used to calculate homology or line up sequences (such as identifying equivalent residues or corresponding sequences (typically on their default settings)), for example as described in Altschul S. F. (1993) J Mol Evol 36:290- 300; Altschul, S.F et al (1990) J Mol Biol 215:403-10. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http://www.ncbi.nlm.nih.gov/).
The RT primer sites and/or the portion or region of the RT primers may be any length, such as at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at
least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 21 nucleotides, at least about 25 nucleotides or at least about 25 nucleotides in length. The RT primer sites and the portion or region of the RT primers are preferably the same length.
The first adaptors used in each sample may use the same RT primers sites. This allows the method to use a standard set of first adaptors and RT primers in which only the first and second barcodes differ between samples.
The first adaptors used in each first sample comprise the reverse complement of a different first barcode. The first adaptors comprise the reverse complement of different barcodes because the reverse complement is reverse transcribed in the second round to create the first different barcodes. The first adaptors used in each first sample comprise the reverse complement of a unique first barcode. The RNA molecules in each first sample are labelled with a different or unique first barcode. The first adaptors used in each first sample comprise the reverse complement of a first barcode that is different from the first barcode used in all the other first samples. The RNA molecules in each first sample are labelled with a first barcode which is different from the first barcode used to label the RNA molecules in all the other first samples. The first adaptors used in each first sample comprise the reverse complement of a first barcode that is different from any other barcode or all the other barcodes used in the method. The RNA molecules in each first sample are labelled with a first barcode which is different from any other barcode or all the other barcodes used in the method. Barcodes are discussed in more detail below.
The first adaptors are typically polynucleotides. The first adaptors may be any type of polynucleotide. A polynucleotide, such as a nucleic acid, is a macromolecule comprising two or more nucleotides. A polynucleotide can be single-stranded or double-stranded. A doublestranded polynucleotide is made of two single stranded polynucleotides hybridised together. The first adaptor can be a single-stranded polynucleotide or a double-stranded polynucleotide.
A polynucleotide may comprise any combination of any nucleotides. The nucleotides can be naturally occurring or artificial. A nucleotide typically contains a nucleobase, a sugar and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine (A), guanine (G), thymine (T), uracil (U) and cytosine (C).
The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably a deoxyribose. The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and/or thymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC).
The nucleotide is typically a ribonucleotide or deoxyribonucleotide. The nucleotide typically contains a monophosphate, diphosphate, or triphosphate. The nucleotide may comprise more than three phosphates, such as 4 or 5 phosphates. Phosphates may be attached on the 5' or 3' side of a nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5- hydroxy methylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP and dUMP.
A nucleotide may be abasic (/.e., lack a nucleobase). A nucleotide may also lack a nucleobase and a sugar (/.e., is a C3 spacer).
The nucleotides in the polynucleotide may be attached to each other in any manner. The nucleotides are typically attached by their sugar and phosphate groups as in nucleic acids. The nucleotides may be connected via their nucleobases as in pyrimidine dimers.
The polynucleotide can be a nucleic acid, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). The polynucleotide can comprise one strand of RNA hybridized to one strand of DNA. The polynucleotide may be any synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), bridged nucleic acid (BNA) or other synthetic polymers with nucleotide side chains. The PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone is composed of repeating glycol units linked by phosphodiester bonds. The TNA backbone is composed of repeating threose sugars linked together by phosphodiester bonds. LNA is formed from ribonucleotides as discussed above having an extra bridge connecting the 2’ oxygen and 4’ carbon in the ribose moiety.
The polynucleotide is preferably DNA, RNA or a DNA or RNA hybrid, most preferably DNA. A DNA/RNA hybrid may comprise DNA and RNA on the same strand. Preferably, the DNA/RNA hybrid comprises one DNA strand hybridized to an RNA strand.
The backbone of the polynucleotide can be altered to reduce the possibility of strand scission. For example, DNA is known to be more stable than RNA under many conditions. The backbone of the polynucleotide strand can be modified to avoid damage caused by e.g., harsh chemicals such as free radicals. DNA or RNA that contains unnatural or modified bases can be produced by amplifying natural DNA or RNA polynucleotides in the presence of modified NTPs using an appropriate polymerase.
The first adaptors are typically synthetic or semi-synthetic. For example, DNA or RIMA may be purely synthetic, synthesised by conventional DNA synthesis methods such as phosphoramidite based chemistries. Synthetic polynucleotides subunits may be joined together by known means, such as ligation or chemical linkage, to produce longer strands. Internal self-forming structures (e.g., hairpins, quadruplexes) can be designed into the substrate, e.g., by ligating appropriate sequences. Synthetic polynucleotides can be copied and scaled up for production by means known in the art, including PCR, incorporation into bacterial factories, and the like.
The first adaptors can be any length. For example, the first adaptors can be at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 56 nucleotides, at least about 60, or at least 60 nucleotides or nucleotide pairs in length.
The first adaptors are preferably double stranded. The first adaptors are preferably double stranded DNA. The first adaptors may be called barcoded cDNA reverse transcription adaptors (barcoded CDRAs).
The first adaptors are preferably double stranded and comprise overhangs which are capable of hybridising or specifically hybridising to the 3' ends of the RNA molecules. The non-overhanging strands of such adaptors are typically ligated to the RNA molecules. The skilled person is capable of designing overhangs capable of hybridising or specifically hybridising to the 3' ends of the RNA molecules. For instance, the first adaptors used in each first sample are preferably double stranded and comprise overhangs comprising every possible combination of sequences based on nucleotides comprising adenosine (A), thymine (T), uracil (U), guanine (G) and cytosine (C). Such overhangs are capable of hybridising or specifically hybridising to the 3' ends of the RNA molecules. First adaptors comprising overhangs can be produced randomly to include the combinations of sequences. Step (a) preferably comprises hybridising or specifically hybridising overhangs of double stranded first adaptors to the 3' ends of the RNA molecules in the plurality of first samples, ligating the non-overhanging strands of the first adaptors to the 3' ends of the RNA molecules, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode.
Eukaryotic RNA molecules typically have poly(A) tails (see Figure 1). The first adaptors are preferably double stranded and comprise overhangs which are capable of hybridising or specifically hybridising to the poly(A) tails of the RNA molecules. The non-overhanging strands of such adaptors are typically ligated to the RNA molecules. An example of this is shown in Figure 1. Step (a) preferably comprises hybridising or specifically hybridising overhangs of double stranded first adaptors to poly(A) tails of the RNA molecules in the
plurality of first samples, ligating the non-overhanging strands of the first adaptors to the poly(A) tails of the RIMA molecules, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode.
Non-eukaryotic RNA molecules typically do not have poly(A) tails. Step (b) may further comprise polyadenylating the 3' ends of the RNA molecules and using double stranded first adaptors which comprise overhangs which are capable of hybridising or specifically hybridising to the polyadenylated 3' ends of the RNA molecules. The non-overhanging strands of such adaptors are typically ligated to the RNA molecules. Step (a) may comprise polyadenylating the 3' ends of the RNA molecules, hybridising or specifically hybridising overhangs of double stranded first adaptors to polyadenylated 3' ends of the RNA molecules in the plurality of first samples, ligating the non-overhanging strands of the first adaptors to the polyadenylated 3' ends of the RNA molecules, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode. The RNA molecules may be polyadenylated using any known method, such as using E. coli Poly(A) polymerase (EPAP).
The overhangs can be any length. For example, the overhangs can be at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 15, at least about 20, or at least 25 nucleotides in length. The overhangs are preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length. Hybridisation and specific hybridisation are discussed above and any of those embodiments apply to the overhangs to the 3' ends, the poly(A) tails or the polyadenylated 3' ends.
The overhangs preferably comprise one or more of (i) thymine-containing nucleotides, (ii) uracil-containing nucleotides and (iii) universal nucleotides. Overhangs comprising (i) and (ii) are capable of specifically hybridising to the poly(A) tails or polyadenylated 3' ends as described above. Nucleotides are defined above. A universal nucleotide is one which will hybridise or bind to some degree to all of the nucleotides in a polynucleotide. A universal nucleotide is preferably one which will hybridise or bind to some degree to nucleotides comprising the nucleosides adenosine (A), thymine (T), uracil (U), guanine (G) and cytosine (C). The universal nucleotide may hybridise or bind more strongly to some nucleotides than to others. For instance, a universal nucleotide (I) comprising the nucleoside, 2'- deoxyinosine, will show a preferential order of pairing of I-C>I-A>I-G approximately=I-T.
The universal nucleotide preferably comprises one of the following nucleobases: hypoxanthine, 4-nitroindole, 5-nitroindole, 6-nitroindole, formylindole, 3-nitropyrrole,
nitroimidazole, 4-nitropyrazole, 4-nitrobenzimidazole, 5-nitroindazole, 4- aminobenzimidazole or phenyl (C6-aromatic ring). The universal nucleotide more preferably comprises one of the following nucleosides: 2'-deoxyinosine, inosine, 7-deaza-2'- deoxyinosine, 7-deaza-inosine, 2-aza-deoxyinosine, 2-aza-inosine, 2-O'-methylinosine, 4- nitroindole 2'-deoxyribonucleoside, 4-nitroindole ribonucleoside, 5-nitroindole 2'- deoxyribonucleoside, 5-nitroindole ribonucleoside, 6-nitroindole 2'-deoxyribonucleoside, 6- nitroindole ribonucleoside, 3-nitropyrrole 2'-deoxyribonucleoside, 3-nitropyrrole ribonucleoside, an acyclic sugar analogue of hypoxanthine, nitroimidazole 2'- deoxyribonucleoside, nitroimidazole ribonucleoside, 4-nitropyrazole 2'-deoxyribonucleoside, 4-nitropyrazole ribonucleoside, 4-nitrobenzimidazole 2'-deoxyribonucleoside, 4- nitrobenzimidazole ribonucleoside, 5-nitroindazole 2'-deoxyribonucleoside, 5-nitroindazole ribonucleoside, 4-aminobenzimidazole 2'-deoxyribonucleoside, 4-aminobenzimidazole ribonucleoside, phenyl C-ribonucleoside, phenyl C-2'-deoxyribosyl nucleoside, 2'- deoxynebularine, 2'-deoxyisoguanosine, K-2'-deoxyribose, P-2'-deoxyribose and pyrrolidine. The universal nucleotide more preferably comprises 2'-deoxyinosine. The universal nucleotide is more preferably IMP or dIMP. The universal nucleotide is most preferably dPMP (2'-Deoxy-P-nucleoside monophosphate) or dKMP (N6-methoxy-2, 6- diaminopurine monophosphate).
The overhangs preferably comprise (in the 5' to 3' direction) one or more consecutive repeating units of UTT, such as 2 or more, 3 or more, 4 or more or 5 more consecutive repeating units of UTT.
The first adaptors used in each sample may use the same overhangs. This allows the method to use a standard set of first adaptors in which only the first barcodes differ between first samples.
If the first adaptors are double stranded, the reverse complements of the different first barcodes are preferably present in the strands of the first adaptors which are ligated to the RIMA molecules. The reverse complements of the different first barcodes are preferably present in the opposite strands from the strands with the overhangs which hybridise or specifically hybridise to the 3' ends, the poly(A) tails or the polyadenylated 3' ends of the RNA molecules. If the reverse complements of the different first barcodes are ligated to the RNA molecules, the different first barcodes will be present in the cDNA product of the reverse transcription in the second round (see for example Figure 2).
The nucleotides at the 3' ends of the first adaptors are preferably dideoxy cytosine (ddC) or inverted thymidine (inverted dT). The nucleotides at the 3' ends of one of both strands of the double stranded first adaptors are preferably ddC or inverted dT. The nucleotides at the 3' ends of the strand of the double stranded first adaptors are preferably ddC or inverted dT. These nucleotides prevent DNA polymerase from further extending the cDNA sequence.
They also protect the oligonucleotide from 3' exonuclease cleavage. The use of these types of modifications in adaptors, especially on strands that are not directly incorporated into cDNA synthesised in the second round (see, for examples, Figures 1 and 2), prevents unintended crosstalk and miss-priming during amplification of the cDNA.
Exemplary first adaptors are shown in the Examples.
The non-ligated strands of the first adaptors are preferably removed before step (d). If the first adaptors are double stranded and comprise overhangs which are capable of hybridising or specifically hybridising to the 3' ends, the poly(A) tails or the polyadenylated 3' ends of the RIMA molecules, the overhanging strands of the adaptors are preferably removed before step (d). The non-ligated/overhanging strands may be removed using any suitable method, including digestion. The strands are preferably removed using USER digestion. USER enzyme mix is a blend of Uracil DNA glycosylase (UDG) and the DNA glycosylase-lyase Endonuclease VIII. As explained above, this mitigates the possibility of re-barcoding or crosstalk upon pooling.
Second round
The second round of the method of the invention comprises steps (c) and (d). Step (c) comprises pooling the plurality of first samples and dividing the pool of cells into a plurality of second samples. The pool may be divided into any number of second samples. The pool is preferably divided into any number of second samples discussed above with respect to the first samples. The pool is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 second samples. Each second sample is preferably divided into about 24, about 48, about 96, about 384, about 1536 or about 3456 second samples. The number of second samples is typically identical to the number of first samples.
Each second sample typically comprises at least about 5.00 x 102 cells. Each second sample more preferably comprises at least about 5.70 x 102 cells, at least about 1.00 x 103 cells, at least about 2.00 x 103 cells, at least about 2.30 x 103 cells, at least about 5.00 x 103 cells, at least about 9.00 x 103 cells, at least about 9.20 x 103 cells, at least about 1.00 x 104 cells, at least about 2.00 x 104, at least about 5.00 x 104, at least about 1.00 x 105, at least about 1.40 x 105 cells, at least about 2.00 x 105, at least about 5.00 x 105, at least about 1.00 x 106, at least about 2.00 x 106, at least about 2.30 x 106 cells, at least about
5.00 x 106 cells, at least about 1.00 x 107 cells, at least about 1.10 x 107 cells, at least about 2.00 x 107 cells, at least about 3.00 x 107 cells, or at least about 5.00 x 107 cells. Each second sample may contain even more cells, such as at least about 108 cells or at least about 109 cells.
Each second sample preferably comprises about 5.00 x 107 or fewer cells. Each second sample more preferably comprises such as about 3.00 x 107 or fewer cells, about 2.00 x 107 or fewer cells, about 1.19 x 107 or fewer cells, about 1.15 x 107 or fewer cells, about 1.10 x 107 or fewer cells, about 1.00 x 107 or fewer cells, about 5.00 x 106 or fewer cells, about 2.35 x 106 or fewer cells, about 2.30 x 106 or fewer cells, about 2.00 x 106 or fewer cells, about 1.00 x 106 or fewer cells, about 5.00 x 105 or fewer cells, about 2.00 x 105 cells or fewer cells, about 1.50 x 105 or fewer cells, about 1.47 x 105 or fewer cells, about 1.45 x 105 or fewer cells, about 1.40 x 105 or fewer cells, 1.00 x 105 or fewer cells, about 5.00 x 104 or fewer cells, about 2.00 x 104 or fewer cells, about 1.00 x 104 or fewer cells, about 9.21 x 103 or fewer cells, about 9.20 x 103 or fewer cells, about 9.00 x 103 or fewer cells, about 5.00 x 103 or fewer cells, about 2.30 x 103 or fewer cells, about 2.00 x 103 or fewer cells, about 1.00 x 103 or fewer cells, about 5.76 x 102 or fewer cells, about 5.70 x 102 or fewer cells or about 5.00 x 102 or fewer cells.
The method may be conducted in a 24-well plate. Each second sample preferably comprises about 576 or fewer cells, about 570 or fewer cells, about 550 or fewer or about 500 or fewer cells.
The method may be conducted in a 48-well plate. Each second sample preferably comprises about 2304 or fewer cells, about 2300 or fewer cells or about 2000 or fewer cells. Each second sample may comprise any of the number of cells for a 24-well plate.
The method may be conducted in a 96-well plate. Each second sample preferably comprises about 9216 or fewer cells, about 9200 or fewer cells, or about 9000 or fewer cells. Each second sample may comprise any of the number of cells for a 24-well plate or a 48-well plate.
The method may be conducted in a 384-well plate. Each second sample preferably comprises about 147,456 or fewer, about 147,000 or fewer, about 1450,000 or fewer, or about 140,000 cells or fewer. Each second sample may comprise any of the number of cells for a 24-well plate, 48-well plate, or 96-well plate.
The method may be conducted in a 1536-well plate. Each second sample preferably comprises about 2,359,296 or fewer cells, about 2,300,000 or fewer cells, or about 2,000,000 or fewer cells. Each second sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, or 384-well plate.
The method may be conducted in a 3456-well plate. Each second sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells. Each second sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, 384-well plate, or 1536-well plate.
The number of cells in each second sample is typically approximately or about the same. The number of cells is each second sample is typically approximately or about the same as the number of cells in each first sample. The skilled person will appreciate how the numbers of cells may differ between samples based on standard techniques for measuring numbers of cells and dividing cells into different samples.
Step (d) comprises reverse transcribing the RIMA molecules in the plurality of second samples. Methods for conducting reverse transcription are well known in the art and any suitable conditions may be used. Preferred conditions are described in the Examples.
Reverse transcription involves the use of the enzyme reverse transcriptase to convert RNA into cDNA Reverse transcriptases are commercially available (e.g. Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753), Superscript® II reverse transcriptase (Invitrogen) and Affinity script (Agilent)). The reverse transcriptase is preferably Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753). The doubles stranded constructs produced in step (b) typically comprise the RNA molecules hybridised to cDNA strands produced by reverse transcription.
Step (d) uses the primer sites in the first adaptors and RT primers to form double stranded constructs. The skilled person is capable of designing suitable RT primers for use in the invention. The RT primers are typically polynucleotide RT primers and may be any of the polynucleotides discussed above with reference to the first adaptors. The RT primers are typically synthetic or semi-synthetic.
The RT primers are preferably single stranded polynucleotide RT primers. The RT primers may be any length. The RT primers are preferably at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 32 nucleotides, at least about 35 nucleotides, or at least about 40 nucleotides in length. The RT primers are preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length.
Step (d) preferably comprises hybridising the RT primers to the first adaptors. Step (d) preferably comprises hybridising the RT primers to the RT primers sites in the first adaptors. Step (d) preferably comprises hybridising RT primers to the first adaptors or to the RT primers sites in the first adaptors and reverse transcribing the RNA molecules in the
plurality of second samples using the primer sites and RT primers to form double stranded constructs. Hybridisation is preferably specific hybridisation. As explained above, the RT primer sites are typically sequences in the first adaptors which specifically hybridise to a portion or region of the RT primers used in the second round. RT primer sites and portions or regions of the RT primers are discussed above with reference to the first round. The RT primers may comprise any of the portions or regions discussed above. The RT primers may use the same portion or region which specifically hybridises to the RT primer sites. This allows the method to use a standard set of RT primers in which only the second barcodes differ between second samples.
The RT primers used in each second sample comprise a different second barcode. The RT primers used in each second sample comprise a unique second barcode. The RNA molecules in each second sample are labelled with a different or unique second barcode. The RT primers used in each second sample comprise a second barcode that is different from the second barcode used in all the other second samples. The RNA molecules in each second sample are labelled with a second barcode which is different from the second barcode used to label the RNA molecules in all the other second samples. The RT primers used in each second sample comprise a second barcode that is different from any other barcode or all the other barcodes used in the method. The RNA molecules in each second sample are labelled with a second barcode which is different from any other barcode or all the other barcodes used in the method. Barcodes are discussed in more detail below.
The RT primers preferably comprise sequences at or near their 5' ends which are capable of specifically hybridising to the second adaptors used in step (f). The sequences preferably specifically hybridise to overhangs in the second adaptors as discussed in more detail below. Specific hybridisation is defined above. The sequences may be any length. The sequences preferably at least about 5 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 12 nucleotides, at least about 15 nucleotides, or at least about 20 nucleotides in length. The sequences are preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length. The RT primers may use the same sequences at or near their 5' ends which are capable of specifically hybridising to the second adaptors used in step (f). This allows the method to use a standard set of RT primers in which only the second barcodes differ between second samples.
Exemplary RT primers are shown in the Examples.
The RT in step (d) preferably produces strands comprising the different first barcodes, the different second barcodes and sequences complementary to the RNA molecules. The different first barcodes are transcribed from the reverse complements in the first adaptors.
An example of this is shown in Figure 2. The components are preferably in the following order in the strands 5' to 3': different second barcodes, different first barcodes and complementary sequences. These strand constructs are double barcoded.
Step (d) preferably further comprises (i) hybridising strand switching primers (SSPs) to overhangs at the 3' end of the strands in the double stranded constructs produced by reverse transcription of the RIMA molecules and (ii) reverse transcribing the SSPs to elongate the double stranded constructs. An example of this is shown in Figure 2. The SSPs are preferably specifically hybridised to the overhangs. The strands in the double constructs produced by reverse transcription are preferably cDNA strands. Overhangs, hybridisation, and specific hybridisation are discussed above and any of those embodiments apply to the SSP embodiments. The overhangs preferably comprise consecutive cytosine-containing nucleotides, such as consecutive deoxycytidine (dC)-containing nucleotides. Such overhangs can be created through the addition of non-template dC to the synthesised cDNA by reverse transcriptase when it reaches the end of an RNA template. The SSPs may be any of the types of polynucleotides and/or any of the lengths discussed above with reference to the RT primers used in the invention. The SSPs preferably comprise one or more unique molecular identifiers (UMIs) to facilitate characterisation. The SSPs may comprise any number of UMIs, such as 2 or more, 3 or more, 4 or more, 5 or more or 10 or more. An example of a UMI is TTVVVVTT. This UMI structure is optimised for characterisation using nanopores by reducing homopolymers. The SSPs facilitate deduplication of PCR replicates. By including one or more UMIs in the SSPs, they can be removed from the second adapters used in the third round (unlike in the RLL approach) to save on production cost and complexity of the second adapters. Any method of reverse transcription, including any of those discussed above, may be used in step (ii).
Exemplary SSPs are shown in the Examples.
Third round
The third round of the method of the invention comprises steps (e) and (f). Step (e) comprises pooling the plurality of second samples and dividing the pool of cells into a plurality of third samples. The pool may be divided into any number of third samples. The pool is preferably divided into any number of third samples discussed above with respect to the first samples. The pool is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 third samples. Each
third sample is preferably divided into about 24, about 48, about 96, about 384, about 1536 or about 3456 third samples. The number of third samples is typically identical to the number of first and/or second samples. The number of first samples, the number of second samples and the number of third samples is preferably the same.
Each third sample typically comprises at least about 5.00 x 102 cells. Each third sample more preferably comprises at least about 5.70 x 102 cells, at least about 1.00 x 103 cells, at least about 2.00 x 103 cells, at least about 2.30 x 103 cells, at least about 5.00 x 103 cells, at least about 9.00 x 103 cells, at least about 9.20 x 103 cells, at least about 1.00 x
104 cells, at least about 2.00 x 104, at least about 5.00 x 104, at least about 1.00 x 105, at least about 1.40 x 105 cells, at least about 2.00 x 105, at least about 5.00 x 105, at least about 1.00 x 106, at least about 2.00 x 106, at least about 2.30 x 106 cells, at least about 5.00 x 106 cells, at least about 1.00 x 107 cells, at least about 1.10 x 107 cells, at least about 2.00 x 107 cells, at least about 3.00 x 107 cells, or at least about 5.00 x 107 cells. Each third sample may contain even more cells, such as at least about 108 cells or at least about 109 cells.
Each third sample preferably comprises about 5.00 x 107 or fewer cells. Each third sample more preferably comprises such as about 3.00 x 107 or fewer cells, about 2.00 x 107 or fewer cells, about 1.19 x 107 or fewer cells, about 1.15 x 107 or fewer cells, about 1.10 x 107 or fewer cells, about 1.00 x 107 or fewer cells, about 5.00 x 106 or fewer cells, about 2.35 x 106 or fewer cells, about 2.30 x 106 or fewer cells, about 2.00 x 106 or fewer cells, about 1.00 x 106 or fewer cells, about 5.00 x 105 or fewer cells, about 2.00 x 105 cells or fewer cells, about 1.50 x 105 or fewer cells, about 1.47 x 105 or fewer cells, about 1.45 x
105 or fewer cells, about 1.40 x 105 or fewer cells, 1.00 x 105 or fewer cells, about 5.00 x 104 or fewer cells, about 2.00 x 104 or fewer cells, about 1.00 x 104 or fewer cells, about 9.21 x 103 or fewer cells, about 9.20 x 103 or fewer cells, about 9.00 x 103 or fewer cells, about 5.00 x 103 or fewer cells, about 2.30 x 103 or fewer cells, about 2.00 x 103 or fewer cells, about 1.00 x 103 or fewer cells, about 5.76 x 102 or fewer cells, about 5.70 x 102 or fewer cells or about 5.00 x 102 or fewer cells.
The method may be conducted in a 24-well plate. Each third sample preferably comprises about 576 or fewer cells, about 570 or fewer cells, about 550 or fewer or about 500 or fewer cells.
The method may be conducted in a 48-well plate. Each third sample preferably comprises about 2304 or fewer cells, about 2300 or fewer cells or about 2000 or fewer cells. Each third sample may comprise any of the number of cells for a 24-well plate.
The method may be conducted in a 96-well plate. Each third sample preferably comprises about 9216 or fewer cells, about 9200 or fewer cells, or about 9000 or fewer cells. Each third sample may comprise any of the number of cells for a 24-well plate or a 48-well plate.
The method may be conducted in a 384-well plate. Each third sample preferably comprises about 147,456 or fewer, about 147,000 or fewer, about 1450,000 or fewer, or about 140,000 cells or fewer. Each third sample may comprise any of the number of cells for a 24- well plate, 48-well plate, or 96-well plate.
The method may be conducted in a 1536-well plate. Each third sample preferably comprises about 2,359,296 or fewer cells, about 2,300,000 or fewer cells, or about 2,000,000 or fewer cells. Each third sample may comprise any of the number of cells for a 24-well plate, 48- well plate, 96-well plate, or 384-well plate.
The method may be conducted in a 3456-well plate. Each third sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells. Each third sample may comprise any of the number of cells for a 24-well plate, 48-well plate, 96-well plate, 384-well plate, or 1536-well plate.
The number of cells in each third sample is typically approximately or about the same. The number of cells is each third sample is typically approximately about the same as the number of cells in each first sample and/or the same as the number of cells in each second sample. The number of cells is each first sample, in each second sample and in each third sample is typically approximately or about the same. The skilled person will appreciate how the numbers of cells may differ between samples based on standard techniques for measuring numbers of cells and dividing cells into different samples.
Step (f) comprises ligating second adaptors to the double stranded constructs in the plurality of third samples to form triple barcoded constructs. The triple barcoded constructs comprise DNA sequences transcribed from the RIMA molecules. Ligation is discussed above and any of the methods above may be used. The skilled person is capable of designing suitable second adaptors for use in the invention. The second adaptors are typically polynucleotide adaptors and may be any of the polynucleotides discussed above with reference to the first adaptors. The second adaptors are typically synthetic or semisynthetic.
The second adaptors are preferably double stranded polynucleotide adaptors. The second adaptors may be any length. The second adaptors are preferably at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 32, at least about 35, at least about 40, at least about 45, at least 47, or at least about 50 nucleotides or nucleotide pairs in length.
The second adaptors are preferably ligated to the strands in the double stranded constructs comprising the different first barcodes, different second barcodes and complementary sequences. This allows the third different barcodes to be added to the double barcoded constructs produced in the second round.
Step (f) preferably comprises hybridising, preferably specifically hybridising, the second adaptors to the double stranded constructs in the plurality of third samples and ligating the second adaptors to the double stranded constructs to form triple barcoded constructs. The triple barcoded constructs comprise DNA sequences transcribed from the RIMA molecules. Hybridisation, specific hybridisation, and ligation are discussed above and any of those approaches may be used with the second adaptors. The second adaptors preferably comprise overhangs capable of hybridising, preferably specifically hybridising, to the RT primers. The second adaptors preferably comprise overhangs capable of hybridising, preferably specifically hybridising, to sequences at or near the 5' ends of the RT primers. These sequences are discussed in more detail above with reference to the RT primers. The overhangs may be any length. The overhangs are preferably at least about 5 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 12 nucleotides, at least about 15 nucleotides, or at least about 20 nucleotides in length. The overhangs are preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length. The overhangs are preferably the same length as the sequences at or near the 5' ends of the RT primers. The second adaptors may use the same overhangs. This allows the method to use a standard set of second adaptors in which only the third barcodes differ between third samples.
If the second adaptors are double stranded, the different third barcodes are preferably present in the non-overhanging strands of the second adaptors which are ligated to the double stranded constructs. An example of this is shown in Figure 3. This approach produces constructs which are triple barcoded.
Step (f) is preferably carried out in the presence of 5' phosphorylated blockers which are complementary to the overhangs in the second adaptors. An example of this is shown in Figure 4. This mitigates against false re-barcoding or barcode crosstalk between surplus RT primers after cells are pooled. These blockers may comprise the sequences at or near the 5' ends of the RT primers. An example of this is shown in Figure 3.
The second adaptors preferably comprise additional sequences at their 5' ends which facilitate isolation, amplification and/or characterisation of the triple barcoded constructs. The additional sequences may comprise click chemistry. The additional sequences may comprise or facilitate the addition of sequencing adaptors. Such adaptors are discussed in
more detail below. The second adaptors preferably comprise the same additional sequences. This allows the method to use a standard set of RT primers in which only the second barcodes differ between second samples.
The second adaptors used in each third sample comprise a different third barcode. The second adaptors used in each third sample comprise a unique third barcode. The double stranded constructs in each third sample are labelled with a different or unique third barcode. The second adaptors used in each third sample comprise a third barcode that is different from the third barcode used in all the other third samples. The double stranded constructs in each third sample are labelled with a third barcode which is different from the third barcode used to label the double stranded constructs in all the other third samples. The second adaptors used in each third sample comprise a third barcode that is different from any other barcode or all the other barcodes used in the method. The double stranded constructs in each third sample are labelled with a third barcode which is different from any other barcode or all the other barcodes used in the method. Barcodes are discussed in more detail below.
Exemplary second adaptors are shown in the Examples.
Barcodes
The method involves producing triple barcoded constructs. Step (b) comprises the use of first different barcodes. Step (d) involves the use of different second barcodes. Step (f) comprises the use of different third barcodes. The meaning of "different" or "unique" barcodes is discussed above. The barcodes are typically polynucleotide barcodes. The barcodes may be any of those discussed above. The first, second and third different polynucleotides all typically comprise the same type of polynucleotide. Polynucleotide barcodes are well-known in the art (Kozarewa, I. et al., (2011), Methods Mol. Biol. 733, p279-298). A barcode is a specific sequence of polynucleotide that can be characterised and identified. A barcode is preferably a specific sequence of polynucleotide that affects the current flowing through the pore in a specific and known manner.
The barcode may comprise one or more different nucleotide species. For instance, T k-mers (i.e. k-mers in which the central nucleotide is thymine-based, such as TTA, GTC, GTG and CTA) typically have the lowest current states. Modified versions of T nucleotides may be introduced into the modified polynucleotide to reduce the current states further and thereby increase the total current range seen when the barcode moves through the pore.
G k-mers (i.e. k-mers in which the central nucleotide is guanine-based, such as TGA, GGC, TGT and CGA) tend to be strongly influenced by other nucleotides in the k-mer and so modifying the G nucleotides in the modified polynucleotide may help them to have more independent current positions.
Including three copies of the same nucleotide species instead of three different species may facilitate characterization because it is then only necessary to map, for example, 3- nucleotide k-mers in the modified polynucleotide. However, such modifications do reduce the information provided by the barcode.
One or more abasic nucleotides may be included in the barcode. Using one or more abasic nucleotides results in characteristic current spikes. This allows the clear highlighting of the positions of the one or more nucleotide species in the barcode.
The nucleotide species in the barcode may comprise a chemical atom or group such as a propynyl group, a thio group, an oxo group, a methyl group, a hydroxymethyl group, a formyl group, a carboxy group, a carbonyl group, a benzyl group, a propargyl group or a propargylamine group. The chemical group or atom may be or may comprise a fluorescent molecule, biotin, digoxigenin, DNP (dinitrophenol), a photo-labile group, an alkyne, DBCO, azide, free amino group, a redox dye, a mercury atom, or a selenium atom.
The barcode may comprise a nucleotide species comprising a halogen atom. The halogen atom may be attached to any position on the different nucleotide species, such as the nucleobase and/or the sugar. The halogen atom is preferably fluorine (F), chlorine (Cl), bromine (Br) or iodine (I). The halogen atom is most preferably F or I.
The barcodes may be any length. The barcodes are preferably at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 11 nucleotides, at least about 12 nucleotides, at least about 13 nucleotides, at least about 14 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, or at least about 50 nucleotides in length. The first, second and third different barcodes may the same length or different lengths.
The invention is based on the use of unique or different barcodes for each sample. As explained in more detail below, the method often requires large numbers of unique or different barcodes. The longer the barcodes, the more different possible barcodes that can be generated and used in the method of the invention. The skilled person is capable of designing barcodes and sufficient barcode numbers for use in the method of the invention.
Barcode numbers
The barcode used in each sample is preferably different from any other barcode or all the other barcodes used in the method. The barcode used in each first sample, each second sample and each third sample is preferably different from any other barcode or all the other
barcodes used in the method. This means each sample is contacted with a unique or different barcode.
The number of cells in the population is preferably less than the number of first samples multiplied by the number of second samples multiplied the number of third samples (/.e., first samples x second samples x third samples). If the number of possible barcode combinations used is greater than the number of cells in the population, there is a high probability the RIMA molecules from each cell or the constructs produced from each cell will be labelled with a unique combination of three barcodes. If the number of possible barcode combinations used is greater than the number of cells in the population, there is a high probability the RNA molecules from each cell or the constructs produced from each cell will be labelled with a unique combination of three barcodes. This allows RNA molecules to be associated with individual cells in the population and the transcriptome of each cell or two or more cells to be measured.
Numbers of individual samples are discussed above. The (i) plurality of first samples, (ii) plurality of second samples, and/or (ii) plurality of third samples preferably comprise(s) at least 24 samples. Preferably, (i), (ii), (iii), (i) and (ii), (i) and (iii), (ii) and (iii), or (i), (ii) and (iii) preferably comprises at least about 24 samples. Preferably, (i), (ii), (iii), (i) and (ii), (i) and (iii), (ii) and (iii), or (i), (ii) and (iii) preferably comprises at least 24 samples, at least 48 samples, at least 96 samples, at least 384 samples, at least 1536 samples or at least 3456 samples. The minimum number of unique or different barcodes is preferably at least about 13,824, at least about 110,592, at least about 884,736, at least about 56,623,104, at least about 3,623,878,656 or at least about 41,278,242,816 barcodes. The skilled person is capable of designing suitable barcode numbers and numbers of samples based on the number of cells in the starting population.
Previous steps
The population of cells is preferably fixed and/or permeabilised before step (a). Suitable methods for doing this are known in the art (e.g., Chen et al. (2018). PBMC fixation and processing for Chromium single-cell RNA sequencing. Journal of Translational Medicine, 16(1).) Specific methods are also discussed in the Examples.
Additional steps
The method preferably further comprises repeating steps (e) to (f) at least once comprising the formation of a plurality of fourth or more samples and using third or more adaptors which comprise a different fourth or more barcode for each fourth or more sample to produce quadruple or more barcoded constructs. This repetition may be conducted more than one to produce quintuple, sextuple, septuple or octuple or more barcoded constructs. The skilled person is capable of designing a method uniquely labelling RNA molecules or
producing uniquely barcoded constructs comprising sequences transcribed from RNA molecules with any number of barcodes. Any of the embodiments discussed above with reference to the first, second and third rounds equally apply to any of the repeated steps.
The method produces constructs having three or more barcodes. The method may produce constructs having any number of three or more barcodes, such as 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more or 20 or more barcodes. In the following discussion, all of these constructs will be collectively called "uniquely labelled constructs" or "barcoded constructs".
Characterisation
The method preferably further comprises isolating the barcoded constructs from the cells. Any method of isolation may be used including any of the ones discussed below. Any of the adaptors or RT primers can comprise sequences and/or molecules which facilitate isolation.
The method preferably further comprises amplifying the barcoded constructs. Any amplification method may be used, including polymerase chain reaction (PCR). Any of the adaptors or RT primers can comprise primer sites that allow amplification. This facilitates preparation of the barcoded constructs for amplification.
The method preferably further comprises isolating the barcoded constructs from the cells and amplifying the barcoded constructs.
The method preferably further comprises characterising or sequencing the barcoded constructs. Any method may be used for sequencing or characterising the barcoded constructs, including next generation sequencing. The barcoded constructs are preferably characterised or sequenced using a nanopore. This is discussed in more detail below.
In a preferred embodiment, the invention provides a method of characterising the transcriptome of an individual cell or two or more cells in a population of cells, the method comprising a conducting a method of the invention on the population of cells and characterising the RNA molecules from the individual cell or two or more cells using their unique labelling. The RNA molecules may be characterised using their barcoding. The characterisation preferably comprises sequencing the RNA molecules. This can be achieved by sequencing the barcoded constructs produced from the RNA molecules. The invention also provides a method of characterising the transcriptome of an individual cell or two or more cells in a population of cells, the method comprising producing uniquely labelled constructs transcribed from the RNA molecules in the population of cells using a method of the invention and characterising the uniquely labelled constructs from the individual cell or two or more cells using their unique labelling. The constructs may be characterised using their barcode. As explained above, the method of the invention can be conducted such there
is a high probability the RIMA molecules from each cell or barcoded constructs produced from each cell will be labelled with a unique combination of three barcodes. The characterisation is preferably sequencing of the RNA molecules or barcoded constructs, preferably using a nanopore. Any embodiments discussed above, especially in relation to the first to third rounds, adaptors, or RT primers, equally apply to these embodiments.
These methods may involve characterising the transcriptome of any number of two or more cells, such as 3 or more, 4 or more, 5 or more, 10 or more, 20 or more, 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more or 10,000 or more cells. The number of two or more cells may be any of the numbers discussed above with reference to the population of cells used in the method of the invention.
Sequencing adaptors
The barcoded constructs produced using the method of the invention are preferably characterised or sequenced. The barcoded constructs are preferably modified using sequencing adaptors which facilitate characterisation or sequencing, especially using nanopores. As explained above, the second adaptors may comprise additional sequences which comprise or facilitate the addition of sequencing adaptors.
A sequencing adaptor typically comprises a polynucleotide strand capable of being attached to the end of a target polynucleotide. The target polynucleotide is typically intended for characterisation in accordance with methods disclosed herein, and includes the first adaptors, RT primers, second adaptors and barcoded constructs.
A sequencing adaptor may be added to both ends of the target polynucleotide. Alternatively, different adaptors may be added to the two ends of the target polynucleotide. An adaptor may be added to just one end of the target polynucleotide. Methods of adding adaptors to polynucleotides are known in the art. Adaptors may be attached to polynucleotides, for example, by ligation, by click chemistry, by tagmentation, by topoisomerisation or by any other suitable method.
An adaptor may be synthetic or artificial. Typically, an adaptor comprises a polymer as described herein. The adaptor preferably comprises a polynucleotide. An adaptor may comprise a single-stranded polynucleotide strand. An adaptor may comprise a doublestranded polynucleotide. A sequencing adaptor may comprise any of the polynucleotide discussed above with reference to the first adaptors and includes DNA, RNA, modified DNA (such as a basic DNA), RNA, PNA, LNA, BNA and/or PEG. Usually, the adaptor comprises single stranded and/or double stranded DNA or RNA.
The sequencing adaptors may be Y adaptors. Y adaptors are typically double stranded and comprise (a) at one end, a region where the two strands are hybridised together and (b), at
the other end, a region where the two strands are not complementary. The non- complementary parts of the strands form overhangs. The hybridised stem of the adaptors typically attaches to the 5' end of a first strand of a double-stranded polynucleotide and the 3' end of a second strand of a double-stranded polynucleotide; or to the 3' end of a first strand of a double-stranded polynucleotide and the 5' end of a second strand of a doublestranded polynucleotide. The presence of a non-complementary region in the Y adaptors gives them their Y shape since the two strands typically do not hybridise to each other unlike the double stranded portion. The hybridised stem end of the Y adaptors may also comprise a short overhang that allows them to specifically hybridise to and be attached to the first adaptors, RT primers, second adaptors and barcoded constructs.
Some of the methods of the invention use a polynucleotide binding protein to control the movement of barcoded constructs with respect to a nanopore. A polynucleotide binding protein may bind to an overhang of an adaptor such as a Y adaptor. A polynucleotide binding protein may bind to the double stranded region. A polynucleotide binding protein may bind to a single-stranded and/or a double-stranded region of the adaptor. A first polynucleotide binding protein may bind to the single-stranded region of such an adaptor and a second polynucleotide binding protein may bind to the double-stranded region of the adaptor.
The sequencing adaptors preferably comprise a membrane anchor or a pore anchor. The anchors may be attached to polynucleotides that are complementary to and hence hybridised to the overhangs to which a polynucleotide binding protein is bound.
One of the non-complementary strands of sequencing adaptors, such as Y adaptors, may comprise leader sequences, which when contacted with a nanopore are capable of threading into the nanopore.
The leader sequences typically comprise a polymer such as a polynucleotide, for instance DNA or RNA, a modified polynucleotide (such as abasic DNA), PNA, LNA, polyethylene glycol (PEG) or a polypeptide. The leader sequences preferably comprise a single strand of DNA, such as a poly dT section. The leader sequences can be any length, but are typically 10 to 150 nucleotides in length, such as from 20 to 120, 30 to 100, 40 to 80 or 50 to 70 nucleotides in length.
The sequencing adaptors may be hairpin loop adaptors. Hairpin loop adaptors are adaptors comprising a single polynucleotide strand, wherein the ends of the polynucleotide strand are capable of hybridising to each other, or are hybridized to each other, and wherein the middle section of the polynucleotide forms a loop. Suitable hairpin loop adaptors can be designed using methods known in the art. Typically, the 3' end of a hairpin loop adaptor attaches to the 5' end of a first strand of a double-stranded polynucleotide and the 5' end of
the hairpin loop adaptor attaches to the 3' end of a second strand of a double-stranded polynucleotide; or the 5' end of a hairpin loop adaptor attaches to the 3' end of a first strand of a double-stranded polynucleotide and the 3' end of the hairpin loop adaptor attaches to the 5' end of a second strand of a double-stranded polynucleotide. As explained in more detail below, sequencing adaptors can be attached to a target polynucleotide in order to characterise the target polynucleotide.
Those skilled in the art will also appreciate that when the adaptors comprise a polynucleotide strand, the sequences of the adaptors are typically not determinative and can be controlled or chosen according to the polynucleotide binding protein and other experimental conditions such as any polynucleotides to be characterised. Exemplary sequences are provided solely by way of illustration in the examples. For example, the adaptors may comprise a sequence such as one or more of SEQ ID NOs: 21-26 or 28-33 in WO 2021/255476 (incorporated herein by reference in its entirety) or polynucleotide sequences having at least 20%, such as at least 30%, e.g., at least 40% such as at least 50%, e.g., at least 60% such as at least 70%, e.g., at least 80%, for example at least 90% e.g., at least 95% sequence similarity or identity to one or more of SEQ ID NOs: 21-26 or 28-33 in WO 2021/255476 (incorporated herein by reference in its entirety). The sequences of the adaptors can typically be altered without negatively affecting the efficacy of the method of the invention.
Sequencing adaptors may comprise a loading site for loading the polynucleotide binding protein. The loading site may be for instance a single-stranded region which can targeted by the polynucleotide binding protein. The loading site may be a region of the sequencing adaptor to which an exogenous polynucleotide strand comprising the polynucleotide binding protein can bind in order to transfer the polynucleotide binding protein to the polynucleotide to be assessed in the method of the invention.
The polynucleotide binding protein if present may be provided on sequencing adaptors. WO 2015/110813 and WO 2020/234612 describe the loading of polynucleotide binding proteins onto a target polynucleotide such as an adaptor and are hereby incorporated by reference in their entireties.
Spacers
Any of the polynucleotides described herein, including the first adaptors, RT primers, second adaptors and barcoded constructs, may comprise one or more spacers, e.g., from about one to about 10 spacers, e.g., from about 1 to about 5 spacers, e.g., about 1, 2, 3, 4 or 5 spacers. The spacer may comprise any suitable number of spacer units. A spacer typically provides an energy barrier which impedes movement of a polynucleotide binding protein. For example, a spacer may impede movement of a polynucleotide binding protein by
reducing the traction of the protein, e.g., using an abasic spacer. A spacer may physically block movement of the protein, for instance by introducing a bulky chemical group to physically impede the movement of the polynucleotide binding protein.
One or more spacers are typically included in the polynucleotide or in a sequencing adaptor to provide a distinctive signal when they pass through or across a nanopore. One or more spacers may be used to define or separate one or more regions of a polynucleotide, e.g., to separate an adaptor from the target polynucleotide.
A spacer may comprise a linear molecule, such as a polymer, e.g., a polypeptide or a polyethylene glycol (PEG). Typically, such a spacer has a different structure from the target polynucleotide. For instance, if the target polynucleotide is DNA, the or each spacer typically does not comprise DNA. In particular, if the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the or each spacer preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) or a synthetic polymer with nucleotide side chains. A spacer may comprise one or more nitroindoles, one or more inosines, one or more acridines, one or more 2- aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dTs), one or more inverted dideoxy-thymidines (ddTs), one or more dideoxy-cytidines (ddCs), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-Methyl RNA bases, one or more Isodeoxycytidines (Iso-dCs), one or more Iso-deoxyguanosines (Iso-dGs), one or more C3 (OC3H6OPO3) groups, one or more photo-cleavable (PC) [OC3H6-C(O)NHCH2-C6H3NO2- CH(CH3)OPO3] groups, one or more hexandiol groups, one or more spacer 9 (iSp9) [(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSplS) [(OCH2CH2)6OPO3] groups; or one or more thiol connections. A spacer may comprise any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9 and iSpl8 spacers are all available from IDT®. A spacer may comprise any number of the above groups as spacer units.
A spacer may comprise one or more chemical groups, e.g., one or more pendant chemical groups. The one or more chemical groups may be attached to one or more nucleobases in a sequencing adaptor. The one or more chemical groups may be attached to the backbone of a sequencing adaptor. Any number of appropriate chemical groups may be present, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and/or biotin, cholesterol, methylene blue, dinitrophenols (DNPs), digoxigenin and/or anti-digoxigenin and di benzylcyclooctyne groups.
A spacer may comprise one or more abasic nucleotides (/.e., nucleotides lacking a nucleobase), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more abasic nucleotides. The nucleobase can be replaced by -H (idSp) or -OH in the abasic nucleotide. Abasic spacers
can be inserted into target polynucleotides by removing the nucleobases from one or more adjacent nucleotides. For instance, polynucleotides may be modified to include 3- methyladenine, 7-methylguanine, l,N6-ethenoadenine inosine or hypoxanthine and the nucleobases may be removed from these nucleotides using Human Alkyladenine DNA Glycosylase (hAAG). Alternatively, polynucleotides may be modified to include uracil and the nucleobases removed with Uracil-DNA Glycosylase (UDG). The one or more spacers preferably do not comprise any abasic nucleotides.
Suitable spacers can be designed or selected depending on the nature of the polynucleotide or sequencing adaptor, the polynucleotide binding protein, and the conditions under which the method is to be carried out.
Tags
Any of the polynucleotides, including the first adaptors, RT primers, second adaptors, barcoded constructs and/or sequencing adaptors, used in the invention may comprise a tag or tether. For example, a polynucleotide can bind to a tag on a nanopore, e.g., via its adaptor, and release at some point, e.g., during characterization of the polynucleotide by the nanopore. A strong non-covalent bond (e.g., biotin/avidin) is still reversible and can be useful in some embodiments of the methods described herein.
The pair of pore tag and sequencing adaptor can be configured such that the binding strength or affinity of a binding site on the polynucleotide (e.g., a binding site provided by an anchor or a leader sequence of an adaptor or by a capture sequence within the duplex stem of an adaptor) to a tag on a nanopore is sufficient to maintain the coupling between the nanopore and polynucleotide until an applied force is placed on it to release the bound polynucleotide from the nanopore.
The tags or tethers are preferably uncharged. This can ensure that the tags or tethers are not drawn into the nanopore under the influence of a potential difference.
One or more molecules that attract or bind the polynucleotide or adaptor may be linked to the detector (e.g., the pore). Any molecule that hybridizes to the adaptor and/or target polynucleotide may be used. The molecule attached to the pore may be selected from a PNA tag, a PEG linker, a short oligonucleotide, a positively charged amino acid and an aptamer. Pores having such molecules linked to them are known in the art. For example, pores having short oligonucleotides attached thereto are disclosed in Howarka et al (2001) Nature Biotech. 19: 636-639 and WO 2010/086620, and pores comprising PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J. Am. Chem. Soc. 122(11): 2411-
A short oligonucleotide attached to the detector (e.g., a nanopore), which oligonucleotide comprises a sequence complementary to a sequence in the leader sequence or another single stranded sequence in the adaptor may be used to enhance capture of the target polynucleotide in the methods described herein.
The tag or tether may comprise or be an oligonucleotide (e.g., DNA, RIMA, LNA, BNA, PNA, or morpholino). The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) can have about 10-30 nucleotides in length or about 10-20 nucleotides in length. The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) for use in the tag or tether can have at least one end (e.g., 3'- or 5'-end) modified for conjugation to other modifications or to a solid substrate surface including, e.g., a bead. The end modifiers may add a reactive functional group which can be used for conjugation. Examples of functional groups that can be added include, but are not limited to amino, carboxyl, thiol, maleimide, aminooxy, and any combinations thereof. The functional groups can be combined with different length of spacers (e.g., C3, C9, C12, Spacer 9 and 18) to add physical distance of the functional group from the end of the oligonucleotide sequence.
The tag or tether may comprise or be a morpholino oligonucleotide. The morpholino oligonucleotide can have about 10-30 nucleotides in length or about 10-20 nucleotides in length. The morpholino oligonucleotides can be modified or unmodified. For example, the morpholino oligonucleotide can be modified on the 3' and/or 5' ends of the oligonucleotides. Examples of modifications on the 3' and/or 5' end of the morpholino oligonucleotides include, but are not limited to 3' affinity tag and functional groups for chemical linkage (including, e.g., 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyl dithio, and any combinations thereof); 5' end modifications (including, e.g., 5'-primary ammine, and/or 5'- dabcyl), modifications for click chemistry (including, e.g., 3'-azide, 3'-alkyne, 5'-azide, 5'- alkyne), and any combinations thereof.
The tag or tether may further comprise a polymeric linker, e.g., to facilitate coupling to a detector e.g., a nanopore. An exemplary polymeric linker includes, but is not limited to, polyethylene glycol (PEG). The polymeric linker may have a molecular weight of about 500 Da to about 10 kDa (inclusive), or about 1 kDa to about 5 kDa (inclusive). The polymeric linker (e.g., PEG) can be functionalized with different functional groups including, e.g., but not limited to maleimide, NHS ester, dibenzocyclooctyne (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combinations thereof. The tag or tether may further comprise a 1 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further comprise a 2 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further comprise a 3 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further comprise a 5 kDa PEG with a 5'-maleimide group and a 3'-DBCO group.
Other examples of a tag or tether include, but are not limited to His tags, biotin or streptavidin, antibodies that bind to analytes, aptamers that bind to analytes, analyte binding domains such as DNA binding domains (including, e.g., peptide zippers such as leucine zippers, single-stranded DNA binding proteins (SSB)), and any combinations thereof.
The tag or tether may be attached to the external surface of a nanopore, e.g., on the cis side of a membrane, using any methods known in the art. For example, one or more tags or tethers can be attached to the nanopore via one or more cysteines (cysteine linkage), one or more primary amines such as lysines, one or more non-natural amino acids, one or more histidines (His tags), one or more biotin or streptavidin, one or more antibody-based tags, one or more enzyme modification of an epitope (including, e.g., acetyl transferase), and any combinations thereof. Suitable methods for carrying out such modifications are well-known in the art. Suitable non-natural amino acids include, but are not limited to, 4-azido-L- phenylalanine (Faz) and any one of the amino acids numbered 1-71 in Figure 1 of Liu C. C. and Schultz P. G., Annu. Rev. Biochem., 2010, 79, 413-444.
Where one or more tags or tethers are attached to a nanopore via cysteine linkage(s), the one or more cysteines can be introduced to one or more monomers that form the nanopore by substitution. The nanopore may be chemically modified by attachment of (i) Maleimides including diabromomaleimides such as: 4-phenylazomaleinanil, l.N-(2- Hydroxyethyl)maleimide, N-Cyclohexylmaleimide, 1.3-Maleimidopropionic Acid, 1.1-4- Aminophenyl-lH-pyrrole,2,5,dione, l.l-4-Hydroxyphenyl-lH-pyrrole,2,5,dione, N- Ethylmaleimide, N-Methoxycarbonylmaleimide, N-tert-Butylmaleimide, N-(2- Aminoethyl)maleimide , 3-Maleimido-PROXYL , N-(4-Chlorophenyl)maleimide, l-[4- (dimethylamino)-3,5-dinitrophenyl]-lH-pyrrole-2, 5-dione, N-[4-(2- Benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(l-naphthyl)- maleimide, N-(2,4-xylyl)maleimide, N-(2,4-difluorophenyl)maleimide , N-(3-chloro-para- tolyl)-maleimide, l-(2-amino-ethyl)-pyrrole-2, 5-dione hydrochloride, l-cyclopentyl-3- methyl-2,5-dihydro-lH-pyrrole-2, 5-dione, l-(3-aminopropyl)-2,5-dihydro-lH-pyrrole-2,5- dione hydrochloride, 3-methyl-l-[2-oxo-2-(piperazin-l-yl)ethyl]-2,5-dihydro-lH-pyrrole- 2, 5-dione hydrochloride, l-benzyl-2,5-dihydro-lH-pyrrole-2, 5-dione, 3-methyl-l-(3,3,3- trifluropropyl)-2,5-dihydro-lH-pyrrole-2, 5-dione, l-[4-(methylamino)cyclohexyl]-2,5- dihydro-lH-pyrrole-2, 5-dione trifluroacetic acid, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2, l-benzyl-3- methyl-2,5-dihydro-lH-pyrrole-2, 5-dione, l-(2-fluorophenyl)-3-methyl-2,5-dihydro 1H- pyrrole-2, 5-dione, N-(4-phenoxyphenyl)maleimide , N-(4-nitrophenyl)maleimide (ii) lodocetamides such as :3-(2-Iodoacetamido)-proxyl, N-(cyclopropylmethyl)-2- iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2- trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, N-(4-
(aminosulfonyl)phenyl)-2-iodoacetamide, N-(l,3-benzothiazol-2-yl)-2-iodoacetamide, N- (2,6-diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide, (iii) Bromoacetamides: such as N-(4-(acetylamino)phenyl)-2-bromoacetamide , N-(2- acetylphenyl)-2-bromoacetamide , 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3- (trifluoromethyl)phenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide , 2-bromo-N- (4-fluorophenyl)-3-methylbutanamide, N-Benzyl-2-bromo-N-phenylpropionamide, N-(2- bromo-butyryl)-4-chloro-benzenesulfonamide, 2-Bromo-N-methyl-N-phenylacetamide, 2- bromo-N-phenethyl-acetamide,2-adamantan-l-yl-2-bromo-N-cyclohexyl-acetamide, 2- bromo-N-(2-methylphenyl)butanamide, Monobromoacetanilide, (iv) Disulphides such as: aldrithiol-2 , aldrithiol-4 , isopropyl disulfide, l-(Isobutyldisulfanyl)-2-methylpropane, Dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-Pyridyldithio)propionic acid, 3-(2- Pyridyldithio)propionic acid hydrazide, 3-(2-Pyridyldithio)propionic acid N-succinimidyl ester, am6amPDPl-[3CD and (v) Thiols such as: 4-Phenylthiazole-2-thiol, Purpald, 5, 6, 7, 8- tetrahydro-quinazoline-2-thiol.
The tag or tether may be attached directly to a nanopore or via one or more linkers. The tag or tether may be attached to the nanopore using the hybridization linkers described in WO 2010/086602 (incorporated herein by reference in its entirety). Alternatively, peptide linkers may be used. Peptide linkers are amino acid sequences. The length, flexibility and hydrophilicity of the peptide linker are typically designed such that it does not to disturb the functions of the monomer and pore. Preferred flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10 or 16, serine and/or glycine amino acids. More preferred flexible linkers include (SG)i, (SG)2, (SG)3, (SG)4, (SG)5 and (SG)8 wherein S is serine and G is glycine. Preferred rigid linkers are stretches of 2 to 30, such as 4, 6, 8, 16 or 24, proline amino acids. More preferred rigid linkers include (P)i2 wherein P is proline.
Suitable pore tags are also described in WO 2018/100370, which describes non-hairpin methods for characterising double-stranded polynucleotides and is herein incorporated by reference in its entirety.
Anchor
Any of the polynucleotides, including the first adaptors, RT primers, second adaptors, barcoded constructs and/or sequencing adaptors, used in the invention may comprise a membrane anchor. The anchor typically assists in the characterisation of a target polynucleotide in accordance with the methods disclosed herein. For example, a membrane anchor may promote localisation of the selected polynucleotides around a nanopore.
The anchor may be a polypeptide anchor and/or a hydrophobic anchor that can be inserted into the membrane. The hydrophobic anchor is preferably a lipid, fatty acid, sterol, carbon
nanotube, polypeptide, protein, or amino acid, for example cholesterol, palmitate, or tocopherol. The anchor may comprise thiol, biotin, or a surfactant.
The anchor may be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or a fusion protein), Ni-NTA (for binding to poly-histidine or poly-histidine tagged proteins) or peptides (such as an antigen).
The anchor preferably comprises a linker, or 2, 3, 4 or more linkers. Preferred linkers include, but are not limited to, polymers, such as polynucleotides, polyethylene glycols (PEGs), polysaccharides and polypeptides. These linkers may be linear, branched, or circular. For instance, the linker may be a circular polynucleotide. The adaptor may hybridise to a complementary sequence on a circular polynucleotide linker. The one or more anchors or one or more linkers may comprise a component that can be cut or broken down, such as a restriction site or a photolabile group. The linker may be functionalised with maleimide groups to attach to cysteine residues in proteins. Suitable linkers are described in WO 2010/086602 (incorporated herein by reference in its entirety).
The anchor is preferably cholesterol or a fatty acyl chain. For example, any fatty acyl chain having a length of from 6 to 30 carbon atom, such as hexadecanoic acid, may be used.
Examples of suitable anchors and methods of attaching anchors to adaptors are disclosed in WO 2012/164270 and WO 2015/150786 (incorporated herein by reference in their entireties).
The anchor may consist or comprise a hydrophobic modification to the polynucleotide or sequencing adaptor. The hydrophobic modification may comprise a modified phosphate group comprised within the polynucleotide or polynucleotide anchor. The hydrophobic modification may for example comprise a phosphorothioate such as a charge-neutralized alkyl-phosphorothioate (PPT) as described in Jones et al, J. Am. Chem. Soc. 2021, 143, 22, 8305, the entire contents of which are hereby incorporated by reference. Suitable alkyl groups include for example Ci-Cio alkyl groups such as C2-C6 alkyl groups, e.g., methyl, ethyl, propyl, butyl, pentyl and hexyl groups. Incorporation of the charge-neutralized alkyl- phosphorothioate into a polynucleotide allows for the polynucleotide to anchor to a hydrophobic region such as a lipid bilayer.
Biotin enrichment
Any of the polynucleotides, including the first adaptors, RT primers, second adaptors, barcoded constructs and/or sequencing adaptors, preferably comprise biotin. The biotin may be used to isolate barcoded constructs. Suitable methods for biotin-based enrichment are known in the art. For instance, a surface, such as a bead, comprising avidin/or streptavidin may be used to barcoded constructs comprising biotin.
Characterising methods
The method of the invention preferably further comprises characterising or sequencing the barcoded constructs. This allows the RIMA molecules to be characterised or sequenced. As explained in more detail above, the method of the invention preferably further comprises using sequencing adaptors to characterise or sequence the barcoded constructs. The invention also provides methods of characterising the transcriptome of an individual cell or two or more cells in a population of cells which comprises characterising the RNA molecules from the individual cell or characterising the uniquely labelled constructs from the individual cell or two or more cells using their unique labelling. The RNA molecules or the constructs can be characterised using their barcoding. These methods preferably comprising using sequencing adaptors.
Any method of characterisation may be used. The method preferably uses next generation sequencing (NGS).
The barcoded constructs are preferably moved with respect to a detector such as a nanopore. The detector may be selected from (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube; and (v) a nanopore. Preferably, the detector is a nanopore.
The barcoded constructs may be characterised in the method of the invention in any suitable manner. The barcoded constructs are preferably characterised by detecting an ionic current or optical signal as they move with respect to a nanopore. This is described in more detail herein. The method is amenable to these and other methods of characterising polynucleotides.
In another non-limiting example, the barcoded constructs are characterised by detecting the by-products of a polynucleotide-processing reaction, such as a sequencing by synthesis reaction. The method may thus involve detecting the product of the sequential addition of (poly)nucleotides by an enzyme such as a polymerase to barcoded constructs. The product may be a change in one or more properties of the enzyme such as in the conformation of the enzyme. Such methods may thus comprise subjecting an enzyme such as polymerase or a reverse transcriptase to the barcoded constructs as templates under conditions such that the template-dependent incorporation of nucleotide bases into a growing oligonucleotide strand causes conformational changes in the enzyme in response to sequentially encountering template nucleic acid bases and/or incorporating template-specified natural or analog bases (/.e., an incorporation event), detecting the conformational changes in the enzyme in response to such incorporation events, and thereby detecting the sequence of the templates. In such methods the barcoded constructs may be moved in accordance with the method of the invention. Such methods may involve detecting and/or measuring
incorporation events using methods known to those skilled in the art, such as those described in US 2017/0044605.
In another embodiment, by-products may be labelled so that a phosphate labelled species is released upon the addition of a nucleotide to a synthesised nucleic acid strand that is complementary to the template barcoded constructs, and the phosphate labelled species is detected e.g., using a detector as described herein. The barcoded constructs being characterised in this way may be moved in accordance with the methods herein. Suitable labels may be optical labels that are detected using a nanopore, or a zero-mode wave guide, or by Raman spectroscopy, or other detectors. Suitable labels may be non-optical labels that are detected using a nanopore, or other detectors.
In another approach, nucleoside phosphates (nucleotides) are not labelled and upon the addition of nucleotides to synthesised nucleic acid strands that are complementary to the barcoded constructs, natural by-product species are detected. Suitable detectors may be ion-sensitive field-effect transistors, or other detectors.
These and other detection methods are suitable for use in the methods described herein. Any suitable measurements can be taken using a detector as the barcoded constructs move with respect to the detector.
Nanopore characterisation
The barcoded constructs are preferably characterised using a nanopore.
The method preferably comprises (i) contacting the barcoded constructs with a nanopore such that the barcoded constructs move with respect to the nanopore and (ii) taking one or more measurements as the barcoded constructs move with respect to the nanopore wherein the measurements are indicative of one or more characteristics of the barcoded constructs and thereby characterising the barcoded constructs. The one or more characteristics are preferably selected from (i) the length of the barcoded constructs, (ii) the identity of the barcoded constructs, (iii) the sequence of the barcoded constructs, (iv) the secondary structure of the barcoded constructs and (v) whether or not the barcoded constructs is modified. The barcoded constructs may be modified by methylation, by oxidation, by damage, with one or more proteins or with one or more labels, tags, or spacers. The one or more characteristics of the barcoded constructs are preferably measured by electrical measurement and/or optical measurement. The electrical measurement is preferably a current measurement, an impedance measurement, a tunnelling measurement, or a field effect transistor (FET) measurement.
The method more preferably comprises (i) contacting the barcoded constructs with a nanopore such that the barcoded constructs move through the nanopore and (ii) measuring
the current moving through the nanopore as the barcoded constructs move through the nanopore wherein the current is indicative of one or more characteristics of the barcoded constructs and thereby characterising the barcoded constructs. The one or more characteristics may be any of those described above.
The movement of the barcoded constructs with respect to the nanopore or through the nanopore is preferably controlled using a polynucleotide binding protein. The use of such proteins in nanopore sequencing is known. Examples of suitable proteins are discussed in more detail below.
Any suitable nanopore can be used. The nanopore is preferably a transmembrane pore. A transmembrane pore is a structure that crosses the membrane to some degree. It permits hydrated ions driven by an applied potential to flow across or within the membrane. The transmembrane pore typically crosses the entire membrane so that hydrated ions may flow from one side of the membrane to the other side of the membrane. However, the transmembrane pore does not have to cross the membrane. It may be closed at one end. For instance, the pore may be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions may flow.
The nanopore typically has a first opening and a second opening. The first opening is typically the cis opening and the second opening is typically the trans opening. However, the first opening may be the trans opening and the second opening may be the cis opening. Any polynucleotide binding protein used in the method of the invention is typically provided at the first opening of the nanopore and thus controls the movement of the target polynucleotide in the direction from the second opening of the nanopore towards the first opening of the nanopore.
Any transmembrane pore may be used in the method of the invention. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores and solid-state pores. The pore may be a DNA origami pore (Langecker et al., Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013/083983.
The nanopore is preferably a transmembrane protein pore. A transmembrane protein pore is a polypeptide or a collection of polypeptides that permits hydrated ions, such as polynucleotide, to flow from one side of a membrane to the other side of the membrane. In the method of the invention, the transmembrane protein pore is capable of forming a pore that permits hydrated ions driven by an applied potential to flow from one side of the membrane to the other. The transmembrane protein pore preferably permits polynucleotides to flow from one side of the membrane, such as a triblock copolymer
membrane, to the other. The transmembrane protein pore allows a polynucleotide to be moved through the pore.
The nanopore may be a transmembrane protein pore which is a monomer or an oligomer. The pore is preferably made up of several repeating subunits, such as at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, or at least about 16 subunits. The pore is preferably a hexameric, heptameric, octameric or nonameric pore. The pore may be a homo-oligomer or a hetero-oligomer.
The transmembrane protein pore may comprise a barrel or channel through which the ions may flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane [3-barrel or channel or a transmembrane a-helix bundle or channel.
Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near a constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate the interaction between the pore and nucleotides, polynucleotides, or nucleic acids.
The nanopore may be a transmembrane protein pore derived from p-barrel pores or a-helix bundle pores, p-barrel pores comprise a barrel or channel that is formed from p-strands. Suitable p-barrel pores include, but are not limited to, p-toxins, such as a-hemolysin, anthrax toxin and leukocidins, and outer membrane proteins/porins of bacteria, such as Mycobacterium smegmatis porin (Msp), for example MspA, MspB, MspC or MspD, CsgG, outer membrane porin F (OmpF), outer membrane porin G (OmpG), outer membrane phospholipase A and Neisseria autotransporter lipoprotein (NalP) and other pores, such as lysenin. a-helix bundle pores comprise a barrel or channel that is formed from a-helices. Suitable a-helix bundle pores include, but are not limited to, inner membrane proteins and a outer membrane proteins, such as WZA and ClyA toxin.
The nanopore may be a transmembrane pore derived from or based on Msp, a-hemolysin (a-HL), lysenin, CsgG, ClyA, Spl or haemolytic protein fragaceatoxin C (FraC).
The nanopore may be a transmembrane protein pore derived from CsgG, e.g., from CsgG from E. coli Str. K-12 substr. MC4100. Such a pore is oligomeric and typically comprises 7, 8, 9 or 10 monomers derived from CsgG. The pore may be a homo-oligomeric pore derived from CsgG comprising identical monomers. Alternatively, the pore may be a heterooligomeric pore derived from CsgG comprising at least one monomer that differs from the others. Examples of suitable pores derived from CsgG are disclosed in WO 2016/034591,
WO 2017/149316, WO 2017/149317, WO 2017/149318, and WO 2019/002893 (all of which are incorporated herein by reference in their entireties).
The nanopore may be a transmembrane pore derived from lysenin. Examples of suitable pores derived from lysenin are disclosed in WO 2013/153359 (incorporated herein by reference in its entirety).
The nanopore may be a transmembrane pore derived from or based on a-hemolysin (a-HL). The wild type a-hemolysin pore is formed of 7 identical monomers or sub-units (/.e., it is heptameric). An a-hemolysin pore may be a-hemolysin-NN or a variant thereof. The variant preferably comprises N residues at positions Elll and K147.
The nanopore may be a transmembrane protein pore derived from Msp, e.g., from MspA. Examples of suitable pores derived from MspA are disclosed in WO 2012/107778 (incorporated herein by reference in its entirety).
The nanopore may be a transmembrane pore derived from or based on ClyA.
Membrane
The detector or nanopore is typically present in a membrane. Any suitable membrane may be used.
The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer sub-units that are polymerized together to create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit. However, a block copolymer may have unique properties that polymers formed from the individual sub-units do not possess. Block copolymers can be engineered such that one of the monomer sub-units is hydrophobic (/.e., lipophilic), whilst the other sub-unit(s) are hydrophilic whilst in aqueous media. In this case, the block copolymer may possess amphiphilic properties and may form a structure that mimics a biological membrane. The block copolymer may be a diblock (consisting of two monomer sub-units) but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphipiles. The copolymer may be a triblock, tetrablock or pentablock copolymer. The membrane may be a triblock copolymer membrane.
Archaebacterial bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipid forms a monolayer membrane. These lipids are generally found in extremophiles that survive in harsh biological environments, thermophiles, halophiles and acidophiles. Their stability is believed to derive from the fused nature of the final bilayer. It is straightforward to construct block copolymer materials that mimic these biological entities by creating a triblock polymer that has the general motif hydrophilic-hydrophobic- hydrophilic. This material may form monomeric membranes that behave similarly to lipid bilayers and encompass a range of phase behaviours from vesicles through to laminar membranes. Membranes formed from these triblock copolymers hold several advantages over biological lipid membranes. Because the triblock copolymer is synthesised, the exact construction can be carefully controlled to provide the correct chain lengths and properties required to form membranes and to interact with pores and other proteins.
Block copolymers may also be constructed from sub-units that are not classed as lipid submaterials; for example, a hydrophobic polymer may be made from siloxane or other non- hydrocarbon-based monomers. The hydrophilic sub-section of block copolymer can also possess low protein binding properties, which allows the creation of a membrane that is highly resistant when exposed to raw biological samples. This head group unit may also be derived from non-classical lipid head-groups.
Triblock copolymer membranes also have increased mechanical and environmental stability compared with biological lipid membranes, for example a much higher operational temperature or pH range. The synthetic nature of the block copolymers provides a platform to customise polymer-based membranes for a wide range of applications.
The membrane may be one of the membranes disclosed in International Application No. WO2014/064443 or WO2014/064444 (both of which are incorporated herein by reference in their entireties).
The amphiphilic molecules may be chemically modified or functionalised to facilitate coupling of the polynucleotide. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.
Amphiphilic membranes are typically naturally mobile, essentially acting as two-dimensional fluids with lipid diffusion rates of approximately IO-8 cm s4. This means that the pore and coupled polynucleotide can typically move within an amphiphilic membrane.
The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of
substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer, or a liposome. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008/102121, WO 2009/077734, and WO 2006/100484 (incorporated herein by reference in their entireties).
Methods for forming lipid bilayers are known in the art. Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566).
A lipid bilayer may be formed as described in WO 2009/077734 (incorporated herein by reference in its entirety). In this method, the lipid bilayer is formed from dried lipids. A lipid bilayer may be formed across an opening as described in W02009/077734.
The membrane may comprise a solid-state layer. Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as Si3N4, A12O3, and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses. The solid-state layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009/035647 (incorporated herein by reference in its entirety). If the membrane comprises a solid-state layer, the pore is typically present in an amphiphilic membrane or layer contained within the solid-state layer, for instance within a hole, well, gap, channel, trench or slit within the solid-state layer. The skilled person can prepare suitable solid state/amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009/020682 and WO 2012/005857 (incorporated herein by reference in their entireties). Any of the amphiphilic membranes or layers discussed above may be used.
The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell having a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer may comprise other transmembrane and/or intramembrane proteins as well as other molecules in addition to the pore. Suitable apparatus and conditions are discussed below. The method of the invention is typically carried out in vitro.
Polynucleotide binding protein
As those skilled in the art will appreciate, any suitable polynucleotide binding protein can be used in the methods and products of the invention. The polynucleotide binding protein may be any protein that is capable of binding to a polynucleotide and controlling its movement with respect to a detector, e.g., a nanopore.
In more detail, polynucleotide binding proteins such as helicases can typically control the movement of DNA in at least two active modes of operation (when is provided with all the
necessary components to facilitate movement e.g., ATP and Mg2+) and one inactive mode of operation (when not provided with the necessary components to facilitate movement; or when the polynucleotide binding protein is modified in order to prevent the active mode).
When provided with all the necessary components to facilitate movement, a polynucleotide binding protein may move along a polynucleotide such as DNA in either a 5'-3' direction or a 3'-5' direction. Many polynucleotide binding proteins process polynucleotides such as DNA in a 5'-3' direction. Polynucleotide binding proteins which control the movement of polynucleotides in this manner are typically suitable for use in the method of the invention.
However, when a polynucleotide binding protein is not provided with the necessary components to facilitate movement or is modified in order to prevent it from actively controlling the movement of the polynucleotide with respect to the nanopore, it can still passively control the movement of the polynucleotide with respect to the nanopore. For example, the polynucleotide binding protein can bind to the polynucleotide and act as a brake slowing the movement of the polynucleotide when it is pulled into the pore by an applied field (e.g., by the first force in the method of the invention). In the "inactive" mode it typically does not matter whether the DNA is captured either 3' or 5' down (/.e., moves through the nanopore in a 5'-3' direction or in a 3'-5' direction), as the applied force provides the impetus to move the polynucleotide through the nanopore. However, in such embodiments, the polynucleotide binding protein may still control the movement of the polynucleotide with respect to the nanopore e.g., by acting as a brake. When in the inactive mode the movement control of a polynucleotide by a polynucleotide binding protein can be described in a number of ways including ratcheting, sliding, and braking. Typically the method of the invention do not comprise the use of a polynucleotide binding protein operating in the passive mode. However, when a polynucleotide binding protein the polynucleotide binding protein is used, it may be a polynucleotide binding protein operating in the passive mode.
Some methods of the invention may comprise use of a polynucleotide binding protein as a pausing moiety to impede the movement of the polynucleotide strand through the nanopore. The polynucleotide binding protein may be a protein which binds to polynucleotides but which does not have polynucleotide processing capacity, i.e., it is not a polynucleotide binding protein.
A polynucleotide-handling enzyme is a polypeptide that is capable of interacting with a polynucleotide. The enzyme may modify the polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may modify the polynucleotide by orienting it or moving it to a specific position. A polynucleotide binding protein as used herein may be, or may be derived from a polynucleotide handling
enzyme. A polynucleotide binding protein may be, or may be derived from a polynucleotide- handling enzyme.
The polynucleotide binding protein may be derived from a member of any of the Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30 and 3.1.31.
Typically, the polynucleotide binding protein is a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.
The polynucleotide binding protein may be modified to prevent the polynucleotide binding protein disengaging from the polynucleotide. Thus, the target polynucleotide preferably does not disengage from the polynucleotide binding protein.
As used herein, the term "disengaging" refers to the dissociation of the polynucleotide binding protein from the target polynucleotide. Thus, a polynucleotide binding protein may be modified to prevent it from dissociating from the target polynucleotide, e.g., into the reaction medium. It is important to distinguish potential "disengagement" of a polynucleotide binding protein from "unbinding" of a polynucleotide binding protein from a target polynucleotide. As used herein, "unbinding" refers to the transient release of the target polynucleotide the active site of the polynucleotide binding protein (described in more detail herein) but does not imply disengagement. Thus, for example, a polynucleotide binding protein may be modified to prevent the polynucleotide binding protein from disengaging from a polynucleotide, but without preventing the polynucleotide binding protein from unbinding from the polynucleotide. When unbound, the polynucleotide binding protein remains engaged with the target polynucleotide. For example, the polynucleotide binding protein may remain engaged with the target polynucleotide (/.e., it may be prevented from disengaging from the target polynucleotide) because it is topologically closed around the target polynucleotide. The polynucleotide binding site may remain free to bind or unbind the target polynucleotide such that the polynucleotide binding protein may bind or unbind to the target polynucleotide, whilst the polynucleotide binding protein remains engaged with the target polynucleotide. When the polynucleotide binding protein is unbound from the target polynucleotide it may be able to move on (e.g., along) the target polynucleotide under an applied force and may be capable of re-binding to the target polynucleotide. When engaged on the target polynucleotide but unbound from the target polynucleotide, the polynucleotide binding protein is not capable of dissociating from the target polynucleotide.
The polynucleotide binding protein can be adapted to prevent disengagement in any suitable way. For example, the polynucleotide binding protein can be loaded on the polynucleotide and then modified in order to prevent it from disengaging from the
polynucleotide. Alternatively, the polynucleotide binding protein can be modified to prevent it from disengaging from the polynucleotide before it is loaded onto the polynucleotide. Modification of a polynucleotide binding protein and/or a polynucleotide binding protein in order to prevent it from disengaging from a polynucleotide can be achieved using methods known in the art, such as those discussed in WO 2014/013260, which is hereby incorporated by reference in its entirety, and with particular reference to passages describing the modification of polynucleotide binding proteins such as helicases in order to prevent them from disengaging with polynucleotide strands. For example, a polynucleotide binding protein can be modified by treating with tetramethylazodicarboxamide (TMAD). Various other closing moieties are described in WO 2021/255476 (incorporated herein by reference in its entirety).
For example, a polynucleotide binding protein and/or a polynucleotide binding protein may have a polynucleotide-unbinding opening, e.g., a cavity, cleft or void through which a polynucleotide strand may pass when the polynucleotide binding protein disengages from the strand. The polynucleotide-unbinding opening may be the opening through which a polynucleotide may pass when the polynucleotide binding protein disengages from the polynucleotide. The polynucleotide-unbinding opening for a given polynucleotide binding protein can be determined by reference to its structure, e.g., by reference to its X-ray crystal structure. The X-ray crystal structure may be obtained in the presence and/or the absence of a polynucleotide substrate. The location of a polynucleotide-unbinding opening in a given polynucleotide binding protein may be deduced or confirmed by molecular modelling using standard packages known in the art. The polynucleotide-unbinding opening may be transiently produced by movement of one or more parts e.g., one or more domains of the polynucleotide binding protein.
The polynucleotide binding protein may be modified by closing the polynucleotide-unbinding opening. The polynucleotide-unbinding opening may be closed with a closing moiety.
Closing the polynucleotide-unbinding opening may therefore prevent the polynucleotide binding protein from disengaging from the polynucleotide. For example, the polynucleotide binding protein may be modified by covalently closing the polynucleotide-unbinding opening. However, as explained above closing the polynucleotide-unbinding opening does not necessarily prevent the target polynucleotide from unbinding from the polynucleotide binding site of the polynucleotide binding protein. A preferred protein for addressing in this way is a helicase.
The polynucleotide binding protein may be modified with a closing moiety for (i) topologically closing the polynucleotide binding site of the polynucleotide binding protein around the target polynucleotide and (ii) promoting unbinding of the target polynucleotide from the polynucleotide binding site of the polynucleotide binding protein and/or retarding re-binding of the target polynucleotide to the polynucleotide binding site of the
polynucleotide binding protein. The polynucleotide binding protein may be modified in any suitable manner to facilitate attachment of such a closing moiety.
A closing moiety may comprise a bifunctional cross-linking moiety. The closing moiety may comprise a bifunctional cross-linker. The bifunctional crosslinker may attach at two points on the polynucleotide binding protein and close the polynucleotide-unbinding opening of the polynucleotide binding protein thereby preventing disengagement of the polynucleotide from the polynucleotide binding protein whilst allowing unbinding of the polynucleotide from the polynucleotide-binding site of the polynucleotide binding protein.
The closing moiety may attach at any suitable positions on the polynucleotide binding protein. For example, the closing moiety may crosslink two amino acid residues of the polynucleotide binding protein. Typically, at least one amino acid crosslinked by the closing moiety is a cysteine or a non-natural amino acid. The cysteine or non-natural amino acid may be introduced into the polynucleotide binding protein by substitution or modification of a naturally occurring amino acid residue of the polynucleotide binding protein. Methods for introducing non-natural amino acids are well known in the art and include for example native chemical ligation with synthetic polypeptide strands comprising such non-natural amino acids. Methods for introducing cysteines into a polynucleotide binding protein are likewise within the capability of one of skill in the art, for example using techniques disclosed in references such as Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).
The closing moiety may have a length of from about 1 A to about 100 A. The length of the closing moiety may be calculated according to static bond lengths or more preferably using molecular dynamics simulations. The length may for example be from about 2 A to about 80 A, such as from about 5 A to about 50 A, e.g., from about 8 to about 30 A such as from about 10 to about 25 A or about 20 A, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 A.
Polynucleotide binding proteins suitable for being closed using a closing moiety as described above are discussed in more detail herein. The polynucleotide binding protein is preferably a helicase, e.g., a Dda helicase as described herein.
The polynucleotide binding protein may be or may be derived from an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli, exonuclease III enzyme from E. coli, RecJ from T. thermophilus and bacteriophage lambda exonuclease, TatD exonuclease and variants thereof.
The polynucleotide binding protein may be a polymerase. The polymerase may be PyroPhage® 3173 DNA Polymerase (which is commercially available from Lucigen®
Corporation), SD Polymerase (commercially available from Bioron®), Klenow from NEB or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase or a variant thereof. Modified versions of Phi29 polymerase that may be used in the invention are disclosed in US Patent No. 5,576,204.
The polynucleotide binding protein may be a topoisomerase. In one embodiment, the topoisomerase is a member of any of the Moiety Classification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase may be a reverse transcriptase, which are enzymes capable of catalysing the formation of cDNA from a RNA template. They are commercially available from, for instance, New England Biolabs® and Invitrogen®.
The polynucleotide binding protein is preferably a helicase. Any suitable helicase can be used in accordance with the method of the invention. For example, the or each enzyme used in accordance with the present disclosure may be independently selected from a Hel308 helicase, a RecD helicase, a Tral helicase, a TrwC helicase, an XPD helicase, and a Dda helicase, or a variant thereof. Monomeric helicases may comprise several domains attached together. For instance, Tral helicases and Tral subgroup helicases may contain two RecD helicase domains, a relaxase domain and a C-terminal domain. The domains typically form a monomeric helicase that is capable of functioning without forming oligomers. Particular examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pifl and Tral. These helicases typically work on single stranded DNA. Examples of helicases that can move along both strands of a double stranded DNA include FtfK and hexameric enzyme complexes, or multisubunit complexes such as RecBCD. The polynucleotide binding protein is preferably a Dda (DNA-dependent ATPase) helicase.
Hel308 helicases are described in publications such as WO 2013/057495, the entire contents of which are incorporated by reference. RecD helicases are described in publications such as WO 2013/098562, the entire contents of which are incorporated by reference. XPD helicases are described in publications such as WO 2013/098561, the entire contents of which are incorporated by reference. Dda helicases are described in publications such as WO 2015/055981 and WO 2016/055777, the entire contents of each of which are incorporated by reference.
The helicase may be Trwc Cba or a variant thereof, Hel308 Mbu or a variant thereof or Dda or a variant thereof. Variants may differ from the native sequences in any of the ways discussed herein. An example variant of Dda comprises E94C/A360C. A further example variant of Dda comprises E94C/A360C and then (AM1)G1G2 (/.e., deletion of Ml and then addition of G1 and G2).
General methods
As mentioned above, the method of the invention may be operated using any suitable detector, and as such any suitable apparatus for detecting polynucleotides can be used.
The method of the invention may be carried out using any apparatus that is suitable for nanopore sensing. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier may have an aperture in which a membrane containing a transmembrane pore is formed. Transmembrane pores are described herein.
The methods may be carried out using the apparatus described in WO 2008/102120, WO 2010/122293, or WO 00/28312 (incorporated herein by reference in their entireties). In brief, the binding of a molecule (e.g., a target polynucleotide) in the channel of a pore will have an effect on the open-channel ion flow through the pore, which is the essence of "molecular sensing" of pore channels. Variation in the open-channel ion flow can be measured using suitable measurement techniques by the change in electrical current. The degree of reduction in ion flow, as measured by the reduction in electrical current, is related to the size of the obstruction within, or in the vicinity of, the pore. Binding of a molecule of interest (e.g., the target polynucleotide) in or near the pore therefore provides a detectable and measurable event, thereby forming the basis of a "biological sensor". Detecting the presence of biological molecules finds application in personalised drug development, medicine, diagnostics, life science research, environmental monitoring and in the security and/or the defence industry.
When used to characterize the polynucleotide, the presence, absence or one or more characteristics of the target polynucleotide are determined. The methods may be for determining the presence, absence or one or more characteristics of at least one target polynucleotide. The methods may concern determining the presence, absence or one or more characteristics of two or more target polynucleotide. The methods may comprise determining the presence, absence or one or more characteristics of any number of target polynucleotides, such as 2, 5, 10, 15, 20, 30, 40, 50, 100 or more target polynucleotides. Any number of characteristics of the one or more target polynucleotides may be determined, such as 1, 2, 3, 4, 5, 10 or more characteristics. Characteristics amenable to being detected in the methods provide herein include the identity or sequence of the polynucleotide, the length, of the polynucleotide, whether or not the polynucleotide is modified, etc. In some embodiments the method of the invention are methods of sequencing the barcoded constructs. In some embodiments the sequences of the barcoded constructs may be determined in real-time by aligning real-time signal or basecalling to known references. Exemplary methods of determining a polynucleotide sequence are described in WO 2016/059427 (incorporated herein by reference in its entirety).
When used to characterize the polynucleotide, the methods may involve measuring the ion current flow through the pore, typically by measurement of a current. Alternatively, the ion flow through the pore may be measured optically, such as disclosed by Heron et al: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore, the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore. The characterisation methods may be carried out using a patch clamp or a voltage clamp. The characterisation methods preferably involve the use of a voltage clamp.
The methods may involve measuring an optical signal as described in Chen et al, Nature Communications (2018)9: 1733, the entire contents of which are hereby incorporated by reference. For example, a nanopore such as an optically engineered nanopore structure (e.g., a plasmonic nanoslit) may be used to locally enable single-molecule surface enhanced Raman spectroscopy (SERS) to allow the characterisation of the polynucleotide through direct Raman spectroscopic detection.
The methods may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.
The methods may involve the measuring of a current flowing through the pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV. The voltage used is preferably in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20mV and 0 mV and an upper limit independently selected from + 10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more preferably in the range 100 mV to 240mV and most preferably in the range of 120 mV to 220 mV. It is possible to increase discrimination between different nucleotides by a pore by using an increased applied potential.
The methods are typically carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3-methyl imidazolium chloride. In the exemplary apparatus discussed above, the salt is present in the aqueous solution in the chamber. Potassium chloride (KCI), sodium chloride (NaCI) or caesium chloride (CsCI) is typically used. KCI is preferred. The salt may be an alkaline earth metal salt such as calcium chloride (CaCI2). The salt concentration may be at saturation. The salt concentration may be 3M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is preferably from 150 mM to 1 M.
The method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. High salt concentrations provide a high signal to noise ratio and allow for currents indicative of binding/no binding to be identified against the background of normal current fluctuations.
The methods are typically carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCI buffer. The methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used is preferably about 7.5.
The methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C. The methods are typically carried out at room temperature. The methods are optionally carried out at a temperature that supports enzyme function, such as about 37 °C.
Any of the proteins described herein, such as the protein pores, may be made synthetically or by recombinant means. For example, the pore may be synthesised by in vitro translation and transcription (IVTT). The amino acid sequence of the pore may be modified to include non-naturally occurring amino acids or to increase the stability of the protein. When a protein is produced by synthetic means, such amino acids may be introduced during production. The pore may also be altered following either synthetic or recombinant production.
Any of the proteins described herein, such as the protein pores, can be produced using standard methods known in the art. Polynucleotide sequences encoding a pore or construct may be derived and replicated using standard methods in the art. Polynucleotide sequences encoding a pore or construct may be expressed in a bacterial host cell using standard techniques in the art. The pore may be produced in a cell by in situ expression of the polypeptide from a recombinant expression vector. The expression vector optionally carries an inducible promoter to control the expression of the polypeptide. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.
The pore may be produced in large scale following purification by any protein liquid chromatography system from protein producing organisms or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA systems, the Bio-Cad system, the Bio-Rad BioLogic system, and the Gilson HPLC system.
Kits and Systems
The invention also provides a kit for uniquely labelling RIMA molecules in a population of cells. The kit is preferably for producing uniquely labelled constructs transcribed from the RNA molecules in a population of cells. The uniquely labelled constructs transcribed from the RNA molecules are preferably uniquely labelled constructs comprising DNA sequences transcribed from the RNA molecules.
The kit comprises two or more first adaptors, wherein the two or more first adaptors comprise primer sites for RT, and wherein each of the two or more first adaptors comprises the reverse complement of a different first barcode. The kit preferably comprises at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 first adaptors. The kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536 or at least about 3456 first adaptors.
Any of the embodiments discussed above with reference to the first adaptors used in the method of the invention equally apply to the kits. The kit may comprise any number of first adaptors. The two or more first adaptors are preferably double stranded and comprise overhangs which are capable of hybridising to the 3' ends, the poly(A) tails or the polyadenylated 3' ends of the RNA molecules. The nucleotides at the 3' ends of one of both strands of the double stranded first adaptors are preferably ddC or inverted dT.
The kit preferably further comprises (a) two or more RT primers capable of hybridising to the two or more first adaptors and each comprising a different second barcode and/or (b) two or more second adaptors each comprising a different third barcode. The kit preferably comprises at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000 or at least about 5000 first adaptors. The kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536 or at least about 3456 second adaptors and/or RT primers. The kit preferably comprises the same number of first adaptors, RT primers and second adaptors. Any of the embodiments discussed above with reference to the RT primers and second adaptors used in the method of the invention equally apply to the kits.
Any of the embodiments discussed above with reference to the barcodes used in the method of the invention equally apply to the kits. Each barcode in the kit is preferably different from any other barcode or all the other barcodes in the kit.
The kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536 or at least about 3456 first adaptors, second adaptors and RT primers.
The invention also provides a system for conducting the method of the invention. The system is for uniquely labelling RNA molecules in a population of cells. The system is preferably for producing uniquely labelled constructs transcribed from the RNA molecules in a population of cells. The uniquely labelled constructs transcribed from the RNA molecules are preferably uniquely labelled constructs comprising DNA sequences transcribed from the RNA molecules.
The system comprises (a) two or more first adaptors, wherein the two or more first adaptors comprise primer sites for RT, and wherein each of the two or more first adaptors comprises the reverse complement of a different first barcode, (b) two or more RT primers capable of hybridising to the two or more first adaptors and each comprising a different second barcode, (c) two or more second adaptors each comprising a different third barcode and (d) a nanopore. Any of the embodiments discussed above with reference to the method of the invention equally apply to the system of the invention. The system may comprise any of the numbers of first adaptors, RT primers and second adaptors discussed above with reference to the kits of the invention. The system preferably comprises at same number of first adaptors, RT primers and second adaptors.
The nanopore is preferably present in a membrane. Suitable membranes are discussed above. The system may comprise any of the membranes disclosed above, such as an amphiphilic layer, a triblock copolymer membrane or a solid-state layer. The membrane is typically part of an array of membranes, wherein each membrane preferably comprises a nanopore. The array may be any of those described in WO 2018/060740 (incorporated herein by reference in its entirety).
The system is preferably adapted to apply a voltage across the membrane and to take one or more electrical measurements. Suitable adaptations are discussed in WO 2018/060740 (incorporated herein by reference in its entirety).
The kit or system preferably further comprises one or more sequencing adaptors. The one or more sequencing adaptors may be any of those discussed above with reference to the method of the invention.
The kit or system may further comprise a polynucleotide binding protein. The kit or system may further comprise a microparticle. Any of the embodiments discussed above with reference to the method of the invention equally apply to the system of the invention.
The kit or system may additionally comprise one or more other reagents or instruments which enable any of the embodiments mentioned above to be carried out. Such reagents or instruments include one or more of the following: suitable buffer(s) (aqueous solutions), means to obtain a sample from a subject (such as a vessel or an instrument comprising a needle), means to amplify and/or express polynucleotides, a membrane as defined above or voltage or patch clamp apparatus. Reagents may be present in the kit or system in a dry state such that a fluid sample is used to resuspend the reagents. The kit or system may also, optionally, comprise instructions to enable the kit or system to be used in the methods described herein or details regarding for which organism the method may be used. The kit or system may comprise a magnet or an electromagnet. The kit or system may, optionally, comprise nucleotides.
The following Examples illustrate the invention. It is to be understood that although particular embodiments, specific configurations as well as materials and/or molecules, have been discussed herein for methods according to the invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The following examples are provided to better illustrate particular embodiments, and they should not be considered limiting the application. The application is limited only by the claims.
EXAMPLE 1
This Example describes one exemplary embodiment of the method of the invention. This is called the ligation + RT + ligation (LRL) approach.
The LRL approach applies the first-round barcode to separate sample of cells through ligation of a barcoded cDNA reverse transcription adaptor (CDRA) (Figure 1). After ligation USER enzyme may be used to digest the bottom strand and reveal the primer site for the reverse transcription. The ligation and subsequent USER digestion requires incubation at 37°C for a total of 45 minutes.
Cells are pooled then split out for the second-round barcoding which is applied through a barcoded RT primer which anneals to the top strand of the first adaptor (Figure 2). In the LRL approach, a strand switching primer (SSP) may be used which contains a UMI sequence optimised to reduce homopolymers through a TTVVVVTT structure to facilitate detection through our nanopores (Figure 2). The SSP facilities deduplication of PCR replicates.
After reverse transcription, cells are pooled and then split out for the third-round barcode which is applied through ligation of a barcoded adaptor (Figure 3).
In the LRL approach, the first-round ligation is separated from the third-round ligation by the second-round RT step. This minimises the possibility of barcode crosstalk between the ligations. Regardless, a third-round adaptor blocker has been developed that is fully complementary to the adapter overhang. The blocker itself is 5'phosphorylated and should therefore ligate to surplus adapter permanently once annealed in position (Figure 4).
The fully barcoded molecule is then enriched through biotin-streptavidin pulldown and amplified with PCR primers. These primers may enable the use of click chemistry moieties for rapid attachment chemistry for expedited library prep. Moreover, it is possible to use barcoded universal primers for a potential fourth-round of barcoding. EXAMPLE 2
Oligonucleotides utilised
Resuspend oligos in RNase-free TE buffer (10 mM Tris.HCI, 1 mM EDTA pH 8.0 final) to 100 pM.
Methanol fixation of cells
Based on Chen, J., Cheung, F., Shi, R., Zhou, H., Lu, W., Candia, J., Kotliarov, Y., Stagliano, K. R., & Tsang, J. S. (2018). PBMC fixation and processing for Chromium single-cell RNA sequencing. Journal of Translational Medicine, 16(1). https://doi.org/10.1186/sl2967-018- 1578-4.
1. The process of cell fixation and subsequent rehydration results in an approximate 50% loss of cells. Proceed with a quantity of cells appropriate for the scale of experiment required. For example, if 1E6 cells are required, start with 2E6 cells for methanol fixation.
2. Centrifuge cells at 300 g, 5 min 4°C.
3. Discard supernatant and resuspend in 1 mL of ice-cold IX DPBS (Dulbecco's Phosphate Buffered Saline, Gibco, 14080089) using a wide bore pipette tip.
4. Transfer cells to centrifuge tube appropriate for the input cells quantity:
5. Centrifuge 300 g, 5 min 4°C.
6. Discard supernatant and resuspend in 1 mL of ice-cold IX DPBS (Gibco, 14190144) using a wide bore pipette tip.
7. Remove supernatant, resuspend with ice-cold IX DPBS (Gibco, 14190144) to achieve approximately 5E3 cells/pL. For example, 1E6 cells requires 200 pL of IX DPBS (Gibco, 14190144).
8. Transfer cell suspension to a fume hood.
9. Relative to the volume of resuspended cells, take 4x volumes of ice-cold 100% methanol and add drop by drop, gently swirling continuously to achieve 80% methanol final.
10. Chill at -20°C for a minimum of 30 minutes
11. Store at -80°C, or at -20°C for up to 6 weeks. Alternatively proceed immediately to the following steps.
Preparation of barcode oliqos
The steps below describe either a lx, 50x or lOOx scale experiment, this is relative to the input material for the combinatorial barcoding of 10 K, 500 K or 1 M cells respectively. Proceed with the scale of experiment required.
12. Prepare the oligo pools:
13. Take the CDRA adapter and Round 3 adapters, heat to 95c 5 min, cool -0.1°C per second to anneal the strands. Store at 4°C.
Rehydration 14. Prepare rehydration buffer according to quantity of cells:
15. Place the fixed cells on ice to equilibrate 5 min.
16. Centrifuge at 1000 g, 5 min 4°C.
17. Remove the supernatant and resuspend in 1000 pL of rehydration buffer. 18. Centrifuge at 1000 g, 5 min 4°C.
19. Remove supernatant. Assume 50% of the cells have been lost, resuspend the cell pellet in rehydration buffer to achieve 1250 cell/pL. For example, 1E6 cell input would require 400 pL of rehydration buffer.
CDRA ligation (first-round barcoding) 20. Prepare the CDRA ligation mix as below:
0.5E6 1E6
21. On ice, prepare a new plate with 4 pL/well of CDRA barcode.
22. Add the CDRA ligation mix to each of the CDRA adapters:
CDRA ligation 10E3 cells
23. Incubate 37°C for 30 min.
24. Now add the following components to expose the RT priming site:
CDRA digestion 10E3 cells
25. Incubate 37°C for 15 min.
26. Pool and centrifuge at 1000 g, 5 min 4°C.
27. Remove the supernatant and resuspend in 1000 pL of resuspension buffer.
In-situ reverse transcription (second-round barcoding)
28. Prepare the RT mix:
29. On ice, prepare a new plate with 4 pL/well of RT oligo mix.
30. Place the RT barcode plate on a chilled rack and add the following to each well in order: in situ RT 10E3 cells
31. Seal the plate and then briefly centrifuge.
32. Incubate at 42°C for 90 minutes then chill to 4°C immediately.
33. Briefly centrifuge the plate, then place on a chilled rack.
34. Pool the wells together into a single 1.5 mL microcentrifuge tube on ice.
35. Add 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) to the pooled reaction for 0.1% (v/v) ECOSURF EH9 final.
36. Centrifugate at 1000 g, 5 min at 4°C. 37. Discard supernatant and resuspend cells in ice-cold IX NEB buffer r3.1 (NEB,
B6003S) to 500 cells/pL.
Ligation of the third barcode (third-round barcoding)
38. On ice, prepare a new plate with 10 pL/well of Round 3 ligation barcode
39. On ice, add the following components to the resuspended cells to make the third- round ligation mix:
Third round ligation 1E6 mix 10E3 cells 0.5E6 cells cells
40. Add 40 pL of the mixture to each individual well containing 10 pL ligation barcode, mixing by gentle pipetting. Seal the plate.
Third round ligation 10E3 cells
41. Incubate at 37°C for 30 min.
42. Prepare the third-round blocking solution:
Third round blocker 1E6 solution 10E3 cells 0.5E6 cells cells
43. Add 20 |jL third round blocking solution per well:
Third round blocking 10E3 cells
44. Incubate on ice 5 min.
45. Pool cells into a single tube. Lysis
46. Add 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) to the pooled reaction for 0.1% (v/v/) ECOSURF EH9 final.
47. Centrifuge the cells at 1000 g for 5 min at 4°C.
48. Make the washing buffer:
Washing buffer
49. Carefully remove the supernatant. Wash the cells in 1 mL of washing buffer without disturbing the pellet.
50. Centrifuge at 1000 g for 5 min at 4°C.
51. Remove supernatant without disturbing the pellet. Resuspend cells in 30 pL washing buffer per
52. Mix 5 pL of cell suspension with 5 pL Trypan blue to count cells. Expect a few thousand cells in total. The cells can be divided into sub-libraries if needed.
53. Add 30 pL lysis buffer to the remaining 25 pL cell suspension.
10E3 0.5E6
Lysis buffer cells cells 1E6 cells
54. Incubate at 55°C for 1 hour.
55. Freeze the lysate at -20°C, if needed. cDNA purification
56. Take 5 pL lysate and dilute it to 25 pL with nuclease-free H2O.
57. Purify it with NEB monarch PCR clean up kit (NEB, T1030S) using a 5: 1 buffer:sample ratio as described by the manufacturer.
58. Elute in 52 pL nuclease-free- H2O. Quantify the eluate by Qubit fluorometer (Invitrogen).
Pull-down
Preparing the beads
59. Prepare 4 mL of 2X wash/bind buffer (10 mM Tris.HCI pH 7.5, 2 M NaCI, 1 mM EDTA).
60. Use 3.5 mL of 2X wash/bind buffer and add 3.5 mL of nuclease-free H2O to make 7 mL of IX wash/bind buffer (5 mM Tris.HCI pH 7.5, 1 M NaCI, 0.5 mM EDTA).
61. Take 25 pL of stock MyOne Cl streptavidin beads (10 pg beads/pL in PBS pH 7.4, 0.1% BSA, 0.02% sodium azide, 65001 Invitrogen), ensure the stock is well resuspended.
62. Wash the beads 3 times with 1 mL of IX wash/bind buffer, vortex for 5 sec and then collect on magnet for 2 minutes between each wash.
63. Resuspend the beads in 50 pL (twice the original volume of beads) of 2X wash/bind buffer to achieve a final bead concentration of 5 pg/pL of beads.
Sample binding
64. Take 50 pL of biotinylated cDNA and combine with 50 pL of 5 pg/pL prepared beads (250 pg beads total) to achieve a final buffer concentration of 1 M NaCI, optimal for binding.
65. Incubate at room temperature for 20 minutes with gentle rotation.
66. Wash the beads 3 times with 1 mL of IX wash/bind buffer. Gently vortex 5 seconds and collect on magnet for 3 minutes between each wash. Discard the supernatant after each wash. Take care not to aspirate any of the beads.
67. Wash the beads once in 200 pL of 10 mM Tris.HCI pH 7.5, vortex 5 seconds, briefly spin the sample down and collect on magnet for 3 minutes. Discard the supernatant.
68. Resuspend pellet in 80 pL of nuclease-free H2O, vortex 5 seconds and then briefly spin down to collect the amplicon-bead conjugate.
Post-pull-down PCR
69. Set up the PCR reaction below, split across 2 aliquots of 100 pL:
Run in a thermocycler with the parameters below:
SPRI clean
70. Pool the PCR reactions.
71. Carry out a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), wash the pellet twice with 70% EtOH without resuspension.
72. Elute in 60 pL nuclease-free H2O and quantify with Qubit Fluorimeter (Invitrogen).
Sequencing library preparation
73. Continue to prepare a sequencing library using Oxford Nanopore Technologies SQK- LSK114 ligation sequencing kit. Follow the manufacturer's protocol and use 200 fmols of amplified cDNA as input. Sequence on a PromethlON with a FLO-PRO114M flowcell.
EXAMPLE 3 - RNA INTEGRITY
Integrity during methanol fixation
To test whether methanol fixation was suitable for preserving RNA integrity, a cell culture stock of GM12878 (Coriel Institute, GM12878) was fixed in methanol for testing.
■ 2E7 cells GM12878 were washed in 1 mL of ice-cold IX DPBS (Gibco, 14080089) and pelleted at 300 g. The supernatant was aspirated, pellet washed and pelleted once again.
■ The supernatant was aspirated, resuspending the pellet to approximately 5E3 cells/pL in ice-cold IX DPBS.
■ The cell suspension was then mixed with 4x volumes ice-cold 100% methanol to achieve a final concentration of 80% methanol. The suspension was chilled at -20°C for 30 minutes.
■ The methanol fixed cell suspension was then split into 4 equal aliquots of 5E6 fixed cell, incubated on ice for 5 minutes, then pelleted at 1000 g.
■ The supernatant was aspirated, and the cell pellet was resuspended in a rehydration buffer of either 3X SSC (Invitrogen, AM9770) or IX DPBS (Gibco, 14080089) supplemented with 1 mM DTT (Sigma, 43816), 200 ng/pL recombinant albumin (NEB, B9200s), and 0.2 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 1 mL final.
■ The cell suspension was pelleted again at 1000 g.
■ The supernatant was aspirated, then the cell pellet (assuming a 50% loss of cells) was resuspended in 2000 pL in a rehydration of the same composition as previous.
■ Rehydrated cells were then pelleted at 1000 g, the supernatant was aspirated.
■ Total RNA was then immediately extracted from the cell pellet following a TRIzol reagent (Invitrogen, 15596026) based extraction following the method described by Workman et al. (2018). Briefly, 400 pL of TRIzol reagent was added to each sample, incubating for 5 minutes at room temperature. 80 pL of chloroform was then added and the sample vortexed before incubating for 5 minutes at room temperature. The sample was vortexed then pelleted at 2000 g, 4°C. The supernatant was transferred to a 5 mL centrifuge tube and mixed with an equal volume of isopropanol. The tube mixed with inversion and incubated at room temperature for 15 minutes before centrifugation at 2000 g, 4°C for 20 minutes. The supernatant was aspirated, and the pellet washed with 400 pL of ice-cold 70% ethanol before further centrifugation at 2000 g, 4°C for 5 minutes. The supernatant was aspirated, and the pellet resuspended in 10 pL of TE buffer.
■ RNA integrity was then assessed using the Agilent 2100 Bioanalyzer (Agilent, G2939BA) with the RNA Nano 6000 kit and chip (Agilent, 5067-1511).
The results are shown in Figure 5.
Integrity during incubation
To test whether RNA integrity would remain stable during the 37°C incubation steps required for the ODRA adapter ligation and USER digestion in the first round of combinatorial barcoding in the LRL approach, a cell culture stock of GM12878 (Coriel Institute, GM12878) was fixed in methanol for testing.
■ 2E7 cells GM12878 were washed in 1 mL of ice-cold IX DPBS (Gibco, 14080089) and pelleted at 300 g. The supernatant was aspirated, pellet washed and pelleted once again.
■ The supernatant was aspirated, and then the pellet resuspended to approximately 5E3 cells/pL in ice-cold IX DPBS.
■ The cell suspension was then mixed with 4x volumes ice-cold 100% methanol to achieve a final concentration of 80% methanol. The suspension was chilled at -20°C for 30 minutes.
■ A fraction of the methanol fixed cell suspension was then split into 2 equal aliquots of 5E6 fixed cells and incubated on ice for 5 minutes then pelleted at 1000 g.
■ The supernatant was aspirated, and the cell pellet was resuspended in a rehydration buffer of 3X SSC (Invitrogen, AM9770) supplemented with 1 mM DTT (Sigma, 43816), 200 ng/pL recombinant albumin (NEB, B9200s), and 0.2 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 1 mL final.
■ The cell suspension was pelleted again at 1000 g.
■ The supernatant was aspirated and then the cell suspension (assuming a 50% loss of cells) was resuspended in 2000 pL in a rehydration of the same composition as previous.
■ At this point one aliquot of cells was immediately extracted as below for a nonincubated control. The second aliquot was incubated at 37°C for 45 minutes, as would be required for the CDRA ligation and USER digestion for the LRL method.
■ Rehydrated cells were then pelleted at 1000 g, the supernatant was aspirate.
■ Total RNA was then immediately extracted from the cell pellet following a TRIzol reagent (Invitrogen, 15596026) based extraction following the method described by Workman et al. (2018). Briefly, 400 pL of TRIzol reagent was added to each sample, incubating for 5 minutes at room temperature. 80 pL of chloroform was then added and the sample vortexed before incubating for 5 minutes at room temperature. The sample was vortexed then pelleted at 2000 g, 4°C. The supernatant was transferred to a 5 mL centrifuge tube and mixed with an equal volume of isopropanol. The tube mixed with inversion and incubated at room temperature for 15 minutes before centrifugation at 2000 g, 4°C for 20 minutes. The supernatant was aspirated, and the pellet washed with 400 pL of ice-cold 70% ethanol before further centrifugation at 2000 g, 4°C for 5 minutes. The supernatant was aspirated, and the pellet resuspended in 10 pL of TE buffer.
■ RNA integrity was then assessed using the Agilent 2100 Bioanalyzer (Agilent, G2939BA) with the RNA Nano 6000 kit and chip (Agilent, 5067-1511).
The results are shown in Figure 6.
EXAMPLE 4 - BARCODE CROSSTALK
To assess manipulation of RIMA on a molecular level, we used a specific synthetic RNA analyte derived from the enolase II gene of Saccharomyces cerevisiae (ENO2, YHR174W SGDID: S000001217, chrVIII :451327..452640). The ENO2 gene was amplified from S. cerevisiae genomic DNA using primers with a gene specific portion complimentary to the first and last 22 nt of the coding sequence and a 5' flanking portion of 27 nt of exogenic sequence for downstream application (Table 1).
Table 1, YHR174W primers; the enolase II gene specific portion (green) and 5' flank (blue).
The amplicon product was then further amplified using an upstream primer flanked with a 5' T7 RNA polymerase promoter sequence and a downstream primer flanked with a 5' poly(T) tail to yield an in vitro transcription template (see Table 2).
Table 2, IVT template primers; the T7 RNAP promoter sequence (red) and 3'poly(T) tail
(yellow). Transcription start site underlined.
The in vitro transcription template was then used for an in vitro transcription reaction using T7 RNAP to yield the RNA control strand (RCS) test analyte (SEQ ID NO: 18), as below:
>RCS
GGGAGAGCCATCAGATTGTGTTTGTTAGTCGCTATGGCTGTCTCTAAAGTTTACGCTAGATCCGTCT
ACGACTCCCGTGGTAACCCAACCGTCGAAGTCGAATTAACCACCGAAAAGGGTGTTTTCAGATCCAT
TGTTCCATCTGGTGCCTCCACCGGTGTCCACGAAGCTTTGGAAATGAGAGATGAAGACAAATCCAAG
TGGATGGGTAAGGGTGTTATGAACGCTGTCAACAACGTCAACAACGTCATTGCTGCTGCTTTCGTCA
AGGCCAACCTAGATGTTAAGGACCAAAAGGCCGTCGATGACTTCTTGTTGTCTTTGGATGGTACCGC
CAACAAGTCCAAGTTGGGTGCTAACGCTATCTTGGGTGTCTCCATGGCCGCTGCTAGAGCCGCTGCT
GCTGAAAAGAACGTCCCATTGTACCAACATTTGGCTGACTTGTCTAAGTCCAAGACCTCTCCATACGT
TTTGCCAGTTCCATTCTTGAACGTTTTGAACGGTGGTTCCCACGCTGGTGGTGCTTTGGCTTTGCAAG
AATTCATGATTGCTCCAACTGGTGCTAAGACCTTCGCTGAAGCCATGAGAATTGGTTCCGAAGTTTAC
CACAACTTGAAGTCTTTGACCAAGAAGAGATACGGTGCTTCTGCCGGTAACGTCGGTGACGAAGGT
GGTGTTGCTCCAAACATTCAAACCGCTGAAGAAGCTTTGGACTTGATTGTTGACGCTATCAAGGCTG
CTGGTCACGACGGTAAGGTCAAGATCGGTTTGGACTGTGCTTCCTCTGAATTCTTCAAGGACGGTAA
GTACGACTTGGACTTCAAGAACCCAGAATCTGACAAATCCAAGTGGTTGACTGGTGTCGAATTAGCT
GACATGTACCACTCCTTGATGAAGAGATACCCAATTGTCTCCATCGAAGATCCATTTGCTGAAGATGA
CTGGGAAGCTTGGTCTCACTTCTTCAAGACCGCTGGTATCCAAATTGTTGCTGATGACTTGACTGTCA
CCAACCCAGCTAGAATTGCTACCGCCATCGAAAAGAAGGCTGCTGACGCTTTGTTGTTGAAGGTTAA
CCAAATCGGTACCTTGTCTGAATCCATCAAGGCTGCTCAAGACTCTTTCGCTGCCAACTGGGGTGTT
ATGGTTTCCCACAGATCTGGTGAAACTGAAGACACTTTCATTGCTGACTTGGTTGTCGGTTTGAGAAC
TGGTCAAATCAAGACTGGTGCTCCAGCTAGATCCGAAAGATTGGCTAAGTTGAACCAATTGTTGAGA
ATCGAAGAAGAATTGGGTGACAAGGCTGTCTACGCCGGTGAAAACTTCCACCACGGTGACAAGTTG TAACATCGTCGTGAGTAGTGAACCGTAAGCAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA
First-round barcode
There is the possibility that barcodes in the first round of barcoding could be shared between cells when cells are pooled together. To investigate this, we carried out a barcode competition assay in vitro using the RCS test analyte.
■ Using the assumption that 1 cell yields approximately 1.25 pg of poly adenylated RIMA, and that each single-plex workflow requires 10 K cells, 12.5 ng of RCS test analyte was used as template for the in vitro barcode competition assay.
■ 5 reactions each composed of 12.5 ng RCS synthetic analyte was resuspended in a rehydration buffer of 3X SSC (Invitrogen, AM9770) supplemented with 1 mM DTT (Sigma, 43816), 200 ng/pL recombinant albumin (NEB, B9200S), and 0.2 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 16 pL final to mimic an input of 10 K rehydrated cells.
■ 2 reactions were then mixed with CDRA BC01 (2.27 pM SEQ ID NO: 1 top strand hybridised with 2.27 pM SEQ ID NO: 3 bottom strand), 2 reactions mixed with CDRA BC02 (2.27 pM SEQ ID NO: 2 top strand with 2.27 pM SEQ ID NO: 3 bottom strand), 1 reaction mixed with nuclease-free water as a control for no CDRA adapter ligation, and 0.625 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694), IX T4 ligation buffer (NEB, M0202L) and 25 U/pL T4 ligase (NEB, M0202L) final in 20 pL and incubated for 30 minutes at 37°C. Meanwhile, separate competing barcoded CDRA adapter ligation reactions were assembled as above without RNA template and incubated on ice.
■ Following ligation, each reaction was supplemented with 0.25 U/pL Lambda exonuclease (NEB, M0262S) and 0.05 U/pL USER mix (NEB, M5505S) final in 22 pL and incubate for 15 minutes at 37°C, then 5 minutes on ice.
■ To simulate potential competition between CDRA barcodes during pooling, RNA template CDRA ligation reactions were either mixed with an equal volume of reaction without competing CDRA adapter (no competition), or with an equal volume reaction with competing CDRA adapter (competition). The pooled reactions were then incubated on ice for 5 minutes.
■ To simulate cell pooling, centrifugation and resuspension, each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
■ Each reaction was then supplemented with 25 pM RT primer BC03 (SEQ ID NO: 4), 25 pM strand switching primer (SEQ ID NO: 6), IX RT buffer (ThermoFisher Scientific, EP0753), 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694),
500 pM dNTPs (NEB, N0447S), 10 U/|_il_ Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753) in 20 pL final and incubated for 90 minutes at 42°C, then 5 minutes on ice.
■ Each reaction was then supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) to the pooled reaction for 0.1% (v/v) ECOSURF EH9 final. To simulate cell pooling, centrifugation and resuspension, each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 20 pL of IX NEB buffer r3.1 (NEB, B6003S).
■ Each reaction was then supplemented with ligation barcode BC05 (2.16 pM SEQ ID NO: 9 top strand hybridised with 2.4 pM SEQ ID NO: 7 bottom strand) IX T4 ligation buffer (NEB, M0202L), 20 U/pL T4 ligase (NEB, M0202L) in 50 pL final and incubated for 30 minutes at 37°C.
■ Each reaction was then supplemented with 3.24 pM blocking oligo (SEQ ID NO: 11) and 31.25 mM EDTA in 70 pL final and incubated on ice for 5 minutes.
■ To simulate cell lysis, each reaction was supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) then subjected to a 1.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 10 pL nuclease-free water.
■ Samples were amplified over 16x cycles using universal primers (SEQ ID NO: 12 and SEQ ID NO: 13) with Hotstart LongAmp Taq (NEB, M0533S). Amplicons were subjected to a 0.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 50 pL nuclease-free.
■ 200 fmols amplified cDNA prepared for sequencing using Oxford Nanopore Technologies SQK-LSK114 ligation sequencing kit, and sequence on a PromethlON with a FLO-PRO114M flowcell for each sample. Reads were aligned to a reference of RCS using the MinKNOW software.
Looking at the results in Figure 7 where there was no competing CDRA adapter added there was only background misaligned detection of the competing barcode (<4%). However, where there was competing barcode added, there was still the same misaligned detection of the competing barcode. These data indicate there was no significant crosstalk between the CDRA barcodes, most likely due to the USER enzyme mix digestion, which destroys the bottom strand of the adapter that is required to splint the ligation of the 3' terminus of the RNA poly(A) tail to the 5' terminus of the CDRA adapter top strand.
Where no CDRA adapter was added to the ligation prior to the addition of competing barcode on sample pooling, there was 11.4% detection of the competing background, greater than the background detection in the other samples. However, the PCR yield of this reaction was 8.5-fold less than that of the other samples, indicating that there was still very little crosstalk of the adapter upon sample pooling.
Together these results show that the design and application of the CDRA adapter for the first round of combinatorial barcoding in the LRL approach is robust against barcode crosstalk.
Second-round barcode
There is the possibility that barcodes in the second round of barcoding could be shared between cells when cells are pooled together. To investigate this, we carried out a barcode competition assay in vitro using the RCS test analyte.
■ Using the assumption that 1 cell yields approximately 1.25 pg of poly adenylated RNA, and that each single-plex workflow requires 10 K cells, 12.5 ng of RCS test analyte was used as template for the in vitro barcode competition assay.
■ 4 reactions each composed of 12.5 ng RCS synthetic analyte was resuspended in a rehydration buffer of 3X SSC (Invitrogen, AM9770) supplemented with 1 mM DTT (Sigma, 43816), 200 ng/pL recombinant albumin (NEB, B9200S), and 0.2 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 16 pL final to mimic an input of 10 K rehydrated cells.
■ 4 reactions were then mixed with CDRA BC01 (2.27 pM SEQ ID NO: 1 top strand hybridised with 2.27 pM SEQ ID NO: 3 bottom strand), 0.625 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694), IX T4 ligation buffer (NEB, M0202L) and 25 U/pL T4 ligase (NEB, M0202L) final in 20 pL and incubated for 30 minutes at 37°C.
■ Following ligation, each reaction was supplemented with 0.25 U/pL Lambda exonuclease (NEB, M0262S) and 0.05 U/pL USER mix (NEB, M5505S) final in 22 pL and incubate for 15 minutes at 37°C, then 5 minutes on ice.
■ To simulate cell pooling, centrifugation and resuspension, each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
■ Reactions were then supplemented with 25 pM RT primer BC03 or BC04 (SEQ ID NO: 4or SEQ ID NO: 5 respectively), 25 pM strand switching primer (SEQ ID NO: 6), IX RT buffer (ThermoFisher Scientific, EP0753), 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694), 500 pM dNTPs (NEB, N0447S), 10 U/pL Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753) in 20 pL final and incubated for 90 minutes at 42°C, then 5 minutes on ice. Meanwhile, separate competing barcoded RT primer reactions were assembled as above without RNA template and incubated on ice.
■ Each reaction was then supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) to the pooled reaction for 0.1% (v/v) ECOSURF EH9 final. To simulate potential competition between RT primers during pooling, RNA template RT
reactions were either mixed with an equal volume of reaction without competing RT primer (no competition), or with an equal volume reaction with competing RT primer (competition). The pooled reactions were then incubated on ice for 5 minutes.
■ To simulate cell pooling, centrifugation and resuspension, each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 20 pL of IX NEB buffer r3.1 (NEB, B6003S).
■ Each reaction was then supplemented with ligation barcode BC05 (2.16 pM SEQ ID NO: 9 top strand hybridised with 2.4 pM SEQ ID NO: 7 bottom strand) IX T4 ligation buffer (NEB, M0202L), 20 U/pL T4 ligase (NEB, M0202L) in 50 pL final and incubated for 30 minutes at 37°C.
■ Each reaction was then supplemented with 3.24 pM blocking oligo (SEQ ID NO: 11) and 31.25 mM EDTA in 70 pL final and incubated on ice for 5 minutes.
■ To simulate cell lysis, each reaction was supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) then subjected to a 1.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 10 pL nuclease-free water.
■ Samples were amplified over 16x cycles using universal primers (SEQ ID NO: 12 and SEQ ID NO: 13) with Hotstart LongAmp Taq (NEB, M0533S). Amplicons were subjected to a 0.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 50 pL nuclease-free.
■ 200 fmols amplified cDNA prepared for sequencing using Oxford Nanopore Technologies SQK-LSK114 ligation sequencing kit, and sequence on a PromethlON with a FLO-PRO114M flowcell for each sample. Reads were aligned to a reference of RCS using the MinKNOW software.
Looking at the results in Figure 8 where there was no competing RT primer added there was only background misaligned detection of the competing barcode (<2%). However, where there was competing barcode added, there was still the same misaligned detection of the competing barcode (<2%). These data indicate there was no significant crosstalk between the RT barcodes, most likely since the reverse transcriptase has very minimal activity when incubated on ice and would be unable to extend any cross-primed templates. The PCR yield was consistent across the samples, indicating there was no significant difference in cDNA synthesis with or without RT competition.
Together these results show that the design and application of the RT barcode, specifically complementary to the top strand of the CDRA adapter, applied in the second round of combinatorial barcoding in the LRL approach is robust against barcode crosstalk.
Third-round barcode
There is the possibility that barcodes in the third round of barcoding could be shared between cells when cells are pooled together. The application of a blocking oligo can mitigate this crosstalk. To investigate this, we carried out a barcode competition assay in vitro using the RCS test analyte, using either no blocking oligo, a blocking oligo utilised in the RLL approach, or an alternative blocking oligo.
■ Using the assumption that 1 cell yields approximately 1.25 pg of poly adenylated RNA, and that each single-plex workflow requires 10 K cells, 12.5 ng of RCS test analyte was used as template for the in vitro barcode competition assay.
■ 9 reactions each composed of 12.5 ng RCS synthetic analyte was resuspended in a rehydration buffer of 3X SSC (Invitrogen, AM9770) supplemented with 1 mM DTT (Sigma, 43816), 200 ng/pL recombinant albumin (NEB, B9200S), and 0.2 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 16 pL final to mimic an input of 10 K rehydrated cells.
■ 4 reactions were then mixed with CDRA BC01 (2.27 pM SEQ ID NO: 1 top strand hybridised with 2.27 pM SEQ ID NO: 3 bottom strand), 0.625 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694), IX T4 ligation buffer (NEB, M0202L) and 25 U/pL T4 ligase (NEB, M0202L) final in 20 pL and incubated for 30 minutes at 37°C.
■ Following ligation, each reaction was supplemented with 0.25 U/pL Lambda exonuclease (NEB, M0262S) and 0.05 U/pL USER mix (NEB, M5505S) final in 22 pL and incubate for 15 minutes at 37°C, then 5 minutes on ice.
■ To simulate cell pooling, centrifugation and resuspension, each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in IX NEB buffer r3.1 (NEB, B6003S) supplemented with 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694) in 8 pL final.
■ Reactions were then supplemented with 25 pM RT primer BC03 (SEQ ID NO: 4), 25 pM strand switching primer (SEQ ID NO: 6), IX RT buffer (ThermoFisher Scientific, EP0753), 0.5 U/pL SUPERnaseln RNase Inhibitor (Invitrogen, AM2694), 500 pM dNTPs (NEB, N0447S), 10 U/pL Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753) in 20 pL final and incubated for 90 minutes at 42°C, then 5 minutes on ice. Meanwhile, separate competing barcoded RT primer reactions were assembled as above without RNA template and incubated on ice.
■ Each reaction was then supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) to the pooled reaction for 0.1% (v/v) ECOSURF EH9 final and incubated on ice for 5 minutes.
■ To simulate cell pooling, centrifugation and resuspension, each reaction was subjected to a 0.8X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 20 pL of IX NEB buffer r3.1 (NEB, B6003S).
■ Each reaction was then supplemented with either no third-round ligation barcode, BC05 (2.16 pM SEQ ID NO: 9 top strand hybridised with 2.4 pM SEQ ID NO: 7 bottom strand), or BC06 (2.16 pM SEQ ID NO: 9 top strand hybridised with 2.4 pM SEQ ID NO: 8 bottom strand), IX T4 ligation buffer (NEB, M0202L), 20 U/pL T4 ligase (NEB, M0202L) in 50 pL final and incubated for 30 minutes at 37°C.
■ Each reaction was then supplemented with either no blocking oligo, 3.24 pM standard blocking oligo (SEQ ID NO: 10), or 3.24 pM alternative blocking oligo (SEQ ID NO: 11), and 31.25 mM EDTA in 70 pL final and incubated on ice for 5 minutes.
■ To simulate potential competition between 3rd round adapters during pooling, cDNA template RT reactions were either mixed with an equal volume of reaction without competing 3rd round barcode (no competition), or with an equal volume reaction with competing 3rd round barcode (competition). The pooled reactions were then incubated on ice for 5 minutes.
■ To simulate cell lysis, each reaction was supplemented with 0.01 volumes 10% (v/v) ECOSURF EH9 (Sigma, STS0006) then subjected to a 1.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 10 pL nuclease-free water.
■ Samples were amplified over 16x cycles using universal primers (SEQ ID NO: 12 and SEQ ID NO: 13) with Hotstart LongAmp Taq (NEB, M0533S). Amplicons were subjected to a 0.5X SPRI clean with AMPure XP Reagent (Beckman Coulter, A63882), 70% ethanol washes and resuspended in 50 pL nuclease-free.
■ 200 fmols amplified cDNA prepared for sequencing using Oxford Nanopore Technologies SQK-LSK114 ligation sequencing kit, and sequence on a PromethlON with a FLO-PRO114M flowcell for each sample. Reads were aligned to a reference of RCS using the MinKNOW software.
Looking at the results in Figure 9 where there was no competing third-round ligation barcode there was only background misaligned detection of the competing barcode (<0.2%). However, if no blocking oligo was utilised where there is competing barcode added, there was a significant increase in the detection of the competing barcode (10.7%). Likewise, if there was no third-round ligation barcode present in the ligation before competition on pooling, there was still significant incorporation of the competing barcode (51.1%) and a correspondingly high amplicon yield from the PCR. This indicates that there is the potential for barcode crosstalk during the pooling of the third-round ligation reactions, even in the presence of 30 mM EDTA.
Looking at Figure 10 if the standard blocking oligo of the RLL approach was utilised, there was still some degree of barcode crosstalk with the competing barcode (3.4%). If there was no third-round ligation barcode present in the ligation before competition on pooling, only was there still incorporation of the competing barcode (27.0%), but with a PCR yield 4.4-
fold lower than those where a third-round ligation barcode was present in the ligation prior to pooling.
In Figure 11 the alternative blocking oligo of the LRL approach was utilised, there was still some degree of barcode crosstalk with the competing barcode (4.0%). If there was no third-round ligation barcode present in the ligation before competition on pooling, there was still incorporation of the competing barcode (43.3.0%), but with a PCR yield 13.3-fold lower than those where a third-round ligation barcode was present in the ligation prior to pooling.
Focusing on just the conditions where no third-round ligation barcode was present before competition on pooling, Figure 12 demonstrates most clearly the effectiveness of using a blocking oligo prior to sample pooling. Both the standard and alternative blocking oligo resulted in decreased incorporation of the competing barcode with diminishing PCR yields, but the alternative blocking oligo resulted in the biggest decrease of competing barcode amplicons.
Together these results show that the design and application of the third-round ligation barcode is capable of crosstalk between barcodes upon sample pooling. However, the utilisation of a blocking oligo complementary to the top strand of the third-round ligation barde barcode, specifically the alternative blocking oligo utilised in the LRL approach, can help to decrease the potential crosstalk.
Claims
1. A method of uniquely labelling RIMA molecules in a population of cells, the method comprising:
(a) dividing the population of cells into a plurality of first samples;
(b) ligating first adaptors to the 3' ends of the RNA molecules in the plurality of first samples, wherein the first adaptors comprise primer sites for reverse transcription (RT) and wherein the first adaptors used in each first sample comprise the reverse complement of a different first barcode;
(c) pooling the plurality of first samples and dividing the pool into a plurality of second samples;
(d) reverse transcribing the RNA molecules in the plurality of second samples using the primer sites and RT primers to form double stranded constructs, wherein the RT primers used in each second sample comprise a different second barcode;
(e) pooling the plurality of second samples and dividing the pool into a plurality of third samples; and
(f) ligating second adaptors to the double stranded constructs in the plurality of third samples to form triple barcoded constructs, wherein the second adaptors used in each third sample comprise a different third barcode.
2. A method according to claim 1, wherein the first adaptors are double stranded and comprise overhangs which are capable of hybridising to the 3' ends of the RNA molecules.
3. A method according to claim 2, wherein the overhangs comprise one or more of (i) thymine-containing nucleotides, (ii) uracil-containing nucleotides and (iii) universal nucleotides.
4. A method according to claim 2 or 3, wherein the reverse complements of the different first barcodes are present in the strands of the first adaptors which are ligated to the RNA molecules.
5. A method according to any one of claims 2-4, wherein the nucleotides at the 3' ends of one of both strands of the double stranded first adaptors are dideoxy cytosine (ddC) or inverted dT.
6. A method according to any one of claims 2-5, wherein the non-ligated strands of the first adaptors are removed before step (d).
7. A method according to any one of the preceding claims, wherein step (d) comprises hybridising the RT primers to the first adaptors.
8. A method according to claim 7, wherein the RT in step (d) produces strands comprising the different first barcodes, the different second barcodes and sequences complementary to the RNA molecules.
9. A method according to claim 8, wherein the components are in the following order in the strands 5' to 3': different second barcodes, different first barcodes and complementary sequences.
10. A method according to claim 8 or 9, wherein the second adaptors are ligated to the strands comprising the different first barcodes, different second barcodes and complementary sequences.
11. A method according to claim 10, wherein the second adaptors comprise overhangs capable of hybridising to the RT primers.
12. A method according to claim 11, wherein the different third barcodes are present in the non-overhanging strands of the second adaptors which are ligated to the double stranded constructs.
13. A method according to claim 11 or 12, wherein step (f) is carried out in the presence of 5' phosphorylated blockers which are complementary to the overhangs.
14. A method according to any one of the preceding claims, wherein the barcode used in each sample is different from any other barcode or all the other barcodes used in the method.
15. A method according to any one of the preceding claims, wherein the number of cells in the population is less than the number of first samples multiplied by the number of second samples multiplied the number of third samples.
16. A method according to any one of the preceding claims, wherein (i) the plurality of first samples, (ii) the plurality of second samples, and/or (ii) the plurality of third samples comprises at least about 24 samples.
17. A method according to any one of the preceding claims, wherein the population of cells is fixed and/or permeabilised before step (a).
18. A method according to any one of the preceding claims, wherein the method further comprises repeating steps (e) to (f) at least once comprising the formation of a plurality of fourth or more samples and using third or more adaptors which comprise a different fourth
or more barcode for each fourth or more sample to produce quadruple or more barcoded constructs.
19. A method according to any one of the preceding claims, wherein the method further comprises isolating the barcoded constructs from the cells and/or amplifying the barcoded constructs.
20. A method according to any of the preceding claims, wherein the method further comprises characterising or sequencing the barcoded constructs, preferably using a nanopore.
21. A method of characterising the transcriptome of an individual cell or two or more cells in a population of cells, the method comprising:
(a) conducting a method according to any one of the preceding claims on the population of cells; and
(b) characterising the RIMA molecules from the individual cell or two or more cells using their unique labelling.
22. A method according to claim 21, wherein the characterisation comprises sequencing the RNA molecules, preferably using a nanopore.
23. A kit for uniquely labelling RNA molecules in a population of cells, the kit comprising two or more first adaptors, wherein the two or more first adaptors comprise primer sites for RT, wherein each of the two or more first adaptors comprises the reverse complement of a different first barcode.
24. A kit according to claim 23, wherein the two or more first adaptors are double stranded and comprise overhangs which are capable of hybridising to the 3' ends of the RNA molecules.
25. A kit according to claim 24, wherein the overhangs comprise one or more of (i) thymine-containing nucleotides, (ii) uracil-containing nucleotides and (iii) universal nucleotides.
26. A kit according to claim 24 or 25, wherein the reverse complements of the different first barcodes are present in the strands of the two or more first adaptors which are ligated to the RNA molecules.
27. A kit according to any one of claims 24-26, wherein the nucleotides at the 3' ends of one of both strands of the double stranded two or more first adaptors are dideoxycytosine (ddC) or inverted dT.
28. A kit according to any one of claims 23-27, wherein the kit further comprises (a) two or more RT primers capable of hybridising to the two or more first adaptors and each comprising a different second barcode and/or (b) two or more second adaptors each comprising a different third barcode.
29. A kit according to any one of claims 23-28, wherein (a) each barcode in the kit is different from any other barcode or all the other barcodes in the kit and/or (b) the kit comprises at least 96 first adaptors, at least 96 RT primers and/or at least 96 second adaptors.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GBGB2304324.3A GB202304324D0 (en) | 2023-03-24 | 2023-03-24 | Method and kits |
| PCT/EP2024/057803 WO2024200280A1 (en) | 2023-03-24 | 2024-03-22 | Method and kits |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4689159A1 true EP4689159A1 (en) | 2026-02-11 |
Family
ID=86228075
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24718330.4A Pending EP4689159A1 (en) | 2023-03-24 | 2024-03-22 | Method and kits |
Country Status (5)
| Country | Link |
|---|---|
| EP (1) | EP4689159A1 (en) |
| CN (1) | CN120898002A (en) |
| AU (1) | AU2024245816A1 (en) |
| GB (1) | GB202304324D0 (en) |
| WO (1) | WO2024200280A1 (en) |
Family Cites Families (41)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5198543A (en) | 1989-03-24 | 1993-03-30 | Consejo Superior Investigaciones Cientificas | PHI29 DNA polymerase |
| US6267872B1 (en) | 1998-11-06 | 2001-07-31 | The Regents Of The University Of California | Miniature support for thin films containing single channels or nanopores and methods for using same |
| GB0505971D0 (en) | 2005-03-23 | 2005-04-27 | Isis Innovation | Delivery of molecules to a lipid bilayer |
| US20110121840A1 (en) | 2007-02-20 | 2011-05-26 | Gurdial Singh Sanghera | Lipid Bilayer Sensor System |
| WO2009020682A2 (en) | 2007-05-08 | 2009-02-12 | The Trustees Of Boston University | Chemical functionalization of solid-state nanopores and nanopore arrays and applications thereof |
| WO2009035647A1 (en) | 2007-09-12 | 2009-03-19 | President And Fellows Of Harvard College | High-resolution molecular graphene sensor comprising an aperture in the graphene layer |
| GB0724736D0 (en) | 2007-12-19 | 2008-01-30 | Oxford Nanolabs Ltd | Formation of layers of amphiphilic molecules |
| KR20110125226A (en) | 2009-01-30 | 2011-11-18 | 옥스포드 나노포어 테크놀로지즈 리미티드 | Hybridization linker |
| GB0901588D0 (en) | 2009-02-02 | 2009-03-11 | Itis Holdings Plc | Apparatus and methods for providing journey information |
| CN102405410B (en) | 2009-04-20 | 2014-06-25 | 牛津楠路珀尔科技有限公司 | Lipid bilayer sensor array |
| US8828211B2 (en) | 2010-06-08 | 2014-09-09 | President And Fellows Of Harvard College | Nanopore device with graphene supported artificial lipid membrane |
| KR101939420B1 (en) | 2011-02-11 | 2019-01-16 | 옥스포드 나노포어 테크놀로지즈 리미티드 | Mutant pores |
| EP4737389A2 (en) | 2011-05-27 | 2026-05-06 | Oxford Nanopore Technologies PLC | Coupling method |
| CN104039979B (en) | 2011-10-21 | 2016-08-24 | 牛津纳米孔技术公司 | Hole and Hel308 unwindase is used to characterize the enzyme method of herbicide-tolerant polynucleotide |
| GB201120910D0 (en) | 2011-12-06 | 2012-01-18 | Cambridge Entpr Ltd | Nanopore functionality control |
| US9617591B2 (en) | 2011-12-29 | 2017-04-11 | Oxford Nanopore Technologies Ltd. | Method for characterising a polynucleotide by using a XPD helicase |
| AU2012360244B2 (en) | 2011-12-29 | 2018-08-23 | Oxford Nanopore Technologies Limited | Enzyme method |
| KR102083695B1 (en) | 2012-04-10 | 2020-03-02 | 옥스포드 나노포어 테크놀로지즈 리미티드 | Mutant lysenin pores |
| CA2879261C (en) | 2012-07-19 | 2022-12-06 | Oxford Nanopore Technologies Limited | Modified helicases |
| GB201313121D0 (en) | 2013-07-23 | 2013-09-04 | Oxford Nanopore Tech Ltd | Array of volumes of polar medium |
| JP6375301B2 (en) | 2012-10-26 | 2018-08-15 | オックスフォード ナノポール テクノロジーズ リミテッド | Droplet interface |
| CN117947149A (en) | 2013-10-18 | 2024-04-30 | 牛津纳米孔科技公开有限公司 | Modified enzymes |
| CN111534504B (en) | 2014-01-22 | 2024-06-21 | 牛津纳米孔科技公开有限公司 | Methods for attaching one or more polynucleotide binding proteins to a target polynucleotide |
| WO2015150786A1 (en) | 2014-04-04 | 2015-10-08 | Oxford Nanopore Technologies Limited | Method for characterising a double stranded nucleic acid using a nano-pore and anchor molecules at both ends of said nucleic acid |
| GB201417712D0 (en) | 2014-10-07 | 2014-11-19 | Oxford Nanopore Tech Ltd | Method |
| CA2959220A1 (en) | 2014-09-01 | 2016-03-10 | Vib Vzw | Mutant csgg pores |
| EP4397972A3 (en) | 2014-10-16 | 2024-10-09 | Oxford Nanopore Technologies PLC | Alignment mapping estimation |
| CA3021580A1 (en) | 2015-06-25 | 2016-12-29 | Barry L. Merriman | Biomolecular sensors and methods |
| CN116200476A (en) | 2016-03-02 | 2023-06-02 | 牛津纳米孔科技公开有限公司 | Target analyte determination methods, mutant CsgG monomers, constructs, polynucleotides and oligo-wells thereof |
| US11174503B2 (en) * | 2016-09-21 | 2021-11-16 | Predicine, Inc. | Systems and methods for combined detection of genetic alterations |
| GB201616590D0 (en) | 2016-09-29 | 2016-11-16 | Oxford Nanopore Technologies Limited | Method |
| GB201620450D0 (en) | 2016-12-01 | 2017-01-18 | Oxford Nanopore Tech Ltd | Method |
| CN117106038B (en) | 2017-06-30 | 2025-12-09 | 弗拉芒区生物技术研究所 | Novel protein pores |
| ES2994431T3 (en) | 2017-09-22 | 2025-01-23 | Univ Washington | In situ combinatorial labeling of cellular molecules |
| SG11202008080RA (en) * | 2018-02-22 | 2020-09-29 | 10X Genomics Inc | Ligation mediated analysis of nucleic acids |
| US11519027B2 (en) * | 2018-03-30 | 2022-12-06 | Massachusetts Institute Of Technology | Single-cell RNA sequencing using click-chemistry |
| JP7358388B2 (en) * | 2018-05-03 | 2023-10-10 | ベクトン・ディキンソン・アンド・カンパニー | Molecular barcoding at opposite transcript ends |
| GB201907244D0 (en) | 2019-05-22 | 2019-07-03 | Oxford Nanopore Tech Ltd | Method |
| JP2023530155A (en) | 2020-06-18 | 2023-07-13 | オックスフォード ナノポール テクノロジーズ ピーエルシー | Methods for Characterizing Polynucleotides Translocating Through Nanopores |
| LT4211260T (en) * | 2020-09-11 | 2025-06-25 | New England Biolabs, Inc. | Application of immobilized enzymes for nanopore library construction |
| US12378596B2 (en) * | 2020-12-03 | 2025-08-05 | Roche Sequencing Solutions, Inc. | Whole transcriptome analysis in single cells |
-
2023
- 2023-03-24 GB GBGB2304324.3A patent/GB202304324D0/en not_active Ceased
-
2024
- 2024-03-22 WO PCT/EP2024/057803 patent/WO2024200280A1/en not_active Ceased
- 2024-03-22 AU AU2024245816A patent/AU2024245816A1/en active Pending
- 2024-03-22 CN CN202480021003.5A patent/CN120898002A/en active Pending
- 2024-03-22 EP EP24718330.4A patent/EP4689159A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| CN120898002A (en) | 2025-11-04 |
| AU2024245816A1 (en) | 2025-09-04 |
| GB202304324D0 (en) | 2023-05-10 |
| WO2024200280A1 (en) | 2024-10-03 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11390904B2 (en) | Nanopore-based method and double stranded nucleic acid construct therefor | |
| EP2895618B1 (en) | Sample preparation method | |
| JP6776366B2 (en) | Mutant pore | |
| EP3033435B1 (en) | Method for fragmenting nucleic acid by means of transposase | |
| JP6408494B2 (en) | Enzyme stop method | |
| EP3259281B1 (en) | Hetero-pores | |
| CN114262735B (en) | Adaptors for characterising polynucleotides and uses thereof | |
| US20210381041A1 (en) | Enzymatic Enrichment of DNA-Pore-Polymerase Complexes | |
| US20250361552A1 (en) | Method and adaptors | |
| EP4689159A1 (en) | Method and kits | |
| WO2024131530A1 (en) | Preparation method for sequencing library | |
| WO2026057755A1 (en) | Adaptors and kits for rna molecules labelling and characterising | |
| WO2025242713A1 (en) | A method for sequencing chromatin |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251020 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |