WO2025259693A1 - High-throughput unbiased identification of double-stranded dna breaks - Google Patents

High-throughput unbiased identification of double-stranded dna breaks

Info

Publication number
WO2025259693A1
WO2025259693A1 PCT/US2025/033041 US2025033041W WO2025259693A1 WO 2025259693 A1 WO2025259693 A1 WO 2025259693A1 US 2025033041 W US2025033041 W US 2025033041W WO 2025259693 A1 WO2025259693 A1 WO 2025259693A1
Authority
WO
WIPO (PCT)
Prior art keywords
grna
cells
dso
library
genomic loci
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/US2025/033041
Other languages
French (fr)
Inventor
Shengdar TSAI
Azusa MATSUBARA
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
St Jude Childrens Research Hospital
Original Assignee
St Jude Childrens Research Hospital
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by St Jude Childrens Research Hospital filed Critical St Jude Childrens Research Hospital
Publication of WO2025259693A1 publication Critical patent/WO2025259693A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/1034Isolating an individual clone by screening libraries
    • C12N15/1065Preparation or screening of tagged libraries, e.g. tagged microorganisms by STM-mutagenesis, tagged polynucleotides, gene tags
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/1034Isolating an individual clone by screening libraries
    • C12N15/1082Preparation or screening gene libraries by chromosomal integration of polynucleotide sequences, HR-, site-specific-recombination, transposons, viral vectors
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/11DNA or RNA fragments; Modified forms thereof; Non-coding nucleic acids having a biological activity
    • C12N15/113Non-coding nucleic acids modulating the expression of genes, e.g. antisense oligonucleotides; Antisense DNA or RNA; Triplex- forming oligonucleotides; Catalytic nucleic acids, e.g. ribozymes; Nucleic acids used in co-suppression or gene silencing
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N9/00Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
    • C12N9/14Hydrolases (3)
    • C12N9/16Hydrolases (3) acting on ester bonds (3.1)
    • C12N9/22Ribonucleases [RNase]; Deoxyribonucleases [DNase]
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2310/00Structure or type of the nucleic acid
    • C12N2310/10Type of nucleic acid
    • C12N2310/20Type of nucleic acid involving clustered regularly interspaced short palindromic repeats [CRISPR]
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2320/00Applications; Uses
    • C12N2320/10Applications; Uses in screening processes
    • C12N2320/11Applications; Uses in screening processes for the determination of target sites, i.e. of active nucleic acids
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2740/00Reverse transcribing RNA viruses
    • C12N2740/00011Details
    • C12N2740/10011Retroviridae
    • C12N2740/16011Human Immunodeficiency Virus, HIV
    • C12N2740/16041Use of virus, viral particle or viral elements as a vector
    • C12N2740/16043Use of virus, viral particle or viral elements as a vector viral genome or elements thereof as genetic vector
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N2800/00Nucleic acids vectors
    • C12N2800/90Vectors containing a transposable element

Definitions

  • RNA-guided gene editing is critical components in RNA-guided gene editing, particularly within the CRISPR-Cas system, as they directly influence the precision and efficacy of genetic modifications.
  • High specificity ensures that the guide RNA (gRNA) targets only the intended DNA sequence, minimizing off-target effects that could cause unintended genetic alterations and potential adverse effects. This is crucial for applications in therapeutic gene editing, where precision is paramount to avoid unwanted mutations that could lead to disease or dysfunction.
  • Activity refers to the efficiency with which the nuclease-gRNA complex induces the desired genetic changes. High activity levels ensure that the target gene is effectively and reliably edited, making the process more efficient and reducing the time and resources needed for successful gene modification. Together, specificity and activity are fundamental to achieving accurate, safe, and efficient gene editing, enabling advancements in genetic research, biotechnology, and personalized medicine.
  • this disclosure provides a method for identifying one or more highly active and/or highly specific guide RNA (gRNA).
  • the method comprises: (a) transfecting cells comprising an RNA-guided nuclease (RGN) and a double- stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO; (b) detecting genomic loci of the cells that comprise the integrated DSO; and (c) identifying one or more highly active and/or highly specific gRNA from the gRNA library using the genomic loci detected in (b).
  • RGN RNA-guided nuclease
  • DSO double- stranded oligodeoxynucleotide
  • this disclosure provides a method for identifying one or more highly active guide RNA (gRNA).
  • the method comprises: (a) transfecting cells comprising an RNA-guided nuclease (RGN) and a double- stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO; (b) detecting genomic loci of the cells that comprise the integrated DSO; and (c) identifying one or more highly active from the gRNA library using the genomic loci detected in (b).
  • RGN RNA-guided nuclease
  • DSO double- stranded oligodeoxynucleotide
  • this disclosure provides a method for identifying one or more highly specific guide RNA (gRNA).
  • the method comprises: (a) transfecting cells comprising an RNA-guided nuclease (RGN) and a double- stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO; (b) detecting genomic loci of the cells that comprise the integrated DSO; and (c) identifying one or more highly specific from the gRNA library using the genomic loci detected in (b).
  • RGN RNA-guided nuclease
  • DSO double- stranded oligodeoxynucleotide
  • an edit distance between any two different gRNAs is greater than a fixed edit distance.
  • the fixed edit distance is 5, 6, or 7 (e.g., a hamming distance between any two different gRNAs that is 5, 6, or 7 nucleotides).
  • genomic target regions in the cells are separated from each other by at least 30 nucleotide base pairs.
  • the gRNA library comprises at least 100, at least 1000, or at least 5000 different gRNAs, optionally all targeting different sequences within a gene.
  • detecting of the genomic loci is performed without intentionally shearing DNA of the cells.
  • detecting of the genomic loci comprises extracting DNA from the cells and processing the DNA using DNA tagmentation.
  • detecting of the genomic loci of the cells comprises amplifying and/or sequencing DNA of the cells, optionally using next generation sequence (NGS) technology.
  • amplifying the DNA of the cells comprising use of a pair of primers respectively complementary to a region of the DSO that is at least three nucleotides from the 3’ and 5’ terminus of the DSO.
  • amplifying the DNA of the cells comprises performing a polymerase chain reaction (PCR), optionally no more than a single PCR.
  • the gRNA library comprises a lentiviral library encoding the different gRNAs.
  • the different gRNAs are encoded on minus strands of lentiviruses of the lentiviral library.
  • the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1.
  • MOI multiplicity of infection
  • the RGN is a Cas nuclease, optionally a Cas9 nuclease.
  • the Cas nuclease comprises multiple nuclear localization sequences, optionally three nuclear localization sequences.
  • transfecting the cells with the DSOs comprises transfecting the cells with about 50 pmol to about 500 pmol of DSO per 0.5 million cells to about 1 million cells, optionally about 100-200 pmol of DSO per 0.6 million cells.
  • each of at least a subset of the different gRNAs of the gRNA library is operably linked to a U6 promoter and/or human cysteine tRNA (hCtRNA).
  • hCtRNA human cysteine tRNA
  • the method further comprises associating one or more of the different gRNAs of the gRNA library with one or more genomic loci comprising a DSO integration.
  • associating one or more of the different gRNAs of the gRNA library comprises, for each different gRNA, associating the different gRNA with one or more genomic loci comprising an integrated DSO if the homology region of the different gRNA has the highest nucleotide percent identity to the one or more genomic loci compared to the homology regions of the other different gRNAs of the gRNA library.
  • determining highest nucleotide percent identity comprises determining highest nucleotide percent identity to a region of the one or more genomic loci that is within 25 base pairs of the integrated DSO.
  • identifying a highly active and/or highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) an amount (e.g., a number) of on-target genomic locus comprising the DSO integration in the cells, and (ii) an amount (e.g., a number) of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value
  • identifying a highly active gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) an amount (e.g., a number) of on-target genomic locus comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the amount of on-target genomic locus in the cells is higher than a first threshold value.
  • an amount e.g., a number
  • identifying a highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) an amount (e.g., a number) of on-target genomic locus comprising the DSO integration in the cells, and (ii) an amount (e.g., a number) of off- target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
  • identifying a highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining an amount (e.g., a number) of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as highly specific gRNA the amount of off-target genomic loci in the cells is higher than a second threshold value.
  • determining the number of on-target genomic loci comprises determining a relative number using (i) a number of sequencing reads covering the on-target genomic loci comprising the DSO integration and (ii) a frequency of the gRNA in the gRNA library. In some embodiments, determining the relative amount of on target and off-target genomic loci comprises determining using (i) a number of sequencing reads covering the on- target genomic loci comprising the DSO integration, and (ii) a number of sequencing reads covering the off-target genomic loci comprising the DSO integration. In some embodiments, the first threshold value is the 90th percentile of relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the second threshold value is the 90th percentile of relative amounts of on target and off-target genomic loci for each guide RNA in the gRNA library.
  • associating one or more of the different gRNAs of the gRNA library comprises associating using single cell sequencing, wherein the genomic locus comprising a DSO integration with the highest percent identity to the gRNA in a given cell is an on-target genomic locus, and wherein other genomic loci comprising a DSO integration in the given cell are off-target genomic loci.
  • identifying a highly active and/or highly specific gRNA from the gRNA library comprises: determining (i) the number of cells comprising an integrated DSO in the on-target genomic locus and (ii) the number of cells comprising an integrated DSO in the off-target genomic loci; and identifying the gRNA as a highly active gRNA if the number of cells comprising the integrated DSO in the on-target genomic locus is higher than a first threshold value, and/or identifying the gRNA as a highly specific gRNA if the number of off- target genomic loci in the cells is lower than a second threshold value.
  • this disclosure provides a method for identifying a guide RNA (gRNA) for editing a gene of interest, the method comprising: (a) transfecting cells comprising an RNA- guided nuclease (RGN) and a double-stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs targeting different genomic regions of a gene of interest to produce cells comprising genomic DSBs tagged by an integrated DSO in the gene of interest; (b) detecting genomic loci in the gene of interest that comprise the integrated DSO; and (c) identifying from the gRNA library a gRNA that is highly active and/or highly specific for the gene of interest using the genomic loci detected in (b).
  • RGN RNA- guided nuclease
  • DSO double-stranded oligodeoxynucleotide
  • this disclosure provides a method for identifying a Cas enzyme variant gene editing system, the method comprising: (a) transfecting cells with a candidate Cas enzyme variant, a double- stranded oligodeoxynucleotide (DSO), and a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO; (b) detecting genomic loci of the cells that comprise the integrated DSO; and (c) identifying from the gRNA library one or more gRNAs as being highly active and/or highly specific for a genomic loci detected in (b), thereby identifying a Cas enzyme variant gene editing system comprising a Cas enzyme variant and one or more gRNAs.
  • DSO double- stranded oligodeoxynucleotide
  • this disclosure provides a method for identifying a Cas enzyme ortholog gene editing system, the method comprising: (a) transfecting first cells with a first Cas enzyme ortholog, a double-stranded oligodeoxynucleotide (DSO), and a first gRNA library comprising different gRNAs to produce first cells comprising genomic DSBs tagged by an integrated DSO; (b) transfecting second cells with a second Cas enzyme ortholog, a double- stranded oligodeoxynucleotide (DSO), and a second gRNA library comprising different gRNAs to produce second cells comprising genomic DSBs tagged by an integrated DSO; (c) combining, after (a) and (b), the first cells and the second cells; (d) detecting genomic loci of the first cells that comprise the integrated DSO and detecting genomic loci of the second cells that comprise the integrated DSO; and (e) identifying, using one or more genomic loci detected in (d), the
  • this disclosure provides a method for identifying one or more highly active and/or highly specific guide RNA (gRNA), comprising: (a) transfecting cells comprising a Cas nuclease and a double-stranded oligodeoxynucleotide (DSO) with a lentiviral library encoding at least 100 different gRNAs to produce cells comprising genomic DSBs tagged by integrated DSOs, wherein an edit distance between any two different gRNAs is greater than 5; (b) extracting DNA from the cells and processing the DNA using DNA tagmentation to produce a DNA library; (c) amplifying DNA of the DNA library using a polymerase chain reaction (PCR), wherein the PCR comprises a pair of primers respectively complementary to a region of the DSO that is at least three nucleotides from the 3’ and 5’ terminus of the DSO; (d) sequencing amplified DNA of (c) to detect genomic loci of the cells that comprise the integrated DSOs; and (e) as
  • identifying a highly active and/or highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
  • genomic target regions in the cells are separated from each other by at least 30 nucleotide base pairs.
  • sequencing of the amplified DNA is performed using next generation sequencing. In some embodiments, no more than a single PCR is performed.
  • the Cas nuclease is a Cas9 nuclease, optionally comprising multiple nuclear localization sequences, optionally three nuclear localization sequences.
  • transfecting the cells with the DSO comprises transfecting the cells with about 50 pmol to about 500 pmol of DSO per 0.5 million cells to about 1 million cells, optionally about 100-200 pmol of DSO per 0.6 million cells.
  • each of at least a subset of the different gRNAs of the library is operably linked to a U6 promoter and/or human cysteine tRNA (hCtRNA).
  • the cells comprise mammalian cells. In some embodiments, the mammalian cells comprise human cells.
  • FIGs. 1A-1D show development of lentivirus sgRNA delivery to activated human T- cells.
  • FIG. 1A shows a comparison of transduction enhancers and timing of transduction to express CTLA4 site09 guide with lentivirus.
  • Purified 2xNLS Cas9 protein was subsequently nucleofected.
  • LENTIBOOST® enhancer after 24 hours of T-cell activation provided the best editing efficiency. From left to right in each condition: 24 hours, 48 hours, 72 hours.
  • FIG. IB shows the effects of nucleofecting different amounts of Cas9 concentration with 100 pmol DSO. Regardless of puromycin condition, 75 pM Cas9 provided the best DSO integration efficiency.
  • FIG. 1C shows a comparison of DSO concentration on indel frequency and DSO integration efficiency. Increased DSO concentration correlated with increased DSO integration; however, decreased indel frequency was observed with Cas9 independent integration at 200 pmol. 100 pmol DSO provided maximum editing efficiency and DSO integration. From left to right in each condition: control, Puromycin selection BEFORE nucleofection, Puromycin selection BEFORE and AFTER nucleofection.
  • FIGs. 2A-2D show a Pooled-GUIDE-seq-2 design and analysis pipeline.
  • FIG. 2A shows Pooled-GUIDE-seq-2 experiment workflow.
  • FIG. 2B shows criteria of the sgRNA pooled sequence design.
  • FIG. 2C shows distribution of guide library target sequence location. 5000 guides were selected across the whole genome from both DNA strands.
  • FIG. 2D shows a schematic of the Pooled-GUIDE-seq-2 analysis pipeline. Sequences adjacent to the DSO tag were mapped to the genome, and DSB sequences were extracted based on the window size. The window sequence was then subjected to sequence similarity analysis against the sgRNA protospacer list created in FIG. 2A, providing an output of on/off-target location and mismatch information.
  • FIGs. 3A-3C show utilization of U6 promoter or hCtRNA promoter to express single guide RNAs (sgRNAs).
  • FIG. 3A shows lentiviral vectors used in preparing sgRNA library. Left: U6 promoter, right: hCtRNA driving the sgRNA transcription. Cellular nucleases RNaseP and RNaseZ released sgRNA post transcription.
  • FIG. 3B shows a comparison of editing efficiency between U6 promoter and hCtRNA to express 6 different sgRNAs.
  • FIG. 3C shows that a library of 4435 sgRNA without a “G” constraint was generated. Coverage of each sgRNA sequence was analyzed by amplicon sequencing from the lentiviral plasmid library (x-axis) and the genome post transduction (y-axis).
  • FIGs. 4A-4G show Pooled-GUIDE-seq-2 with 500 sgRNA pool.
  • FIG. 4A shows distribution of the number of GUIDE- seq-2 sites identified for each guide from a pooled library of 500.
  • FIG. 4B shows an example of the output information obtained for each guide.
  • FIG. 4C shows reproducibility of two replicate Pooled-GUIDE-seq-2 experiments. Number of the read counts for each identified site is plotted.
  • FIG. 4D shows distribution of on-target activity of the 20 guides sampled from the 500 guide library compared with single guide GUIDE-seq-2.
  • FIG. 4E shows an example of one of the guides analyzed in Pooled-GUIDE-seq-2 (left) and GUIDE- seq-2 (right).
  • FIG. 4F shows a comparison of the ranking of guides from Pooled-GUIDE-seq-2 (x-axis) or single GUIDE-seq-2 (y-axis) based on its on-target activity or specificity.
  • FIG. 4G shows the number of guides detected by Pooled-GUIDE-seq-2 with different mismatch filtering.
  • FIGs. 5A-5D show Pooled-GUIDE-seq-2 with different filtering threshold. Results of two replicates are shown.
  • FIG. 5A shows total read counts. Each point represents a guide. The x-axis shows replicase 1 and the y-axis shows replicase 2.
  • FIG. 5B shows number of sites identified. Each point represents a guide.
  • FIG. 5C shows histograms of guide specificity from round2.
  • FIG. 5D shows number of read counts. Each point represents the identified on- or off- target site and is greyscaled by the number of mismatches.
  • FIGs. 6A-6B show sequence feature analysis. FIG.
  • FIG. 6A shows all GUIDE-seq-2 sites from Pooled-GUIDE-seq-2 from the dataset with 500 library roundl (top panels) and 4501 library (bottom panels) that were analyzed for the number of mismatches, PAM preference, and sequence tolerability.
  • FIG. 6B shows sequence features of the guide sequences from all (left), top 10 high activity (e.g., highly active) and specificity (e.g., highly specific) (middle) and bottom 10 least activity and specificity (right) from the dataset with 500 library roundl (top panels) and 4501 library (bottom panels).
  • FIG. 7 shows different applications of Pooled-GUIDE-seq-2.
  • Pooled-GUIDE-seq-2 can be used for target based (top) and unbiased (bottom) applications.
  • target based applications a focused library on specific disease targets can be designed, and Pooled-GUIDE-seq-2 can be performed using a clinically relevant cell type to allow rapid selection of candidate guides with high activity and high specificity in the relevant genomic context.
  • a genome wide sgRNA library can be used to compare results from different cell types, different Cas9 orthologues and other CRISPR editors.
  • FIG. 8 shows a Manhattan plot depicting each guide at the on-target chromosome (X- axis, separated by color) and the specificity (Y-axis) from a pool of 500 guides.
  • the size of the markers depicts the activity of each guide. Large markers with specificity near 1.0 are high specificity and high activity guide RNAs.
  • FIG. 9 shows a schematic depicting identification of the most highly active and specific targets within a specific disease-associated genes.
  • FIG. 10 shows a schematic depicting identification of the most highly active and specific Cas variant among a set of orthologues for a particular gene target.
  • FIG. 11 shows a schematic depicting PCR amplification of blunt double-stranded oligodeoxynucleotide (dsODN) (e.g., a DSO) using the old GUIDE-Seq method compared to the new GUIDE-seq-2 method described herein.
  • dsODN blunt double-stranded oligodeoxynucleotide
  • the new GUIDE-SEQ-2 PCR amplification method has fewer PCR steps and allows for detection of off-target amplification during sequencing.
  • FIGs. 12A-12E show an improvement to GUIDE-Seq-2 when additional nucleotides are added to the i7+ strand primer to match the melting temperate to the i7- strand primer.
  • FIG. 12A shows the original i7+ primer and the new i7+ primer with the additional nucleotides.
  • FIG. 12B shows that using the new i7+ primer improves amplification in GUIDE-Seq-2.
  • FIG. 12C shows that using the new i7+ primer improves the balance of the final GUIDE-Seq-2 mapped regions.
  • FIG. 12D shows that reproducibility between using the original i7+ primer and the new i7+ primer.
  • FIG. 12E shows a correlation of GUIDE-Seq-2 results using the original i7+ primer and the new i7+ primer.
  • FIGs. 13A-13C show results for several Pooled-GUIDE-seq-2 experiments.
  • FIG. 13A shows results from a Pooled-GUIDE-seq-2 experiment on 4,501 guides.
  • FIG. 13B shows results from a Pooled-GUIDE-seq-2 experiment on 4,335 guides using a tRNA promoter instead of a U6 promoter to express the guides, which allows use of guides that start with a “G” nucleotide.
  • FIG. 13C shows results from a Pooled-GUIDE-seq-2 experiment on 552 guides that start with a “G” using a U6 promoter to express the guides.
  • RNA-guided nuclease such as a CRISPR-Cas nuclease
  • CRISPR-Cas nuclease can bind to a gRNA to form a complex, which can then bind to a target genomic locus that includes a sequence complementary to the gRNA. That locus is then edited by the RNA-guided nuclease.
  • the gRNA of that complex is considered to be highly active if the frequency of on-target mutation produced by the complex is high.
  • the gRNA of that complex is considered to have high specificity if the frequency of off-target mutation produced by the complex is low.
  • Identifying a gRNA that is both highly active and highly specific for a target genomic locus can be challenging, and identifying multiple highly active and highly specific gRNAs for multiple genomic loci is even more challenging.
  • the gene-editing field needs a streamlined method for massively parallel profiling of cellular genome-wide activity of genome editors, such as CRISPR genome editors, which enables accurate and efficient identification of highly active and highly specific gRNAs.
  • the cells are harvested, and their genomic DNA is extracted.
  • Primers complementary to the DSO are used to amplify the regions of the genome where the DSO has integrated, and the amplified DNA fragments are sequenced using high- throughput sequencing methods. Finally, the sequencing data is analyzed to identify the genomic locations of the DSO, which correspond to the sites of DSBs. This allows for the identification of both on-target and off-target sites of genome editing.
  • GUIDE-seq is effective for assessing a single gRNA
  • GUIDE-seq throughput for assessing specificity and activity of many gRNAs is low and can take many months to perform. This is partly because GUIDE-seq lacks a means for normalizing activity among different gRNAs and as such cannot accurately measure activity and specificity of many different gRNAs simultaneously in the same experiment. Instead, GUIDE-seq requires an independent set of experiments for each gRNA.
  • gRNA For each gRNA, cells are transfected, a DSB is made, a DSO is integrated, cells are harvested, DNA is extracted, several rounds of nested PCR are used to amplify the DNA, high-throughput sequencing is performed, and then the sequencing data is analyzed. Repeating this set of experiments for each gRNA of a library of hundreds of gRNAs, for example, would take months or longer. Thus, without cost- and time-prohibitive experimentation, one cannot use the current GUIDE-seq protocol to identify from a pool of gRNAs which have the highest specificity and activity.
  • This disclosure provides, in some aspects, methods for broadly evaluating and designing gRNAs at scale. These methods can be used to simultaneously and accurately measure the activity and/or specificity of many gRNAs, for example, hundreds or even thousands of gRNAs, in a single experiment performed in a relatively short amount of time.
  • the methods developed herein, in some aspects include lentiviral transfection to increase cell transfection efficiency, tagmentation to improve the quality of the DNA extracted from the cells, and a single amplification reaction to limit the rate of false positives.
  • methods herein use gRNA libraries designed to have a fixed edit distance (e.g., measured using Hamming Distance) such that an individual DBS mutation tagged by a DSO at a given genomic locus can be accurately matched to its targeting gRNA.
  • This fixed edit distance introduces sufficient nucleotide differences among the gRNA sequences of a library such that one can be confident that the gRNA with the highest nucleotide percent identity to a given genomic locus is the most active and most specific targeting gRNA.
  • genomic target regions in the cells are separated from each other by at least 10 nucleotide base pairs (e.g., at least 20 nucleotide base pairs, at least 30 nucleotide base pairs, at least 40 nucleotide base pairs, at least 50 nucleotide base pairs, or at least 100 nucleotide base pairs). In some embodiments, genomic target regions in the cells are separated from each other by 10-100, 20-90, 20-40, or 25- 35 base pairs.
  • genomic locus includes nucleotides in a specific region in the genome.
  • a genomic locus encodes a gene.
  • a genomic locus encodes a functional RNA (e.g., a microRNA, a long non-coding RNA, or a ribosomal RNA).
  • a genomic locus is a non-coding region of the genome (e.g., an enhancer or promoter).
  • a genomic locus is at least 50 base pairs long (e.g., at least 100 base pairs long, at least 250 base pairs long, at least 500 base pairs long, at least 1000 base pairs long, at least 5000 base pairs long, or at least 10,000 base pairs long). “Genomic loci” refers to more than one genomic locus.
  • this disclosure provides a method for identifying one or more highly active and/or highly specific guide RNA (gRNA).
  • gRNA highly specific guide RNA
  • a “guide RNA” or gRNA includes an RNA polynucleotide that is capable of directing an RNA-guided nuclease (RGN) to a target polynucleotide (e.g., a genomic locus).
  • RGN RNA-guided nuclease
  • a gRNA is typically a short synthetic RNA composed of two parts: (1) a CRISPR RNA (crRNA) that contains a sequence (a homology region) of about 20 nucleotides (though it can be longer or shorter) that is complementary (e.g., wholly complementary) to a target DNA sequence, and (2) a trans-activating crRNA (tracrRNA) that binds to the crRNA and to a CRISPR-associated (Cas) protein, for example, Cas9.
  • crRNA CRISPR RNA
  • tracrRNA trans-activating crRNA
  • a gRNA is a sgRNA.
  • a gRNA directs the Cas protein to the specific DNA sequence, where the Cas protein induces a DSB.
  • the term “gRNA” includes a two-part gRNA as well as a sgRNA unless stated otherwise.
  • homology regions are designed to be complementary to a specific polynucleotide sequence (also referred to as an on- target polynucleotide sequence) and not complementary to other polynucleotide sequences (also referred to as an off-target polynucleotide sequence).
  • a gRNA is operably linked to a promoter.
  • a promoter is a constitutive promoter (e.g., a SV40, CMV, UBC, U6, EF1A, PGK or CAGG promoter).
  • a promoter is an inducible promoter (e.g., a TET promoter).
  • a promoter is a U6 promoter.
  • a promoter is a tRNA promoter.
  • the tRNA promoter is a eukaryotic tRNA promoter.
  • the tRNA promoter is a prokaryotic tRNA promoter.
  • the tRNA promoter is human tRNA promoter. In some embodiments, the tRNA promoter is an Arabidopsis tRNA promoter. In some embodiments, the tRNA promoter is a glycine or alanine promoter. In some embodiments, a promoter is a human cysteine tRNA (hCtRNA) promoter.
  • hCtRNA human cysteine tRNA
  • “Complementary” refers to the relationship between two polynucleotides (DNA or RNA) in which each nucleotide on one strand pairs (binds) specifically with a corresponding nucleotide on another strand.
  • This complementary base pairing is driven by hydrogen bonds: A forms two hydrogen bonds with T (or U in RNA), and C forms three hydrogen bonds with G.
  • a gRNA is 100% complementary to a target gene sequence. That is, each nucleotide of the gRNA is paired (bound) to the target gene sequence.
  • a gRNA is less than 100% complementary to target gene sequence. For example, there can be one or more mismatches between a gRNA and a target gene sequence.
  • a “highly active guide RNA” includes a gRNA that, when used with an RGN, induces a DSB in an on-target polynucleotide sequence (e.g., a location in a gene locus to which the gRNA was designed to target/bind) in the cells more frequently than most (e.g., greater than 50%) other gRNAs of the gRNA library induce DSBs in corresponding on-target polynucleotide sequences.
  • an on-target polynucleotide sequence e.g., a location in a gene locus to which the gRNA was designed to target/bind
  • a highly active guide RNA includes a gRNA that, when used with an RGN, induces a DSB in an on-target polynucleotide sequence (e.g., a location in a gene locus to which the gRNA was designed to target/bind) in the cells more frequently than most other gRNAs of the gRNA library that target the on-target polynucleotide sequence.
  • a highly active guide RNA includes a gRNA that, when used with an RGN, induces a DSB in an on-target polynucleotide sequence in the cells more frequently than 75%, 85%, 90%, 95%, 98%, or 99% of the other gRNAs of the gRNA library.
  • Comparing the frequency with which two or more guide RNAs induce a DSB in an on-target polynucleotide can be performed in any suitable means including those described herein (e.g., by comparing the relative activity of two or more gRNAs).
  • a highly active gRNA when used with an RGN, induces a DSB in an on-target polynucleotide sequence in at least 20% (e.g., at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, and at least 80%, at least 90%, at least 95%, at least 98%, or at least 99%) of cells transfected with the highly active gRNA.
  • Methods of determining if a given gRNA is a highly active gRNA are described herein including in the section entitled “Identifying a Highly Active and/or Highly Specific gRNA.”
  • a “highly specific guide RNA” includes a gRNA that, when used with an RGN, induces a DSB in an on-target polynucleotide sequence more frequently than it induces a DSB in an off- target polynucleotide sequence.
  • a highly specific guide RNA has higher relative specificity (e.g., as described herein) than at least 75% (e.g., at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%) of guide RNAs of a guide RNA library comprising the highly specific guide RNA.
  • a highly specific gRNA when inducing an RGN DSB, induces the DSB in the on-target polynucleotide sequence at least 51% of the times that the RGN induces a DSB in a respective genome of the cells (e.g., at least 60% of the times, at least 70% of the times, at least 80% of the times, at least 90% of the times, at least 95% of the times, at least 98% of the time, at least 99% of the times, or at least 99.5% of the times).
  • a highly specific gRNA when inducing an RGN DSB, induces the DSB in the on-target polynucleotide sequence 100% of the times that the RGN induces a DSB in a respective genome of the cells.
  • a highly specific gRNA when inducing an RGN DSB, induces the DSB in an off-target polynucleotide sequence less than 50% of the times that the RGN induces a DSB (e.g., less than 40% of the times, less than 30% of the times, less than 20% of the times, less than 10% of the times, less than 5% of the times, less than 2% of the times, less than 1% of the times, or less than 0.5% of the times).
  • a highly specific gRNA when inducing an RGN DSB, does not induce the DSB in an off-target polynucleotide. Methods of determining if a given gRNA is a highly specific gRNA are described herein including in the section entitled “Identifying a Highly Active and/or Highly Specific gRNA.”
  • a method comprises: (a) transfecting cells comprising an RGN and a DSO with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO.
  • Transfecting includes introducing a polynucleotide (e.g., a gRNA) into a cell.
  • transfecting comprises transduction.
  • transfecting comprises transformation.
  • transfection comprises mechanical transfection (e.g., electroporation).
  • transfection comprises non-viral transfection (e.g., lipofection).
  • transfection comprises nucleofection.
  • transfection comprises viral transfection.
  • viral transfection comprises transfecting with a retrovirus that comprises or encodes a polynucleotide that is introduced into the cells.
  • transfecting with a retrovirus comprises transfecting with a murine leukemia virus, Amphotropic murine retroviruses, Gibbon ape leukemia virus (GALV), or a Reticuloendothelial virus.
  • transfecting comprises transfecting with a lentivirus.
  • transfecting further comprises transfecting using a transfection enhancing reagent (e.g., polybrene, LENTIBOOST®, protamine).
  • transfection comprises transfecting using a lentivirus and a transfection enhancing agent.
  • transfection comprises transfecting using a lentivirus and LENTIBOOST®.
  • transfecting with a lentivirus comprises transfecting with a transfection enhancer (e.g., LENTIBOOST®) at a dilution of 1/100 to 1/300 of total volume of the transfection culture (a composition comprising culture media and cells transfected with a lentivirus, for example).
  • transfecting with a lentivirus comprises transfecting with a transfection enhancer (e.g., LENTIBOOST®) at a dilution of about 1/200 of total volume of the transfection culture.
  • transfecting comprises transfecting a gRNA or a DNA encoding a gRNA; a DNA or RNA encoding an RGN; and/or a DSO. In some embodiments, transfecting comprises transfecting with a multiplicity of infection (MOI) that is less than 1. In some embodiments, transfecting comprises transfecting with a multiplicity of infection (MOI) that is about 0.8. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.6. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.4. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.35.
  • MOI multiplicity of infection
  • MOI multiplicity of infection
  • transfecting comprises transfecting with a MOI that is about 0.6. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.4. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.35.
  • transfecting comprises transfecting with a MOI that is about 0.3. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.2. In some embodiments, transfecting comprises transfecting with a MOI that is less than 0.6 (e.g., less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1). “About” refers to within 5% of a numerical value. In some embodiments, transfecting comprising transfection a gRNA library into cells.
  • transfecting with a lentivirus comprises transfecting with a MOI that is between 0.2-1 and a transfection enhancer (e.g., LENTIBOOST®) at a dilution of 1/100 to 1/300 of total volume of the transfection culture.
  • transfecting with a lentivirus comprises transfecting with a MOI that is between 0.3-0.8 and a transfection enhancer (e.g., LENTIBOOST®) at a dilution of about 1/200 of total volume of the transfection culture.
  • the method comprises transfection using a selection cassette and corresponding selection reagent (e.g., a puromycin resistance cassette encoded in the transfected polynucleotide and selection with puromycin).
  • transfecting with a lentivirus comprises transfecting with a MOI that less than 0.5 (e.g., about 0.35) and a transfection enhancer (e.g., LENTIBOOST®) at a dilution of 1/100 to 1/300 of total volume of the transfection culture.
  • transfecting with a lentivirus comprises transfecting with a multiplicity of infection (MOI) that is between less than 0.5 (e.g., about 0.35) and a transfection enhancer (e.g., LENTIBOOST®) at a dilution of about 1/200 of total volume of the transfection culture.
  • transfecting with a lentivirus comprises transfecting with a multiplicity of infection (MOI) that results in less than 10% (e.g., less than 7%, less than 5%, less than 2.5%, or less than 1%) of cells receiving multiple viral particles.
  • the gRNA library comprises a lentiviral library encoding the different gRNAs.
  • the different gRNAs are encoded on minus strands of lentiviruses of the lentiviral library, and optionally wherein the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1 (e.g., less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1).
  • MOI multiplicity of infection
  • the different gRNAs are encoded on plus strands of lentiviruses of the lentiviral library, and optionally wherein the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1 (e.g., less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1).
  • MOI multiplicity of infection
  • the different gRNAs are encoded on plus and/or minus strands of lentiviruses of the lentiviral library, and optionally wherein the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1 (e.g., less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1).
  • MOI multiplicity of infection
  • a “cell” can include any cell type.
  • a cell is a eukaryotic cell.
  • the cell is an insect cell, a fungal cell, a reptile cell, a mammalian cell, or an amphibian cell.
  • the cell is an animal cell.
  • the cell is a murine cell.
  • the cell is a human cell.
  • the cell is an immune cell.
  • the cell is a T cell.
  • the cell is a primary T cell.
  • the cell is isolated from an organism. In some embodiments, the cell is isolated from an organism and cultured in vitro.
  • Cells include a plurality of cells of the same type (e.g., a plurality of T cells) or a plurality of cells that comprises different cell types (e.g., a plurality of cells comprises T cells and natural killer cells).
  • a method comprises transfecting by contacting at least 10 million cells (e.g., at least 20 million cells, at least 30 million cells, at least 40 million cells, at least 50 million cells, at least 60 million cells, at least 70 million cells, at least 80 million cells, at least 90 million cells, at least 100 million cells, at least 150 million cells, or at least 200 million cells) with the gRNA library.
  • a method comprises transfecting by contacting about 80 million cells with the gRNA library.
  • the number of cells contacted depends on the number of gRNAs in the gRNA library.
  • transfecting comprises contacting 10-200 million cells with the gRNA library per every 500 different gRNAs in the gRNA library.
  • transfecting comprises contacting at least 10 million cells (e.g., at least 20 million cells, at least 30 million cells, at least 40 million cells, at least 50 million cells, at least 60 million cells, at least 70 million cells, at least 80 million cells, at least 90 million cells, at least 100 million cells, at least 150 million cells, or at least 200 million cells) with the gRNA library per every 500 different gRNAs in the gRNA library.
  • a “guide RNA library” includes a plurality (e.g., two or more) gRNAs.
  • “Different gRNAs” includes gRNAs having different chemical structures. Thus, “different gRNAs” differ relative to each other with respect to at least one characteristic. In some embodiments, different gRNAs have different nucleotide sequences relative to one another. In some embodiments, different gRNAs are approximately the same (or are the same) length relative to one another but have different nucleotide sequences relative to one another. In some embodiments, a gRNA library comprises at least 2 different gRNAs.
  • a gRNA library comprises at least 5 different gRNAs (e.g., at least 10 different gRNAs, at least 100 different gRNAs, at least 500 different gRNAs, at least 1000 different gRNAs, at least 2000 different gRNAs, at least 3000 different gRNAs, at least 4000 different gRNAs, at least 5000 different gRNAs, at least 7500 different gRNAs, at least 10,000 different gRNAs, or at least 20,000 different gRNAs).
  • a gRNA library comprises 100-20,000 different gRNAs.
  • a gRNA library comprises 100-10,000 different gRNAs.
  • a gRNA library comprises 500-10,000 different gRNAs. In some embodiments, a gRNA library comprises 1000-10,000 different gRNAs. In some embodiments, a gRNA library comprises 2500-10,000 different gRNAs. In some embodiments, a gRNA library comprises 100- 5,000 different gRNAs.
  • each different gRNA of the gRNA library differs from every other gRNA of the gRNA library by an edit distance.
  • a “edit distance” includes the number of nucleotides that are different between the region of a given gRNA (e.g., a homology region) that is complementary to a target polynucleotide, and the region of every other different gRNA in the gRNA library that is complementary to a target.
  • a given gRNA can have a homology region that comprises the sequence (AGATAG), and the gRNA library can also comprise a second gRNA having a homology region comprising the sequence of (AGATC) and a third gRNA having a homology region comprising the sequence of (TTATAG).
  • the edit distance between the given gRNA and the second gRNA is 1.
  • the edit distance between the given gRNA and the third gRNA is 2.
  • the edit distance can be determined using a Hamming distance or a Levenshtein distance.
  • the edit distance is used to design gRNA homology regions (of gRNAs of the gRNA library) such the gRNAs can be accurately associated with a given genomic locus (or a specific region of a genomic locus) that comprises an integrated DSO.
  • the edit distance of the gRNAs in the gRNA library is greater than a fixed edit distance (e.g., greater than 3, greater than 4, greater than 5, greater than 6, greater than 7, greater than 8, greater than 9, greater than 10, greater than 11, or greater than 12). In some embodiments, the fixed edit distance of the gRNAs in the gRNA library is greater than 6. In some embodiments, the fixed edit distance of the gRNAs in the gRNA library is greater than 7.
  • a gRNA library comprises different gRNAs that are complementary to genomic target regions in different genomic loci. In some embodiments, a gRNA library comprises different gRNAs that are complementary to different genomic target regions of the same genomic locus (i.e., different regions of the same genomic locus). In some embodiments, a gRNA library comprises different gRNAs that are complementary to different genomic target regions.
  • different gRNAs of a gRNA library are complementary to different genomic target regions that are separated from one another by at least 5 nucleotides (e.g., at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 75 nucleotides, or at least 100 nucleotides).
  • at least 5 nucleotides e.g., at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 75 nucleotides, or at least 100 nucleotides.
  • a method comprises transfecting a gRNA library into cells (e.g., using a method of transfecting as described herein).
  • transfecting a gRNA library into the cells comprises transfecting using a lentiviral library encoding a gRNA library.
  • a lentiviral library can comprise a plurality of lentiviral vectors that collectively encode the gRNAs of a gRNA library.
  • RNA-guided nuclease includes a nuclease that is capable of binding to a gRNA and that is directed to a target polynucleotide by the gRNA.
  • an RGN is a CRIS PR-associated protein (Cas protein) or a variant thereof.
  • a Cas protein is Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxll, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Casl4, C2cl0), Cas 12g, Casl2h, Casl2i, Cas 12k (C2c5), C2c4, C2c8, and C2c9.
  • a Cas protein is Cas9.
  • an RGN comprises a nuclear localization sequence.
  • a “nuclear localization sequence (NLS)” includes an amino acid sequence that, when located in a protein, directs the protein to be imported into the nucleus of a cell e.g., as described in Lu et al. Cell Commun. Signal 19, 60 (2021).
  • the NLS is a monopartite or bipartite NLS.
  • a monopartite NSL typically comprises 4-8 basic amino acids where 4 of the basic amino acids are positively charged.
  • a monopartite NLS comprises a sequence motif of K(K/R)X(K/R) where X is any amino acid residue.
  • a bipartite NSL typically comprises 2-3 clusters of positively charged amino acids that are separated by a 9-12 amino acid region, which comprises a few proline residues.
  • a bipartite NSL comprises a sequence motif of (R/K)(X)io-i2KRXK where X is any amino acid. NLS sequences are further described in Lu, J., Wu, T., Zhang, B. et al. Cell Commun Signal 19, 60 (2021).
  • a nuclear localization sequence comprises a sequence of any one of SEQ ID NOs: 39-41.
  • an RGN comprises 1, 2, 3, 4, 5, or 6 nuclear localization sequences.
  • an RGN comprises 2 nuclear localization sequences.
  • an RGN comprises 3 nuclear localization sequences. In some embodiments, the RGN comprises NLS of SEQ ID NO: 39. In some embodiments, the RGN comprises NLS of SEQ ID NO: 40. In some embodiments, the RGN comprises NLS of SEQ ID NO: 41. In some embodiments, the RGN comprises an NLS SEQ ID NO: 39, SEQ ID NO: 40, and SEQ ID NO: 41. In some embodiments, an RGN is a Cas9 protein comprising 2 nuclear localization sequences. In some embodiments, an RGN is a Cas9 protein comprising 3 nuclear localization sequences.
  • the RGN comprises from n-terminal to c-terminal an amino acid sequence comprising an NLS of SEQ ID NO: 39, a linker of SEQ ID NO: 42, a RGN (e.g., Cas9), a linker of SEQ ID NO: 43, an NLS of SEQ ID NO: 40, a linker of SEQ ID NO: 44, and an NLS of SEQ ID NO: 41.
  • a Cas9 sequence comprises a sequence having at least 80%, at least 85%, at least 90%, or at least 95% identity to the amino acid sequence of SEQ ID NO: 1.
  • a Cas9 sequence comprises the amino acid sequence of SEQ ID NO: 1.
  • a method comprises transfecting DNA encoding an RGN into cells. In some embodiments, a method comprises transfecting RNA encoding an RGN into cells. In some embodiments, a method comprises using cells that express an RGN. In some embodiments, a method comprises nucleofecting an RGN into the cells.
  • a “double- stranded oligodeoxynucleotide (DSO)” includes a double- stranded DNA that is orthologous to the genome of the cells in which the DSO is being integrated (i.e., the strands of the DSO are not present in the genome of the cell and are not complementary to sequences present in the genome of the cell).
  • each strand of a DSO comprises a unique PCR priming sequence.
  • a unique PCR priming sequence is not found in the genome of the host cell.
  • the DSO comprises two PCR primer binding sites, one on each strand).
  • a DSO comprises a restriction enzyme recognition site, optionally a relatively rare site in the genome of the cell.
  • a DSO is about 15-75 base pairs (bp) long. In some embodiments, a DSO is 15-50 bp long, 50-75 bp long, 30-35 bp long, 60-65 bp long, or 50-65 bp long, 15-50 bp long, 20-40 bp long or 30-35 bp long. In some embodiments, a DSO is 32 bp long or 34 bp long. In some embodiments, a DSO comprises a polynucleotide having at least 95% identity to SEQ ID NO: 2, 3, 45 or 46. In some embodiments, a DSO comprises a polynucleotide of SEQ ID NO: 2, 3, 45 or 46.
  • a DSO comprises a first strand of SEQ ID NO: 2 or SEQ ID NO: 45 and a second strand of SEQ ID NO: 3 or SEQ ID NO: 46.
  • a DSO is expressed by the cells.
  • a DSO is integrated into the genome of a cell during GUIDE-SEQ (e.g. Pooled GUIDE-seq-2) via non-homologous end joining (NHEJ) e.g., as described in Tsai, Shengdar Q., et al. Nature biotechnology 33.2 (2015): 187-197.
  • GUIDE-SEQ e.g. Pooled GUIDE-seq-2
  • NHEJ non-homologous end joining
  • a DSO is chemically modified.
  • the 5' end of a DSO is phosphorylated.
  • a phosphorothioate linkage is present on both 3' ends and both 5' ends of a DSO.
  • a DSO is blunt-ended.
  • a DSO comprises randomized 1, 2, 3, 4 or more nucleotide overhangs.
  • a DSO can also include one or more additional modifications, such as those known in the art or described in PCT/US2011/060493.
  • a DSO is biotinylated.
  • a biotinylated version of a DSO is used as a substrate for integration into the DSB site of the genome.
  • Biotin can be anywhere internal to the DSO (e.g., using a modified thymidine residue (biotin-dT), or biotin azide).
  • An “integrated DSO” refers to a DSO that has been inserted into a piece of DNA (e.g., a genomic locus).
  • a method comprises transfecting a DSO into the cells. In some embodiments, a method comprises transfecting about 50-500 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 50-400 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 50-300 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 50-200 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 100-500 pmol of DSO per 0.5-1 million cells.
  • a method comprises transfecting about 100-400 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 100-300 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 100-200 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting 100-200 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting 100-200 pmol of DSO per about 0.6 million cells.
  • a method comprises transfecting at least 50 pmol (e.g., least 50 pmol, least 75 pmol, least 100 pmol, least 125 pmol, least 150 pmol, least 175 pmol, least 200 pmol, least 250 pmol, least 300 pmol, least 400 pmol, or at least 500 pmol) of DSO per about 0.5-1 million cells.
  • a “genomic DSB” refers to a double strand break (DSB) in genomic DNA.
  • a genomic DSB is “tagged by an integrated DSO” when the DSO is inserted into the genomic strand break.
  • the genomic DNA comprises a sequence of:
  • the genomic DNA can comprise a sequence of: 5’-ATC TG-3’
  • the DSO can comprise a sequence of:
  • the genomic DSB tagged by the integrated DSO would have a sequence of: 5’-ATCCC7TAAT7TG-3’ (SEQ ID NO: 69) 3’ -TAGGGAA7TAAAC-5’ (SEQ ID NO: 70)
  • a method further comprises (b) detecting genomic loci of the cells that comprise the integrated DSO. Detecting can be performed using any suitable means. In some embodiments, detecting comprises sequencing the DNA of the cells. Sequencing can be performed using any suitable means including sanger sequencing and/or next-generation sequencing (e.g., ILLUMINA sequencing, PACBIO sequencing, nanopore sequencing, SOLID sequencing, ELEMENT sequencing, or ION TORRENT sequencing). In some embodiments, detecting genomic loci of the cells comprises performing single cell sequencing. In some embodiments, performing single cell sequence comprises sequencing a gRNA sequence of the cells and one or more genomic loci of the cell comprising an integrated DSO.
  • detecting genomic loci of the cells comprises performing single cell sequencing. In some embodiments, performing single cell sequence comprises sequencing a gRNA sequence of the cells and one or more genomic loci of the cell comprising an integrated DSO.
  • detecting comprises preparing the cells for next-generation sequencing.
  • preparing the cells for next-generation sequencing comprises extracting DNA from the cells, fragmenting the DNA, and/or adding sequencing adapter to the DNA (e.g., adding sequencing adapters to the fragmented DNA).
  • DNA of the cells can be extracted from the cells in any suitable way (e.g., by chemical or mechanical cell lysis).
  • DNA of the cells can be fragmented in any suitable manner.
  • fragmenting the DNA comprises mechanically shearing DNA of the cells.
  • fragmenting the DNA does not comprise mechanically shearing DNA of the cells.
  • fragmenting the DNA comprises enzymatically fragmenting the DNA (e.g., using restriction enzymes and/or nicking enzymes).
  • adding sequencing adapters to the DNA comprises ligating sequencing adapters.
  • the type of adapters ligated to the DNA will be determined based on the type of sequencing being used to detect. For example, i5 and i7 primers can be used when performing next-generation sequencing.
  • fragmenting the DNA and adding the sequencing adapters is performed using tagmentation.
  • tagmentation comprises fragmenting and ligating adapter (e.g., a next-generation sequencing adapter sequence, e.g., an i5 or i7 sequencing adapter) to the DNA using a bead-linked Transposome, e.g., as described in Bruinsma S et al. BMC Genomics. 2018 Oct 1 ; 19( 1):722.
  • tagmentation is used to ligate a sequencing adapter of any one of SEQ ID NOs: 4-12 to a DNA fragment (e.g., a DNA fragment comprising a DSO integration).
  • detecting comprises amplifying genomic loci of the cells that comprise the integrated DSO and then sequencing the amplicons. In some embodiments, amplifying is performed after tagmentation. In some embodiments, amplifying comprises amplifying using polymerase chain reaction. In some embodiments, amplifying comprises amplifying using linear amplification. In some embodiments, amplifying comprises amplifying using a pair of primers where one primer of the pair is complementary to the DSO. In some embodiments, amplifying comprises amplifying using a first primer that is complementary to an adapter ligated to a DNA (e.g., an i5 or i7 primer) and a second primer that is complementary to the DSO.
  • a first primer that is complementary to an adapter ligated to a DNA (e.g., an i5 or i7 primer) and a second primer that is complementary to the DSO.
  • the second primer comprises an overhang region comprising an adapter sequence (e.g., and i5 or i7 adapter).
  • tagmentation is used to ligate a first sequencing adapter to the fragmented DNA (e.g., an i5 adapter) and the second primer, during amplification, adds a second sequencing adapter to the fragmented DNA (e.g., an i7 adapter).
  • the second primer is complementary to a region of the DSO that is at least 3 base pairs (e.g., at least 4 base pairs, at least 5 base pairs, at least 6 base pairs, at least 7 base pairs, or at least 8 base pairs) from the 3’ end of the DSO.
  • the first primer comprises the nucleic acid sequence of SEQ ID NO: 13.
  • the second primer comprises the nucleic acid sequence of any one of SEQ ID NOs: 14-37 or 47-58.
  • amplifying comprises amplifying the DNA of the cells comprising using a pair of primers respectively complementary to a region of the DSO that is at least three nucleotides from the 3’ and 5’ terminus of the DSO. In some embodiments, amplifying comprises amplifying the DNA of the cells using of a pair of primers respectively complementary to a region of the DSO that is at least four nucleotides from the 3’ and 5’ terminus of the DSO. In some embodiments, amplifying comprises performing only one PCR reaction. In some embodiments, the method comprises amplifying using an i7 plus (+) strand primer and/or an i7 minus (-) strand primer.
  • the i7 + strand primer and the i7 - strand primer have similar melting temperatures (e.g., within 5 °C, within 4 °C, within 3 °C, within 2 °C, or within 1 °C of one another). Similar melting temperatures may be achieved by modifying the i7 + strand primer and/or the i7 - strand primer length (e.g., by increasing the length of the primer with the lower melting temperature).
  • a method comprises associating a gRNA with a DSO integration (e.g., a DSO integration detected in a genomic locus). In some embodiments, the method comprises associating one or more of the different gRNAs of the gRNA library with one or more genomic loci comprising a DSO integration. In some embodiments, associating a gRNA with a DSO integration comprises identifying a gRNA of a gRNA library that has the highest nucleotide precent identity to the genomic locus comprising the DSO. Highest nucleotide percent identity of a gRNA to a genomic locus can be determined in any suitable way.
  • nucleotide percent identity between a gRNA and a gene locus can be determined using a Hamming distance or Levenshtein distance.
  • associating a gRNA with a DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is within 50 base pairs on either side of the integrated DSO.
  • associating a gRNA with a genomic locus comprising a DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is within 25 base pairs on either side of the integrated DSO. In some embodiments, associating a gRNA with a genomic locus comprising DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus and the region is within sufficient proximity to a protospacer adjacent motif (e.g., prior to DSO integration) to induce a RGN directed DSB.
  • a protospacer adjacent motif e.g., prior to DSO integration
  • the region (e.g., cut site) is within 10 nucleotides (e.g., within 7 nucleotides, within 6 nucleotides, within 5 nucleotides, within 4 nucleotides, within 3 nucleotides or within 2 nucleotides) of the RGN (e.g., Cas9) PAM site.
  • the region (e.g., cut site) is within 10 nucleotides (e.g., within 7 nucleotides, within 6 nucleotides, within 5 nucleotides, within 4 nucleotides, within 3 nucleotides or within 5 nucleotides) of the RGN PAM site and downstream of the RGN PAM site.
  • the region is within 10 nucleotides (e.g., within 7 nucleotides, within 6 nucleotides, within 5 nucleotides, within 4 nucleotides, within 3 nucleotides or within 5 nucleotides) of the RGN PAM site and upstream of the RGN (e.g., Cas9) PAM site.
  • RGN e.g., Cas9
  • upstream and downstream are defined relative to the coding strand.
  • the region (e.g., cut site) is within 6 nucleotides of the RGN (e.g., Cas9) PAM site. In some embodiments, the region (e.g., cut site) is within 6 nucleotides downstream of the RGN (e.g., Cas9) PAM site. In some embodiments, the region (e.g., cut site) is within 6 nucleotides upstream of the RGN (e.g., Cas9) PAM site.
  • the region (e.g., cut site) is within 6 nucleotides upstream of the RGN (e.g., Cas9) PAM site.
  • associating a gRNA with a genomic locus comprising DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is adjacent to the integrated DSO and the region is within sufficient proximity to a protospacer adjacent motif (prior to DSO integration) to induce an RGN directed DSB.
  • associating a gRNA with a genomic locus comprising a DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is within 50 base pairs (e.g., within 25 base pairs) on either side of the integrated DSO.
  • a method comprises associating a plurality of gRNAs with a plurality of a genomic locus comprising a DSO integration.
  • associating a gRNA with a DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is within the Cas RGN cute range from the PAM site on either side of the integrated DSO.
  • multiple guide RNAs can have the highest percent identity to a given genomic locus. In some embodiments, if multiple guide RNAs have the highest percent identity to a given genomic locus then all of the multiple gRNAs are associated with the given genomic locus. In some embodiments, if multiple guide RNAs have the highest percent identity to a given genomic locus then none of the multiple guide RNA are associated with the genomic locus. In some embodiments, the method comprises, before associating, removing genomic loci comprising DSO integrations that have a mismatch number to a given genomic locus of less than 5 (e.g., less than 4, less than 3, or less than 2) to at least two different gRNAs of the gRNA library having different sequences.
  • Mismatch number includes a description of the complementarity between a guide RNA homology region and a target polynucleotide sequence (e.g., a target genomic locus). Mismatch number includes the number of non-complementary base pairs between a guide RNA homology region and a target polynucleotide sequence (e.g., an off-target polynucleotide sequence. For example, if a given gRNA homology region comprises 20 nucleotides and 19 of those nucleotides are complementary to the target polynucleotide then the mismatch number between the given gRNA and the target polynucleotide is 1. If a given gRNA homology region comprises 20 nucleotides, and 17 of those nucleotides are complementary to the target polynucleotide then the mismatch number between the given gRNA and the target polynucleotide is 3.
  • associating comprises determining a gRNA sequence and one or more genomic loci comprising a DSO integration in a given cell that was sequenced using single cell sequencing.
  • transfecting a gRNA library into the cells comprises transfecting with an MOI less than or equal to 1 whereby there is normally 1 gRNA identified per cell.
  • each gRNA of the cell can be associated with corresponding genomic loci and DSO integration based on sequence identity between the genomic loci and the gRNA as described herein (e.g., the gRNA with the highest nucleotide percent identity to the genomic loci is associated with the genomic loci).
  • associating a gRNA with a genomic locus comprises sequencing read support from two molecularly distinct read classes (e.g., bidirectional and/or from both primers). Identifying a Highly Active and/or Highly Specific gRNA
  • identifying a highly active and/or highly specific gRNA from a gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus of the genomic loci with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
  • identifying a highly active gRNA from a gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on- target genomic locus in the cells is higher than a first threshold value.
  • identifying a highly specific gRNA from a gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
  • determining a number of on-target genomic locus comprising an DSO integration in the cells that is associated with a given gRNA comprises determining using next-generation sequencing data of genomic material of the cells. In some embodiments, determining the number of on-target genomic locus comprising the DSO integration comprises determining a number of sequencing reads covering the on-target genomic locus comprising the DSO integration. In some embodiments, the genomic locus is only sequenced when it comprises a DSO integration (e.g., when the DSO integrations are specifically amplified for sequencing).
  • the method comprises determining a relative number of on-target genomic locus comprising the DSO integration by normalizing the number of sequencing reads covering the on-target genomic locus by a relative amount of the given gRNA in the gRNA library (e.g., the gRNA library before transfection or shortly after transfection).
  • the normalization can be used to compare results between different guide RNAs of the guide RNA library.
  • a first gRNA of the gRNA library can comprise a first gRNA at a relative concentration of 3% and a second gRNA at a relative concentration of 6%.
  • Genomic loci comprising DSO integrations associated with the second gRNA can appear twice as frequent and those associated with the first gRNA because the second gRNA was more abundant and thus more frequently transfected than the first gRNA.
  • normalizing by the relative amount of a gRNA in a gRNA library can allow for direct comparison of the activity and specificity of different gRNAs in the library.
  • the relative number of on-target genomic locus comprising a DSO integration is determined by dividing the number of sequencing reads covering the genomic locus comprising the DSO integration and the frequency of the given gRNA in the gRNA library.
  • the relative number of on-target genomic locus comprising the DSO integration identified during sequencing is indicative of the activity (e.g., a relative activity) of the given guide RNA compared to the other guide RNAs in a gRNA library.
  • the method comprises determining a number of off-target genomic loci comprising the DSO integration in the cells that are associated with the given gRNA. In some embodiments, determining a number of off-target genomic loci comprising the DSO integration comprises determining the number of off-target genomic loci (e.g., 1, 2, 3, 4, 5, 6 or more off-target genomic loci) identified when detecting e.g., as described herein. In some embodiments, determining a number of off-target genomic loci comprising the DSO integration comprises determining a number of sequencing reads covering the off-target genomic loci comprising the DSO integration.
  • determining the number of sequencing reads covering the off-target genomic loci comprises determining a relative number of sequencing reads covering the off-target genomic loci. In some embodiments, determining the relative number of sequencing reads comprises determining using the number of sequencing reads covering the off-target genomic loci and the total number of sequencing reads of genomic loci that are associated with the gRNA (i.e., the total number of reads covering the on-target and off-target genomic loci). In some embodiments, determining the relative number of sequencing reads comprises dividing the number of sequencing reads covering off-target genomic loci associated with a given guide RNA by the total number of sequencing reads covering off-target and on-target genomic loci associated with the given guide RNA.
  • the method comprises identifying a given gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value. In some embodiments, the method comprises identifying a given gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value. In some embodiments, the first threshold value is the average relative number of on-target genomic locus associated with each different guide RNA of the gRNA library.
  • the first threshold value is the median relative number of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 75 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 90 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 95 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 98 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library.
  • the first threshold value is the 99 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, identifying a given gRNA as a highly active gRNA comprising identifying the given gRNA as highly active when it has the highest relative activity of the guide RNAs of the gRNA library.
  • the first threshold value is the average relative number of on- target genomic locus associated with each different guide RNAs of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the median relative number of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the 75 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA.
  • the first threshold value is the 90 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the 95 percentile of the relative numbers of on- target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the 98 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA.
  • the first threshold value is the 99 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA.
  • identifying a given gRNA as a highly active gRNA comprising identifying the given gRNA as highly active when it has the highest relative activity of the guide RNAs of the gRNA library that target the same gene locus as the given guide RNA.
  • the method comprises identifying the gRNA as a highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
  • determining the relative amount comprises determining using (i) a number of sequencing reads covering the on-target genomic loci comprising the DSO integration, and (ii) a number of sequencing reads covering the off-target genomic loci comprising the DSO integration.
  • the relative amount is determined by dividing (i) by (i) + (ii) (this is also referred to as relative specificity).
  • determining the second threshold value comprises determining a relative amount for each guide RNA in a gRNA library to produce a plurality of relative amounts. In some embodiments, determining the second threshold value comprises determining a relative amount for each guide RNA in a gRNA library that targets the same genomic locus as the given gRNA to produce a plurality of relative amounts. In some embodiments, the second threshold value is the average of the plurality of relative amounts. In some embodiments, the second threshold value is the median of the plurality of relative amounts. In some embodiments, the second threshold value is the 75 th percentile of the plurality of relative amounts. In some embodiments, the second threshold value is the 90 th percentile of the plurality of relative amounts.
  • the second threshold value is the 95 th percentile of the plurality of relative amounts. In some embodiments, the second threshold value is the 98 th percentile of the plurality of relative amounts. In some embodiments, the second threshold value is the 99 th percentile of the plurality of relative amounts. In some embodiments, identifying a highly active and/or highly specific gRNA from a gRNA library comprises associating one or more of the different gRNAs of the gRNA library using single cell sequencing, wherein the genomic locus comprising a DSO integration with the highest percent identity to the gRNA in a given cell is an on-target genomic locus, and wherein other genomic loci comprising a DSO integration in the given cell are off-target genomic loci.
  • identifying a highly active and/or highly specific gRNA from the gRNA library comprises: determining (i) the number of cells comprising an integrated DSO in the on-target genomic locus and (ii) the number of cells comprising an integrated DSO in the off-target genomic loci; and identifying the gRNA as a highly active gRNA if the number of cells comprising the integrated DSO in the on-target genomic locus is higher than a first threshold value, and/or identifying the gRNA as a highly specific gRNA if the number of off- target genomic loci in the cells is lower than a second threshold value.
  • a method comprises determining (i) the number of cells comprising an integrated DSO in the genomic locus of interest and (ii) the number of other genomic loci in the cells targeted by the gRNA. In some embodiments, determining the number of cells comprising the integrated DSO in the genomic locus of interest comprises determining a relative number of cells (e.g., relative to the total number of cells) that comprise an integrated DSO in the genomic locus of interest. In some embodiments, a method comprises determining the number of cells that comprise the integrated DSO in the genomic locus of interest and the total number of cells. In some embodiments, a method comprises determining a ratio of the number of cells that comprise the integrated DSO in the genomic locus of interest and the total number of cells.
  • determining the number of cells comprises determining, using sequencing data produced by sequencing the cells, (a) the number of reads comprising the integrated DSO in the genomic locus of interest and (a) a total number of reads covering the genomic locus of interest. In some embodiments, a method further comprises determining a ratio between (a) and (b) to determine the relative number of cells comprising an integrated DSO in the genomic locus.
  • a method comprises identifying the gRNA as a highly active gRNA if the number of cells comprising the integrated DSO in the genomic locus of interest is higher than a first threshold value.
  • the number of cells is a relative number of cells (e.g., a fraction or percentage of cells comprising the integrated DSO in the genomic locus).
  • the first threshold value is 15% of cells, 20% of cells, 25% of cells, 30% of cells, 40% of cells, 50% of cells, 60% of cells, 70% of cells, 80% of cells, 90% of cells, 95% of cells, 98% of cells, or 99% of cells comprise the integrated DSO in the genomic locus of interest (or a fractional equivalent thereof).
  • determining the number of cells comprising the integrated DSO in the genomic locus of interest comprises determining if a relative number of next-generation sequencing reads of the genomic locus that comprise the integrated DSO is higher than a first threshold value.
  • the first threshold value is 15% of reads, 20% of reads, 25% of reads, 30% of reads, 40% of reads, 50% of reads, 60% of reads, 70% of reads, 80% of reads, 90% of reads, 95% of reads, 98% of reads, or 99% of reads of the genomic locus of interest comprise the integrated DSO.
  • a method comprises determining the number of other genomic loci in the cells targeted by the gRNA.
  • a genomic locus can be “targeted” by a gRNA when the gRNA (in conjunction with an RGN) introduces a DSB into the genomic locus (e.g., the gRNA is associated with the genomic locus).
  • the number of other genomic loci is a relative number (e.g., relative to the total number of genomic loci associated with the gRNA).
  • determining the number of other genomic loci in the cells targeted by the gRNA comprises determining a number of genomic loci in the cells that are associated with the gRNA and comprise an integrated DSO, optionally excluding the genomic locus of interest.
  • a method comprises identifying the gRNA as a highly specific gRNA if the number of other genomic loci in the cells is lower than a second threshold value.
  • the second threshold value is 1, 2, 3, 4, or 5 of other genomic loci.
  • the second threshold value is 1 other genomic loci.
  • the number of other genomic loci in the cells is a relative number. For example, (a) a number of sequencing reads covering the other genomic loci in the cells targeted by gRNA relative to (b) a number of sequencing reads covering all the genomic loci in the cells targeted by gRNA.
  • the second threshold value is 1%, 2%, 3%, 4%, 5%, 10%, 25%, or 50% of reads of covering the genomic loci targeted by gRNA cover the other genomic loci.
  • a method comprises determining if a plurality of gRNAs (e.g., gRNAs of a gRNA library) are highly active and/or highly specific for a corresponding plurality of genomic loci of interest.
  • gRNAs e.g., gRNAs of a gRNA library
  • this disclosure provides a method for identifying a gRNA for editing a gene of interest. For example, identifying a gRNA that is highly specific and highly active for a gene locus of a gene of interest. “Editing a gene of interest” includes changing the chemical structure of the nucleotides encoding the gene.
  • editing a gene of interest can comprise introducing a mutation into the gene of interest.
  • the mutation is an insertion, deletion, single nucleotide polymorphism, a frameshift, or any combination thereof.
  • a gene of interest can be any gene.
  • a gene of interest can include a gene comprising a mutation that is associated with a disease or disorder.
  • editing the gene of interest comprises correcting a mutation that is associated with the disease or disorder.
  • editing the gene of interest comprises making a mutation in the gene of interest that decreases or eliminates the function of the protein encoded by the gene.
  • editing a gene of interest can comprise introducing a frameshift mutation into the gene of interest that eliminates the function of a protein encoded by the gene of interest.
  • a gene of interest can include a wildtype gene.
  • editing the gene of interest comprises introducing a mutation into a wildtype gene.
  • a method comprises (a) transfecting cells comprising an RGN and a DSO with a gRNA library comprising different gRNAs targeting different genomic regions of a gene of interest. In some embodiments, transfecting the cells produces cells comprising genomic DSBs tagged by an integrated DSO in the gene of interest. In some embodiments, a method further comprises detecting genomic loci in the gene of interest that comprise the integrated DSO. In some embodiments, a method further comprises identifying from a gRNA library a gRNA to be highly active and/or highly specific for the gene of interest using the genomic loci detected (e.g., as described herein).
  • this disclosure provides a method of identifying a Cas enzyme ortholog (e.g., for use with a gRNA targeting a gene of interest).
  • this method can be used to screen the activity and/or specificity of multiple Cas enzyme orthologs on one or more genomic loci to determine which Cas enzyme orthologs have highest activity and/or specificity for the one or more genomic loci.
  • This method can be performed for multiple gRNAs targeting the same genomic locus and/or multiple gRNAs targeting a plurality of genomic loci.
  • Cas enzyme orthologs include RGNs from different organisms.
  • the Cas enzyme ortholog is any Cas enzyme ortholog including a Cas enzyme ortholog described herein.
  • a method comprises identifying a Cas enzyme ortholog that has an improved property relative to a different Cas enzyme ortholog, e.g., increased activity and/or specificity for a given genomic loci. In some embodiments, a method comprises comparing the activity and/or specificity of two or more Cas enzyme orthologs.
  • the method comprises: (a) transfecting first cells with a first Cas enzyme ortholog, a double-stranded oligodeoxynucleotide (DSO), and a first gRNA library comprising different gRNAs to produce first cells comprising genomic DSBs tagged by an integrated DSO; and (b) transfecting second cells with a second Cas enzyme ortholog, a doublestranded oligodeoxynucleotide (DSO), and a second gRNA library comprising different gRNAs to produce second cells comprising genomic DSBs tagged by an integrated DSO.
  • a first Cas enzyme ortholog a double-stranded oligodeoxynucleotide (DSO)
  • a first gRNA library comprising different gRNAs
  • the transfection steps can be performed using any suitable method including those described herein in the section entitled “Transfecting Cells with a Guide RNA Library.”
  • the first cells are transfected with a different DSO than the second cells.
  • Different DSOs include DSOs that have different sequences.
  • the different DSOs differ by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more base pairs.
  • the different DSO differ by such an amount that they are distinguishable by sequencing (e.g., nextgeneration sequencing).
  • the first cells can be any cells including the cells described herein.
  • the second cells can be any cells including the cells described herein. In some embodiments, the first cells and the second cells are the same type of cells. In some embodiments, the first cells and the second cells are the different types of cells. In some embodiments, the first cells and the second cells are T cells.
  • the first Cas enzyme ortholog is any one of Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxl l, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Cas 14, C2cl0), Cas 12g, Casl2h, Casl2i, Cas 12k (C2c5), C2c4, C2c8, and C2c9.
  • the guide RNAs of the first gRNA library are typically compatible with the first Cas enzyme.
  • the second Cas enzyme ortholog is any one of Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxl l, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Cas 14, C2cl0), Cas 12g, Casl2h, Casl2i, Cas 12k (C2c5), C2c4, C2c8, and C2c9.
  • the guide RNAs of the second gRNA library are typically compatible with the second Cas enzyme.
  • the method comprises: (c) combining, after (a) and (b), the first cells and the second cells.
  • Combining can be performed in any suitable manner.
  • combining comprises combining the first cells and second cells in a petri dish or cell culture.
  • combining further comprises culturing the combined first cells and second cells.
  • combining comprises extracting DNA from the first cells, extracting DNA from the second cells, and combining the extracted DNA from the first cells and the extracted DNA from the second cells.
  • the method comprises (d) detecting genomic loci of the first cells that comprise the integrated DSO and detecting genomic loci of the second cells that comprise the integrated DSO. Detecting the genomic loci of the first cells that comprise the integrated DSO can be performed using any suitable means including those described herein in the section entitled, “Detecting Genomic Loci.” Detecting the genomic loci of the second cells that comprise the integrated DSO can be performed using any suitable means including those described herein in the section entitled, “Detecting Genomic Loci.” In some embodiments, detecting the genomic loci further comprises associating one or more genomic loci with the first Cas enzyme ortholog or the second Cas enzyme ortholog (e.g., based on the sequence a guide RNA and/or the sequence of the DSO (if different DSOs are used with the first Cas enzyme ortholog and the second Cas enzyme ortholog)).
  • the method further comprises associating one or more gRNAs of the first gRNA library with one or more genomic loci associated with the first Cas enzyme ortholog (e.g., using a method described herein). In some embodiments, the method further comprises associating one or more gRNAs of the second gRNA library with one or more genomic loci associated with the second Cas enzyme ortholog (e.g., using a method described herein).
  • the method comprises (e) identifying, using one or more genomic loci detected in (d), the Cas enzyme ortholog having the highest activity and/or specificity for the one or more genomic loci.
  • activity e.g., relative activity
  • activity of a given guide RNA can be determined using any suitable means including those described herein in the section entitled, “Identifying a Highly Active and/or Highly Specific gRNA.
  • specificity e.g., relative specificity of a given guide RNA can be determined using any suitable means including those described herein in the section entitled, “Identifying a Highly Active and/or Highly Specific gRNA.
  • this method can be performed with 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50 or more Cas enzyme orthologs by adding a transfecting step for each Cas enzyme ortholog and combining the transfected cells prior to detection. In some embodiments, this method can be performed with at least 2 (e.g., at least 3, at least 5, at least 10, at least 25, or at least 50) Cas enzyme orthologs by adding a transfecting step for each Cas enzyme ortholog and combining the transfected cells prior to detection.
  • this disclosure provides a method of identifying a Cas enzyme variant gene editing system.
  • the method comprises comparing the activity and/or specificity of two or more different Cas enzyme variants (e.g., Cas9 enzyme variants) on the same library of guide RNAs.
  • the library of guide RNAs can be any suitable library of guide RNAs including those described herein.
  • the method comprises identifying a Cas enzyme variant that has improved properties (e.g., improved activity and/or specificity) with a given guide RNA compared to a different Cas enzyme variant.
  • the method comprises (a) transfecting cells with a candidate Cas enzyme variant, a double- stranded oligodeoxynucleotide (DSO), and a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO.
  • the candidate Cas enzyme variant can be any Cas enzyme variant.
  • Transfecting can be performed using any suitable method including those described herein in the section entitled, “Transfecting Cells with a Guide RNA library.”
  • the method comprises performing a first transfection with a first candidate Cas enzyme variant and a second transfection with a second Cas enzyme variant.
  • the method comprises combining the transfected cells from the first transfection and the second transfection (e.g., combining in any suitable way including those described herein).
  • the method further comprises (b) detecting genomic loci of the cells (e.g., the combined cells) that comprise the integrated DSO. Detecting can be performed in any suitable way including those described herein in the section entitled, “Detecting Genomic Loci.”
  • a method further comprises (c) identifying, from a gRNA library, one or more gRNAs as highly active and/or highly specific for a genomic loci detected in (b) (e.g., as described herein).
  • a method comprises repeating (a)-(c) for a plurality of Cas enzyme variant editing systems to obtain a plurality of results.
  • a method comprises analyzing the results to identify one or more gRNAs as highly active and/or highly specific for a given genomic locus with a given Cas enzyme variant editing systems.
  • the method further comprises identifying the Cas enzyme variant gene editing system and gRNA with the corresponding highest activity and/or highest specificity for the given genomic locus.
  • a method comprises repeating step (c) for multiple genetic loci.
  • the method comprises repeating steps (a)-(c) for a plurality of Cas enzymes variant editing system and/or repeating step (a) for a plurality of Cas enzyme variant editing systems, combining the transfected cells and the proceeding with steps (b)-(c).
  • GUIDE-seq-2 is a cell-based method for defining the genome- wide on- and ‘off-target’ activity of genome editors. It is a streamlined version of GUIDE-seq, which is based on the principle of integrating a short end-protected DNA tag into sites of nuclease-induced doublestranded breaks.
  • GUIDE-seq-2 uses a Tn5 -tagmentation library preparation method, and enables selective sequencing of genomic DNA that has incorporated a GUIDE-seq tag. It streamlines library preparation by eliminating a requirement for nested PCR, and also enables primer misamplification detection.
  • GUIDE-seq-2 Scalability of GUIDE-seq-2 has been greatly increased through developing Pooled- GUIDE-seq-2, which allows efficient profiling of cellular Cas9 activities of numerous sgRNAs (e.g., hundreds to thousands) in a high-throughput manner.
  • This method utilizes pooled lentivirus CRISPR sgRNA libraries that contain many different sgRNAs targeting various location of the chromosome together with double- stranded oligodeoxynucleotide (DSO) and Cas9 nucleofection followed by next-generation sequencing of the DSO integration sites.
  • DSO double- stranded oligodeoxynucleotide
  • Lentivirus delivery was further improved by comparing lentivirus transduction enhancer, timing of transduction post T-cell activation, Cas9- NLS variants, and DSO concentration (FIGs. 1A-1C). This resulted in transducing T-cells 24 hours post activation using LENTIBOOST®, using 3xNLS-Cas9, and 100 - 200 pmol of DSO per 0.6 million cells.
  • GUIDE-seq-2 read counts correlated very well between RNP and lentivirus delivery of sgRNA (FIG. ID).
  • a new DSO integration amplification method was also developed.
  • Previous GUIDE-seq methods used nested PCR to amplify a region of a genomic loci comprising a DSO integration. Like many PCR methods, the previous method was prone to mis-priming events during PCR amplification that in turn resulted in false positive identifications of DSO integrations in GUIDE-seq (FIG. 11).
  • a new method was developed that allows for identification of mis- priming events during data analysis, which in turn allows for these mis-priming events to be removed instead of biasing results. This new method replaced the nested PCR with a single PCR step and redesigned primers such that the primer was not complementary the 4 base pairs at a 3’ end of the DSO.
  • the first four base pairs after the primer sequence should be the last four base pairs on a 3’ end of a DSO strand if the primers have bound to the correct location.
  • the guides were intentionally designed to be distinct from each other (FIG. 2B).
  • Up to 5,000 guides were selected from across the genome to have an edit distance of greater than seven (FIG. 2C) (e.g., a Levenshtein distance of greater than seven between all guides in the library of 5,000).
  • GUIDE- seq-2 NGS reads were mapped to the genome, and the genomic sequence was extracted ⁇ 25 bp from the mapped peak. This sequence was then subjected to similarity analysis to re-identify the closest match to the list of guide sequences in the pool. This allowed identification of the on- or off-target sequence, genomic location, frequency of edit, and guide sequence that created the edit and mismatch information (FIG. 2D).
  • the U6 promoter was utilized to express the sgRNA, which artificially constrained the target sequence selection to start with “G”.
  • a lentivirus vector that replaced the U6 promoter was generated with a single human cysteine tRNA (hCtRNA) (FIGs. 3A-3C).
  • hCtRNA human cysteine tRNA
  • the tRNA sequence was cloned into the minus strand to enable lentivirus production (FIG. 3A). Editing efficiency from 6 different protospacer sequences was compared under either the U6 promoter or hCtRNA delivered by lentivirus (FIG. 3B).
  • hCtRNA driven gRNA showed similar editing efficiency.
  • the U6 backbone contained a “G” after the promoter which added a “G” in front of all protospacers.
  • a new lentivirus library of 4435 guides without any G constraints was designed and cloned into the hCtRNA lentivirus vector (FIG. 3C).
  • the number of cells processed was increased from -10,000 cells to 80 million cells per sample. To process this amount, 50 mL tubes were used for bulk tagmentation instead of 96-well plate format and GUIDE-seq-2 PCR was increased from 1 well to 192 wells.
  • FIGs. 4A-4G T-cells transduced by lentivirus
  • FIG. 8 also shows the activity and specificity of the 500 gRNAs.
  • Pooled-GUIDE-SEQ-2 can be performed on a single gene of interest to identify a gRNA that has high specificity (e.g., highly specific) and high activity (e.g., highly active) for that gene (e.g., for use in a therapeutic) or can be performed genome wide (see FIGs. 7 and 9). Pooled- GUIDE-SEQ-2 can also be used to measure the effects of using different Cas9 orthologs with a gRNA library in order to a identify a Cas9 ortholog-gRNA combination that work well in a given application (e.g., knocking out a target gene) (FIG. 10).
  • Pooled-GUIDE-SEQ-2 allows for measuring the activity and specificity of thousands of gRNAs in a single experiments. Performing previous methods, like Guide-SEQ-2, on even a few hundred gRNAs takes many months of laboratory time. In contrast, Pooled-GUIDE-SEQ-2 can be performed in a just a few weeks on thousands of gRNAs.
  • Criteria used to select protospacer sequence of guides were the following. Selected sequences from a commonly expressed exon from each gene, edit distance >7, position distance >30, Repeat mask overlap ⁇ 10. Maximal gRNA number per gene ⁇ 5.
  • oligos were ordered with protospacers flanked with PCR handle, then PCR amplified to produce Gibson overhangs. The amplified oligos were then inserted into BsmBI digested Lentiguide-Puro by Gibson assembly and further amplified using electrocompetent E. coli. Purified plasmids were then transfected into HEK293T cells with packaging vectors to produce lentivirus. Harvested lentivirus were checked for their transduction efficiency by taking the ratio of puromycin selected / non selected cells in activated T-cells.
  • Frozen primary human T-cells were thawed and activated with CD3/CD28 beads (Trans Act) in X-VIVOTM 15, 10% human serum albumin, 10 ng/mL IL-7, and 10 ng/mL IL- 15. Cells were cultured at 37C, 5%CO2, 100% humidity for 24 hours.
  • the cells were transduced with Lentivirus (0.8 MOI) using Lentiboost® (transduction enhancer). After another 24 hours, puromycin was added to start selection. After 72 hours of selection, the cells were nucleofected with 75 pmol Cas9, 100 pmol DSO per 0.6 million cells using Lonza P3 Primary Nucleofector® kit. Three days post nucleofection, cells were harvested and gDNA was extracted from 80 million cells using Puregene® tissue kit. Collected gDNA was then subjected to tagmentation using 8 uL of assembled Tn5 per 100 ng gDNA input.
  • 100 ng of gDNA was tagmented with a custom Tn5-transposome containing Illumina P5 adapter and an 8-nt i5 barcode followed by a 10-nt UMI (one example AATGATACGGCGACCACCGAGATCTACACGTAAGGAGNNNNNNNNNNTCGTCGGCAGCGTCAGA TGTGTATAAGAGACAG ( SEQ ID NO : 9 ) ) to an average length of 600 bp and purified using SPRI- guanidine magnetic beads.
  • a custom Tn5-transposome containing Illumina P5 adapter and an 8-nt i5 barcode followed by a 10-nt UMI (one example AATGATACGGCGACCACCGAGATCTACACGTAAGGAGNNNNNNNNTCGTCGGCAGCGTCAGA TGTGTATAAGAGACAG ( SEQ ID NO : 9 )
  • SEQ ID NO : 9 AATGATACGGCGACCACCGAGATCTACACGTAAGGAGNNNNNNNNNNTCGTC
  • GUIDE-seq-2 primer sequences Below is the common i5 primer sequence used in both (+) and (-) strand PCR. This anneals to the oligo (in above table) added to the gDNA during tagmentation. Below are the (+) and (-) strand primers with 12 different i7 indexes. One is chosen based on the availability and a need to distinguish different sgRNA pool experiments.
  • N701_+ and N701_- will be used if amplification is performed with N701.
  • the different strands are amplified in separate wells and after quantification, the same concentration from each PCR product is combined for sequencing.
  • DSO containing amplicons were then purified with SPRI beads and was eluted with IDTE.
  • Amplicon range of 250 - 400 bp were purified by size selection gels and were quantified with KAPA library quantification kit.
  • the i7_+ primer (i7 primer targeting the plus strand) sequence was elongated 3 nucleotides to match the melting temperature Tm of the i7_- primers (FIG. 12A), which increased the amplification efficiency (FIG. 12B) and the balance of the final GUIDE-seq-2 mapped regions (FIG. 12C).
  • the GUIDE-seq-2 identified sites were reproducible (FIG. 12D) and it correlated well to previous results (FIG. 12E).
  • the new sequence provided better work-flow productivity and sequencing quality.
  • Added filter for read distribution requires read support from two molecularly distinct read classes (i.e., support for forward and reverse reads (bidirectional) or in reads from each i7 primer (both primers)).
  • T transduction efficiency measured by puromycin selection

Landscapes

  • Health & Medical Sciences (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Engineering & Computer Science (AREA)
  • Chemical & Material Sciences (AREA)
  • Organic Chemistry (AREA)
  • Zoology (AREA)
  • Wood Science & Technology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Biomedical Technology (AREA)
  • General Engineering & Computer Science (AREA)
  • Biotechnology (AREA)
  • Molecular Biology (AREA)
  • Microbiology (AREA)
  • General Health & Medical Sciences (AREA)
  • Biochemistry (AREA)
  • Plant Pathology (AREA)
  • Physics & Mathematics (AREA)
  • Biophysics (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Virology (AREA)
  • Medicinal Chemistry (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

In some aspects, this disclosure provides a high throughput method of identifying highly active and/or highly specific guide RNAs.

Description

HIGH-THROUGHPUT UNBIASED IDENTIFICATION OF DOUBEE-STRANDED DNA BREAKS
RELATED APPLICATIONS
This application claims the benefit under 35 U.S.C. § 119(e) to U.S. Provisional Application No. 63/658667, filed June 11, 2024, entitled “High-throughput Unbiased Identification of Double Stranded DNA Breaks”, the entire contents of which are incorporated herein by reference.
REFERENCE TO AN ELECTRONIC SEQUENCE LISTING
The contents of the electronic sequence listing (S232070000WO00-SEQ-ARM.xml; Size: (87,228 bytes; and Date of Creation: June 9, 2025) is herein incorporated by reference in its entirety.
GOVERNMENT LICENSE RIGHTS
This invention was made with government support under AH57189 awarded by National Institutes of Health. The government has certain rights in the invention.
BACKGROUND
Specificity and activity are critical components in RNA-guided gene editing, particularly within the CRISPR-Cas system, as they directly influence the precision and efficacy of genetic modifications. High specificity ensures that the guide RNA (gRNA) targets only the intended DNA sequence, minimizing off-target effects that could cause unintended genetic alterations and potential adverse effects. This is crucial for applications in therapeutic gene editing, where precision is paramount to avoid unwanted mutations that could lead to disease or dysfunction. Activity, on the other hand, refers to the efficiency with which the nuclease-gRNA complex induces the desired genetic changes. High activity levels ensure that the target gene is effectively and reliably edited, making the process more efficient and reducing the time and resources needed for successful gene modification. Together, specificity and activity are fundamental to achieving accurate, safe, and efficient gene editing, enabling advancements in genetic research, biotechnology, and personalized medicine. SUMMARY
In some aspects, this disclosure provides a method for identifying one or more highly active and/or highly specific guide RNA (gRNA). In some embodiments the method comprises: (a) transfecting cells comprising an RNA-guided nuclease (RGN) and a double- stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO; (b) detecting genomic loci of the cells that comprise the integrated DSO; and (c) identifying one or more highly active and/or highly specific gRNA from the gRNA library using the genomic loci detected in (b).
In some aspects, this disclosure provides a method for identifying one or more highly active guide RNA (gRNA). In some embodiments the method comprises: (a) transfecting cells comprising an RNA-guided nuclease (RGN) and a double- stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO; (b) detecting genomic loci of the cells that comprise the integrated DSO; and (c) identifying one or more highly active from the gRNA library using the genomic loci detected in (b).
In some aspects, this disclosure provides a method for identifying one or more highly specific guide RNA (gRNA). In some embodiments the method comprises: (a) transfecting cells comprising an RNA-guided nuclease (RGN) and a double- stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO; (b) detecting genomic loci of the cells that comprise the integrated DSO; and (c) identifying one or more highly specific from the gRNA library using the genomic loci detected in (b).
In some embodiments, an edit distance between any two different gRNAs is greater than a fixed edit distance. In some embodiments, the fixed edit distance is 5, 6, or 7 (e.g., a hamming distance between any two different gRNAs that is 5, 6, or 7 nucleotides). In some embodiments, genomic target regions in the cells are separated from each other by at least 30 nucleotide base pairs. In some embodiments, the gRNA library comprises at least 100, at least 1000, or at least 5000 different gRNAs, optionally all targeting different sequences within a gene. In some embodiments, detecting of the genomic loci is performed without intentionally shearing DNA of the cells. In some embodiments, detecting of the genomic loci comprises extracting DNA from the cells and processing the DNA using DNA tagmentation.
In some embodiments, detecting of the genomic loci of the cells comprises amplifying and/or sequencing DNA of the cells, optionally using next generation sequence (NGS) technology. In some embodiments, amplifying the DNA of the cells comprising use of a pair of primers respectively complementary to a region of the DSO that is at least three nucleotides from the 3’ and 5’ terminus of the DSO. In some embodiments, amplifying the DNA of the cells comprises performing a polymerase chain reaction (PCR), optionally no more than a single PCR. In some embodiments, the gRNA library comprises a lentiviral library encoding the different gRNAs. In some embodiments, the different gRNAs are encoded on minus strands of lentiviruses of the lentiviral library. In some embodiments, the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1.
In some embodiments, the RGN is a Cas nuclease, optionally a Cas9 nuclease. In some embodiments, the Cas nuclease comprises multiple nuclear localization sequences, optionally three nuclear localization sequences.
In some embodiments, transfecting the cells with the DSOs comprises transfecting the cells with about 50 pmol to about 500 pmol of DSO per 0.5 million cells to about 1 million cells, optionally about 100-200 pmol of DSO per 0.6 million cells.
In some embodiments, each of at least a subset of the different gRNAs of the gRNA library is operably linked to a U6 promoter and/or human cysteine tRNA (hCtRNA).
In some embodiments, the method further comprises associating one or more of the different gRNAs of the gRNA library with one or more genomic loci comprising a DSO integration.
In some embodiments, associating one or more of the different gRNAs of the gRNA library comprises, for each different gRNA, associating the different gRNA with one or more genomic loci comprising an integrated DSO if the homology region of the different gRNA has the highest nucleotide percent identity to the one or more genomic loci compared to the homology regions of the other different gRNAs of the gRNA library. In some embodiments, determining highest nucleotide percent identity comprises determining highest nucleotide percent identity to a region of the one or more genomic loci that is within 25 base pairs of the integrated DSO.
In some embodiments, identifying a highly active and/or highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) an amount (e.g., a number) of on-target genomic locus comprising the DSO integration in the cells, and (ii) an amount (e.g., a number) of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
In some embodiments, identifying a highly active gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) an amount (e.g., a number) of on-target genomic locus comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the amount of on-target genomic locus in the cells is higher than a first threshold value.
In some embodiments, identifying a highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) an amount (e.g., a number) of on-target genomic locus comprising the DSO integration in the cells, and (ii) an amount (e.g., a number) of off- target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
In some embodiments, identifying a highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining an amount (e.g., a number) of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as highly specific gRNA the amount of off-target genomic loci in the cells is higher than a second threshold value.
In some embodiments, determining the number of on-target genomic loci comprises determining a relative number using (i) a number of sequencing reads covering the on-target genomic loci comprising the DSO integration and (ii) a frequency of the gRNA in the gRNA library. In some embodiments, determining the relative amount of on target and off-target genomic loci comprises determining using (i) a number of sequencing reads covering the on- target genomic loci comprising the DSO integration, and (ii) a number of sequencing reads covering the off-target genomic loci comprising the DSO integration. In some embodiments, the first threshold value is the 90th percentile of relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the second threshold value is the 90th percentile of relative amounts of on target and off-target genomic loci for each guide RNA in the gRNA library.
In some embodiments, associating one or more of the different gRNAs of the gRNA library comprises associating using single cell sequencing, wherein the genomic locus comprising a DSO integration with the highest percent identity to the gRNA in a given cell is an on-target genomic locus, and wherein other genomic loci comprising a DSO integration in the given cell are off-target genomic loci.
In some embodiments, identifying a highly active and/or highly specific gRNA from the gRNA library comprises: determining (i) the number of cells comprising an integrated DSO in the on-target genomic locus and (ii) the number of cells comprising an integrated DSO in the off-target genomic loci; and identifying the gRNA as a highly active gRNA if the number of cells comprising the integrated DSO in the on-target genomic locus is higher than a first threshold value, and/or identifying the gRNA as a highly specific gRNA if the number of off- target genomic loci in the cells is lower than a second threshold value.
In some aspects, this disclosure provides a method for identifying a guide RNA (gRNA) for editing a gene of interest, the method comprising: (a) transfecting cells comprising an RNA- guided nuclease (RGN) and a double-stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs targeting different genomic regions of a gene of interest to produce cells comprising genomic DSBs tagged by an integrated DSO in the gene of interest; (b) detecting genomic loci in the gene of interest that comprise the integrated DSO; and (c) identifying from the gRNA library a gRNA that is highly active and/or highly specific for the gene of interest using the genomic loci detected in (b).
In some aspects, this disclosure provides a method for identifying a Cas enzyme variant gene editing system, the method comprising: (a) transfecting cells with a candidate Cas enzyme variant, a double- stranded oligodeoxynucleotide (DSO), and a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO; (b) detecting genomic loci of the cells that comprise the integrated DSO; and (c) identifying from the gRNA library one or more gRNAs as being highly active and/or highly specific for a genomic loci detected in (b), thereby identifying a Cas enzyme variant gene editing system comprising a Cas enzyme variant and one or more gRNAs.
In some aspects, this disclosure provides a method for identifying a Cas enzyme ortholog gene editing system, the method comprising: (a) transfecting first cells with a first Cas enzyme ortholog, a double-stranded oligodeoxynucleotide (DSO), and a first gRNA library comprising different gRNAs to produce first cells comprising genomic DSBs tagged by an integrated DSO; (b) transfecting second cells with a second Cas enzyme ortholog, a double- stranded oligodeoxynucleotide (DSO), and a second gRNA library comprising different gRNAs to produce second cells comprising genomic DSBs tagged by an integrated DSO; (c) combining, after (a) and (b), the first cells and the second cells; (d) detecting genomic loci of the first cells that comprise the integrated DSO and detecting genomic loci of the second cells that comprise the integrated DSO; and (e) identifying, using one or more genomic loci detected in (d), the Cas enzyme ortholog having the highest activity and/or specificity for the one or more genomic loci.
In some aspects, this disclosure provides a method for identifying one or more highly active and/or highly specific guide RNA (gRNA), comprising: (a) transfecting cells comprising a Cas nuclease and a double-stranded oligodeoxynucleotide (DSO) with a lentiviral library encoding at least 100 different gRNAs to produce cells comprising genomic DSBs tagged by integrated DSOs, wherein an edit distance between any two different gRNAs is greater than 5; (b) extracting DNA from the cells and processing the DNA using DNA tagmentation to produce a DNA library; (c) amplifying DNA of the DNA library using a polymerase chain reaction (PCR), wherein the PCR comprises a pair of primers respectively complementary to a region of the DSO that is at least three nucleotides from the 3’ and 5’ terminus of the DSO; (d) sequencing amplified DNA of (c) to detect genomic loci of the cells that comprise the integrated DSOs; and (e) associating the at least 100 different gRNAs with genomic loci; and (f) identifying one or more highly active and/or highly specific gRNA from the gRNA library.
In some embodiments, identifying a highly active and/or highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
In some embodiments, genomic target regions in the cells are separated from each other by at least 30 nucleotide base pairs. In some embodiments, sequencing of the amplified DNA is performed using next generation sequencing. In some embodiments, no more than a single PCR is performed. In some embodiments, the Cas nuclease is a Cas9 nuclease, optionally comprising multiple nuclear localization sequences, optionally three nuclear localization sequences. In some embodiments, transfecting the cells with the DSO comprises transfecting the cells with about 50 pmol to about 500 pmol of DSO per 0.5 million cells to about 1 million cells, optionally about 100-200 pmol of DSO per 0.6 million cells. In some embodiments, each of at least a subset of the different gRNAs of the library is operably linked to a U6 promoter and/or human cysteine tRNA (hCtRNA). In some embodiments, the cells comprise mammalian cells. In some embodiments, the mammalian cells comprise human cells.
BRIEF DESCRIPTION OF DRAWINGS
FIGs. 1A-1D show development of lentivirus sgRNA delivery to activated human T- cells. FIG. 1A shows a comparison of transduction enhancers and timing of transduction to express CTLA4 site09 guide with lentivirus. Purified 2xNLS Cas9 protein was subsequently nucleofected. LENTIBOOST® enhancer after 24 hours of T-cell activation provided the best editing efficiency. From left to right in each condition: 24 hours, 48 hours, 72 hours. FIG. IB shows the effects of nucleofecting different amounts of Cas9 concentration with 100 pmol DSO. Regardless of puromycin condition, 75 pM Cas9 provided the best DSO integration efficiency. From left to right in each condition: 75 pM, 100 pM, 125 pM, and 150 pM Cas9 3X NLS. FIG. 1C shows a comparison of DSO concentration on indel frequency and DSO integration efficiency. Increased DSO concentration correlated with increased DSO integration; however, decreased indel frequency was observed with Cas9 independent integration at 200 pmol. 100 pmol DSO provided maximum editing efficiency and DSO integration. From left to right in each condition: control, Puromycin selection BEFORE nucleofection, Puromycin selection BEFORE and AFTER nucleofection. FIG. ID shows a comparison of GUIDE-seq-2 read counts for each off-target site between traditional RNP delivery (x-axis) and lentivirus delivery (y-axis). Both methods correlated well (R2=0.97).
FIGs. 2A-2D show a Pooled-GUIDE-seq-2 design and analysis pipeline. FIG. 2A shows Pooled-GUIDE-seq-2 experiment workflow. FIG. 2B shows criteria of the sgRNA pooled sequence design. FIG. 2C shows distribution of guide library target sequence location. 5000 guides were selected across the whole genome from both DNA strands. FIG. 2D shows a schematic of the Pooled-GUIDE-seq-2 analysis pipeline. Sequences adjacent to the DSO tag were mapped to the genome, and DSB sequences were extracted based on the window size. The window sequence was then subjected to sequence similarity analysis against the sgRNA protospacer list created in FIG. 2A, providing an output of on/off-target location and mismatch information.
FIGs. 3A-3C show utilization of U6 promoter or hCtRNA promoter to express single guide RNAs (sgRNAs). FIG. 3A shows lentiviral vectors used in preparing sgRNA library. Left: U6 promoter, right: hCtRNA driving the sgRNA transcription. Cellular nucleases RNaseP and RNaseZ released sgRNA post transcription. FIG. 3B shows a comparison of editing efficiency between U6 promoter and hCtRNA to express 6 different sgRNAs. FIG. 3C shows that a library of 4435 sgRNA without a “G” constraint was generated. Coverage of each sgRNA sequence was analyzed by amplicon sequencing from the lentiviral plasmid library (x-axis) and the genome post transduction (y-axis).
FIGs. 4A-4G show Pooled-GUIDE-seq-2 with 500 sgRNA pool. FIG. 4A shows distribution of the number of GUIDE- seq-2 sites identified for each guide from a pooled library of 500. FIG. 4B shows an example of the output information obtained for each guide. FIG. 4C shows reproducibility of two replicate Pooled-GUIDE-seq-2 experiments. Number of the read counts for each identified site is plotted. FIG. 4D shows distribution of on-target activity of the 20 guides sampled from the 500 guide library compared with single guide GUIDE-seq-2. FIG. 4E shows an example of one of the guides analyzed in Pooled-GUIDE-seq-2 (left) and GUIDE- seq-2 (right). All off-targets identified in the Pooled-GUIDE-seq-2 had a mismatch of 6 or 7. FIG. 4F shows a comparison of the ranking of guides from Pooled-GUIDE-seq-2 (x-axis) or single GUIDE-seq-2 (y-axis) based on its on-target activity or specificity. FIG. 4G shows the number of guides detected by Pooled-GUIDE-seq-2 with different mismatch filtering.
FIGs. 5A-5D show Pooled-GUIDE-seq-2 with different filtering threshold. Results of two replicates are shown. FIG. 5A shows total read counts. Each point represents a guide. The x-axis shows replicase 1 and the y-axis shows replicase 2. FIG. 5B shows number of sites identified. Each point represents a guide. FIG. 5C shows histograms of guide specificity from round2. FIG. 5D shows number of read counts. Each point represents the identified on- or off- target site and is greyscaled by the number of mismatches. FIGs. 6A-6B show sequence feature analysis. FIG. 6A shows all GUIDE-seq-2 sites from Pooled-GUIDE-seq-2 from the dataset with 500 library roundl (top panels) and 4501 library (bottom panels) that were analyzed for the number of mismatches, PAM preference, and sequence tolerability. FIG. 6B shows sequence features of the guide sequences from all (left), top 10 high activity (e.g., highly active) and specificity (e.g., highly specific) (middle) and bottom 10 least activity and specificity (right) from the dataset with 500 library roundl (top panels) and 4501 library (bottom panels).
FIG. 7 shows different applications of Pooled-GUIDE-seq-2. Pooled-GUIDE-seq-2 can be used for target based (top) and unbiased (bottom) applications. For target based applications, a focused library on specific disease targets can be designed, and Pooled-GUIDE-seq-2 can be performed using a clinically relevant cell type to allow rapid selection of candidate guides with high activity and high specificity in the relevant genomic context. For unbiased applications, a genome wide sgRNA library can be used to compare results from different cell types, different Cas9 orthologues and other CRISPR editors.
FIG. 8 shows a Manhattan plot depicting each guide at the on-target chromosome (X- axis, separated by color) and the specificity (Y-axis) from a pool of 500 guides. The size of the markers depicts the activity of each guide. Large markers with specificity near 1.0 are high specificity and high activity guide RNAs.
FIG. 9 shows a schematic depicting identification of the most highly active and specific targets within a specific disease-associated genes.
FIG. 10 shows a schematic depicting identification of the most highly active and specific Cas variant among a set of orthologues for a particular gene target.
FIG. 11 shows a schematic depicting PCR amplification of blunt double-stranded oligodeoxynucleotide (dsODN) (e.g., a DSO) using the old GUIDE-Seq method compared to the new GUIDE-seq-2 method described herein. The new GUIDE-SEQ-2 PCR amplification method has fewer PCR steps and allows for detection of off-target amplification during sequencing.
FIGs. 12A-12E show an improvement to GUIDE-Seq-2 when additional nucleotides are added to the i7+ strand primer to match the melting temperate to the i7- strand primer. FIG. 12A shows the original i7+ primer and the new i7+ primer with the additional nucleotides. FIG. 12B shows that using the new i7+ primer improves amplification in GUIDE-Seq-2. Each column from left to right: original primer, +3 nt primer. FIG. 12C shows that using the new i7+ primer improves the balance of the final GUIDE-Seq-2 mapped regions. FIG. 12D shows that reproducibility between using the original i7+ primer and the new i7+ primer. FIG. 12E shows a correlation of GUIDE-Seq-2 results using the original i7+ primer and the new i7+ primer.
FIGs. 13A-13C show results for several Pooled-GUIDE-seq-2 experiments. FIG. 13A shows results from a Pooled-GUIDE-seq-2 experiment on 4,501 guides. FIG. 13B shows results from a Pooled-GUIDE-seq-2 experiment on 4,335 guides using a tRNA promoter instead of a U6 promoter to express the guides, which allows use of guides that start with a “G” nucleotide. FIG. 13C shows results from a Pooled-GUIDE-seq-2 experiment on 552 guides that start with a “G” using a U6 promoter to express the guides.
DETAILED DESCRIPTION
Provided herein, in some aspects, are methods for identifying highly active and/or highly specific gRNAs. An RNA-guided nuclease, such as a CRISPR-Cas nuclease, can bind to a gRNA to form a complex, which can then bind to a target genomic locus that includes a sequence complementary to the gRNA. That locus is then edited by the RNA-guided nuclease. The gRNA of that complex is considered to be highly active if the frequency of on-target mutation produced by the complex is high. Likewise, the gRNA of that complex is considered to have high specificity if the frequency of off-target mutation produced by the complex is low. Identifying a gRNA that is both highly active and highly specific for a target genomic locus can be challenging, and identifying multiple highly active and highly specific gRNAs for multiple genomic loci is even more challenging. The gene-editing field needs a streamlined method for massively parallel profiling of cellular genome-wide activity of genome editors, such as CRISPR genome editors, which enables accurate and efficient identification of highly active and highly specific gRNAs.
Current methods for determining activity and/or specificity of a gRNA are limited in throughput. For example, GUIDE-seq (Genome-wide Unbiased Identification of DSBs Enabled by Sequencing; Tsai SG et al. Nature Biotechnology 2015; 33: 187-197) is most often used to determine the specificity of a single gRNA. With this method, a genome-editing nuclease, such as CRISPR-Cas9, is introduced into the cells to induce targeted double-strand breaks (DSBs) in DNA. A short double-stranded oligonucleotide (DSO) with a unique sequence is also introduced into the cells. This DSO integrates into the DSBs created by the nuclease. After allowing sufficient time for the DSO to integrate, the cells are harvested, and their genomic DNA is extracted. Primers complementary to the DSO are used to amplify the regions of the genome where the DSO has integrated, and the amplified DNA fragments are sequenced using high- throughput sequencing methods. Finally, the sequencing data is analyzed to identify the genomic locations of the DSO, which correspond to the sites of DSBs. This allows for the identification of both on-target and off-target sites of genome editing.
While GUIDE-seq is effective for assessing a single gRNA, GUIDE-seq throughput for assessing specificity and activity of many gRNAs, for example, hundreds or even thousands of gRNAs, is low and can take many months to perform. This is partly because GUIDE-seq lacks a means for normalizing activity among different gRNAs and as such cannot accurately measure activity and specificity of many different gRNAs simultaneously in the same experiment. Instead, GUIDE-seq requires an independent set of experiments for each gRNA. For each gRNA, cells are transfected, a DSB is made, a DSO is integrated, cells are harvested, DNA is extracted, several rounds of nested PCR are used to amplify the DNA, high-throughput sequencing is performed, and then the sequencing data is analyzed. Repeating this set of experiments for each gRNA of a library of hundreds of gRNAs, for example, would take months or longer. Thus, without cost- and time-prohibitive experimentation, one cannot use the current GUIDE-seq protocol to identify from a pool of gRNAs which have the highest specificity and activity.
This disclosure provides, in some aspects, methods for broadly evaluating and designing gRNAs at scale. These methods can be used to simultaneously and accurately measure the activity and/or specificity of many gRNAs, for example, hundreds or even thousands of gRNAs, in a single experiment performed in a relatively short amount of time. The methods developed herein, in some aspects include lentiviral transfection to increase cell transfection efficiency, tagmentation to improve the quality of the DNA extracted from the cells, and a single amplification reaction to limit the rate of false positives. Moreover, in some embodiments methods herein use gRNA libraries designed to have a fixed edit distance (e.g., measured using Hamming Distance) such that an individual DBS mutation tagged by a DSO at a given genomic locus can be accurately matched to its targeting gRNA. This fixed edit distance introduces sufficient nucleotide differences among the gRNA sequences of a library such that one can be confident that the gRNA with the highest nucleotide percent identity to a given genomic locus is the most active and most specific targeting gRNA. In some embodiments, genomic target regions in the cells are separated from each other by at least 10 nucleotide base pairs (e.g., at least 20 nucleotide base pairs, at least 30 nucleotide base pairs, at least 40 nucleotide base pairs, at least 50 nucleotide base pairs, or at least 100 nucleotide base pairs). In some embodiments, genomic target regions in the cells are separated from each other by 10-100, 20-90, 20-40, or 25- 35 base pairs.
A “genomic locus” includes nucleotides in a specific region in the genome. In some embodiments, a genomic locus encodes a gene. In some embodiments, a genomic locus encodes a functional RNA (e.g., a microRNA, a long non-coding RNA, or a ribosomal RNA). In some embodiments, a genomic locus is a non-coding region of the genome (e.g., an enhancer or promoter). In some embodiments, a genomic locus is at least 50 base pairs long (e.g., at least 100 base pairs long, at least 250 base pairs long, at least 500 base pairs long, at least 1000 base pairs long, at least 5000 base pairs long, or at least 10,000 base pairs long). “Genomic loci” refers to more than one genomic locus.
In some aspects, this disclosure provides a method for identifying one or more highly active and/or highly specific guide RNA (gRNA).
A “guide RNA” or gRNA includes an RNA polynucleotide that is capable of directing an RNA-guided nuclease (RGN) to a target polynucleotide (e.g., a genomic locus). A gRNA is typically a short synthetic RNA composed of two parts: (1) a CRISPR RNA (crRNA) that contains a sequence (a homology region) of about 20 nucleotides (though it can be longer or shorter) that is complementary (e.g., wholly complementary) to a target DNA sequence, and (2) a trans-activating crRNA (tracrRNA) that binds to the crRNA and to a CRISPR-associated (Cas) protein, for example, Cas9. In practice, these two components are often combined into a single, chimeric RNA molecule referred to as a single-guide RNA (sgRNA). In some embodiments, a gRNA is a sgRNA. A gRNA directs the Cas protein to the specific DNA sequence, where the Cas protein induces a DSB. Herein, the term “gRNA” includes a two-part gRNA as well as a sgRNA unless stated otherwise. Typically, homology regions (e.g., within a crRNA) are designed to be complementary to a specific polynucleotide sequence (also referred to as an on- target polynucleotide sequence) and not complementary to other polynucleotide sequences (also referred to as an off-target polynucleotide sequence).
In some embodiments, a gRNA is operably linked to a promoter. In some embodiments, a promoter is a constitutive promoter (e.g., a SV40, CMV, UBC, U6, EF1A, PGK or CAGG promoter). In some embodiments, a promoter is an inducible promoter (e.g., a TET promoter). In some embodiments, a promoter is a U6 promoter. In some embodiments, a promoter is a tRNA promoter. In some embodiments, the tRNA promoter is a eukaryotic tRNA promoter. In some embodiments, the tRNA promoter is a prokaryotic tRNA promoter. In some embodiments, the tRNA promoter is human tRNA promoter. In some embodiments, the tRNA promoter is an Arabidopsis tRNA promoter. In some embodiments, the tRNA promoter is a glycine or alanine promoter. In some embodiments, a promoter is a human cysteine tRNA (hCtRNA) promoter.
“Complementary” refers to the relationship between two polynucleotides (DNA or RNA) in which each nucleotide on one strand pairs (binds) specifically with a corresponding nucleotide on another strand. This complementary base pairing is driven by hydrogen bonds: A forms two hydrogen bonds with T (or U in RNA), and C forms three hydrogen bonds with G. In some embodiments, a gRNA is 100% complementary to a target gene sequence. That is, each nucleotide of the gRNA is paired (bound) to the target gene sequence. In other embodiments, a gRNA is less than 100% complementary to target gene sequence. For example, there can be one or more mismatches between a gRNA and a target gene sequence.
A “highly active guide RNA” includes a gRNA that, when used with an RGN, induces a DSB in an on-target polynucleotide sequence (e.g., a location in a gene locus to which the gRNA was designed to target/bind) in the cells more frequently than most (e.g., greater than 50%) other gRNAs of the gRNA library induce DSBs in corresponding on-target polynucleotide sequences. In some embodiments, a highly active guide RNA includes a gRNA that, when used with an RGN, induces a DSB in an on-target polynucleotide sequence (e.g., a location in a gene locus to which the gRNA was designed to target/bind) in the cells more frequently than most other gRNAs of the gRNA library that target the on-target polynucleotide sequence. In some embodiments, a highly active guide RNA includes a gRNA that, when used with an RGN, induces a DSB in an on-target polynucleotide sequence in the cells more frequently than 75%, 85%, 90%, 95%, 98%, or 99% of the other gRNAs of the gRNA library. Comparing the frequency with which two or more guide RNAs induce a DSB in an on-target polynucleotide can be performed in any suitable means including those described herein (e.g., by comparing the relative activity of two or more gRNAs). In some embodiments, a highly active gRNA, when used with an RGN, induces a DSB in an on-target polynucleotide sequence in at least 20% (e.g., at least 25%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, and at least 80%, at least 90%, at least 95%, at least 98%, or at least 99%) of cells transfected with the highly active gRNA. Methods of determining if a given gRNA is a highly active gRNA are described herein including in the section entitled “Identifying a Highly Active and/or Highly Specific gRNA.”
A “highly specific guide RNA” includes a gRNA that, when used with an RGN, induces a DSB in an on-target polynucleotide sequence more frequently than it induces a DSB in an off- target polynucleotide sequence. In some embodiments, a highly specific guide RNA has higher relative specificity (e.g., as described herein) than at least 75% (e.g., at least 85%, at least 90%, at least 95%, at least 98%, or at least 99%) of guide RNAs of a guide RNA library comprising the highly specific guide RNA. In some embodiments, a highly specific gRNA, when inducing an RGN DSB, induces the DSB in the on-target polynucleotide sequence at least 51% of the times that the RGN induces a DSB in a respective genome of the cells (e.g., at least 60% of the times, at least 70% of the times, at least 80% of the times, at least 90% of the times, at least 95% of the times, at least 98% of the time, at least 99% of the times, or at least 99.5% of the times). In some embodiments, a highly specific gRNA, when inducing an RGN DSB, induces the DSB in the on-target polynucleotide sequence 100% of the times that the RGN induces a DSB in a respective genome of the cells. In some embodiments, a highly specific gRNA, when inducing an RGN DSB, induces the DSB in an off-target polynucleotide sequence less than 50% of the times that the RGN induces a DSB (e.g., less than 40% of the times, less than 30% of the times, less than 20% of the times, less than 10% of the times, less than 5% of the times, less than 2% of the times, less than 1% of the times, or less than 0.5% of the times). In some embodiments, a highly specific gRNA, when inducing an RGN DSB, does not induce the DSB in an off-target polynucleotide. Methods of determining if a given gRNA is a highly specific gRNA are described herein including in the section entitled “Identifying a Highly Active and/or Highly Specific gRNA.”
Transfecting Cells with a Guide RNA Library
In some embodiments, a method comprises: (a) transfecting cells comprising an RGN and a DSO with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO.
“Transfecting” includes introducing a polynucleotide (e.g., a gRNA) into a cell. In some embodiments, transfecting comprises transduction. In some embodiments, transfecting comprises transformation. In some embodiments, transfection comprises mechanical transfection (e.g., electroporation). In some embodiments, transfection comprises non-viral transfection (e.g., lipofection). In some embodiments, transfection comprises nucleofection. In some embodiments, transfection comprises viral transfection. In some embodiments, viral transfection comprises transfecting with a retrovirus that comprises or encodes a polynucleotide that is introduced into the cells. In some embodiments, transfecting with a retrovirus comprises transfecting with a murine leukemia virus, Amphotropic murine retroviruses, Gibbon ape leukemia virus (GALV), or a Reticuloendothelial virus. In some embodiments, transfecting comprises transfecting with a lentivirus. In some embodiments, transfecting further comprises transfecting using a transfection enhancing reagent (e.g., polybrene, LENTIBOOST®, protamine). In some embodiments, transfection comprises transfecting using a lentivirus and a transfection enhancing agent. In some embodiments, transfection comprises transfecting using a lentivirus and LENTIBOOST®. In some embodiments, transfecting with a lentivirus comprises transfecting with a transfection enhancer (e.g., LENTIBOOST®) at a dilution of 1/100 to 1/300 of total volume of the transfection culture (a composition comprising culture media and cells transfected with a lentivirus, for example). In some embodiments, transfecting with a lentivirus comprises transfecting with a transfection enhancer (e.g., LENTIBOOST®) at a dilution of about 1/200 of total volume of the transfection culture. In some embodiments, transfecting comprises transfecting a gRNA or a DNA encoding a gRNA; a DNA or RNA encoding an RGN; and/or a DSO. In some embodiments, transfecting comprises transfecting with a multiplicity of infection (MOI) that is less than 1. In some embodiments, transfecting comprises transfecting with a multiplicity of infection (MOI) that is about 0.8. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.6. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.4. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.35. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.3. In some embodiments, transfecting comprises transfecting with a MOI that is about 0.2. In some embodiments, transfecting comprises transfecting with a MOI that is less than 0.6 (e.g., less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1). “About” refers to within 5% of a numerical value. In some embodiments, transfecting comprising transfection a gRNA library into cells. In some embodiments, transfecting with a lentivirus comprises transfecting with a MOI that is between 0.2-1 and a transfection enhancer (e.g., LENTIBOOST®) at a dilution of 1/100 to 1/300 of total volume of the transfection culture. In some embodiments, transfecting with a lentivirus comprises transfecting with a MOI that is between 0.3-0.8 and a transfection enhancer (e.g., LENTIBOOST®) at a dilution of about 1/200 of total volume of the transfection culture. In some embodiments, the method comprises transfection using a selection cassette and corresponding selection reagent (e.g., a puromycin resistance cassette encoded in the transfected polynucleotide and selection with puromycin). In some embodiments, transfecting with a lentivirus comprises transfecting with a MOI that less than 0.5 (e.g., about 0.35) and a transfection enhancer (e.g., LENTIBOOST®) at a dilution of 1/100 to 1/300 of total volume of the transfection culture. In some embodiments, transfecting with a lentivirus comprises transfecting with a multiplicity of infection (MOI) that is between less than 0.5 (e.g., about 0.35) and a transfection enhancer (e.g., LENTIBOOST®) at a dilution of about 1/200 of total volume of the transfection culture. In some embodiments, transfecting with a lentivirus comprises transfecting with a multiplicity of infection (MOI) that results in less than 10% (e.g., less than 7%, less than 5%, less than 2.5%, or less than 1%) of cells receiving multiple viral particles.
In some embodiments, the gRNA library comprises a lentiviral library encoding the different gRNAs. In some embodiments, the different gRNAs are encoded on minus strands of lentiviruses of the lentiviral library, and optionally wherein the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1 (e.g., less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1). In some embodiments, the different gRNAs are encoded on plus strands of lentiviruses of the lentiviral library, and optionally wherein the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1 (e.g., less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1). In some embodiments, the different gRNAs are encoded on plus and/or minus strands of lentiviruses of the lentiviral library, and optionally wherein the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1 (e.g., less than 0.6, less than 0.5, less than 0.4, less than 0.3, less than 0.2, or less than 0.1).
A “cell” can include any cell type. In some embodiments, a cell is a eukaryotic cell. In some embodiments, the cell is an insect cell, a fungal cell, a reptile cell, a mammalian cell, or an amphibian cell. In some embodiments, the cell is an animal cell. In some embodiments, the cell is a murine cell. In some embodiments, the cell is a human cell. In some embodiments, the cell is an immune cell. In some embodiments, the cell is a T cell. In some embodiments, the cell is a primary T cell. In some embodiments, the cell is isolated from an organism. In some embodiments, the cell is isolated from an organism and cultured in vitro. “Cells” include a plurality of cells of the same type (e.g., a plurality of T cells) or a plurality of cells that comprises different cell types (e.g., a plurality of cells comprises T cells and natural killer cells). In some embodiments, a method comprises transfecting by contacting at least 10 million cells (e.g., at least 20 million cells, at least 30 million cells, at least 40 million cells, at least 50 million cells, at least 60 million cells, at least 70 million cells, at least 80 million cells, at least 90 million cells, at least 100 million cells, at least 150 million cells, or at least 200 million cells) with the gRNA library. In some embodiments, a method comprises transfecting by contacting about 80 million cells with the gRNA library. In some embodiments, the number of cells contacted depends on the number of gRNAs in the gRNA library. For example, in some embodiments, transfecting comprises contacting 10-200 million cells with the gRNA library per every 500 different gRNAs in the gRNA library. In some embodiments, transfecting comprises contacting at least 10 million cells (e.g., at least 20 million cells, at least 30 million cells, at least 40 million cells, at least 50 million cells, at least 60 million cells, at least 70 million cells, at least 80 million cells, at least 90 million cells, at least 100 million cells, at least 150 million cells, or at least 200 million cells) with the gRNA library per every 500 different gRNAs in the gRNA library.
A “guide RNA library” includes a plurality (e.g., two or more) gRNAs. “Different gRNAs” includes gRNAs having different chemical structures. Thus, “different gRNAs” differ relative to each other with respect to at least one characteristic. In some embodiments, different gRNAs have different nucleotide sequences relative to one another. In some embodiments, different gRNAs are approximately the same (or are the same) length relative to one another but have different nucleotide sequences relative to one another. In some embodiments, a gRNA library comprises at least 2 different gRNAs. In some embodiments, a gRNA library comprises at least 5 different gRNAs (e.g., at least 10 different gRNAs, at least 100 different gRNAs, at least 500 different gRNAs, at least 1000 different gRNAs, at least 2000 different gRNAs, at least 3000 different gRNAs, at least 4000 different gRNAs, at least 5000 different gRNAs, at least 7500 different gRNAs, at least 10,000 different gRNAs, or at least 20,000 different gRNAs). In some embodiments, a gRNA library comprises 100-20,000 different gRNAs. In some embodiments, a gRNA library comprises 100-10,000 different gRNAs. In some embodiments, a gRNA library comprises 500-10,000 different gRNAs. In some embodiments, a gRNA library comprises 1000-10,000 different gRNAs. In some embodiments, a gRNA library comprises 2500-10,000 different gRNAs. In some embodiments, a gRNA library comprises 100- 5,000 different gRNAs.
In some embodiments, each different gRNA of the gRNA library differs from every other gRNA of the gRNA library by an edit distance. A “edit distance” includes the number of nucleotides that are different between the region of a given gRNA (e.g., a homology region) that is complementary to a target polynucleotide, and the region of every other different gRNA in the gRNA library that is complementary to a target. For example, a given gRNA can have a homology region that comprises the sequence (AGATAG), and the gRNA library can also comprise a second gRNA having a homology region comprising the sequence of (AGATC) and a third gRNA having a homology region comprising the sequence of (TTATAG). In this example, the edit distance between the given gRNA and the second gRNA is 1. In this example, the edit distance between the given gRNA and the third gRNA is 2. In some embodiments, the edit distance can be determined using a Hamming distance or a Levenshtein distance.
In some embodiments, the edit distance is used to design gRNA homology regions (of gRNAs of the gRNA library) such the gRNAs can be accurately associated with a given genomic locus (or a specific region of a genomic locus) that comprises an integrated DSO. The smaller the edit distance between gRNA homology regions, the more difficult it can be to determine which gRNA is associated with a given genomic locus comprising an integrated DSO based on highest percent identity of the gRNA homology to the genomic locus. Edit distance between gRNAs can be less important when associating a given guide RNA with its corresponding on-target genomic locus because typically gRNA homology regions are 100% identical to their corresponding on-target genomic locus and guide RNAs are typically not designed to be 100% identical to more than one genomic locus. Edit distance between gRNAs can be more important when associating a given guide RNA with an off-target genomic locus because the smaller the edit distance between gRNAs the more likely that there can be multiple different gRNAs with the same highest percent identity to the genomic locus. In this case, it can be difficult to determine which of the different gRNAs having the same highest percent identity to the genomic is associated with the genomic locus or whether multiple gRNAs are associated with the genomic locus. In this case, other analyses can be used to determine which of the gRNAs is associated with the genomic locus e.g., as described in the section entitled “Associating a gRNA with a Genomic Loci comprising a DSO integration.”
In some embodiments, the edit distance of the gRNAs in the gRNA library is greater than a fixed edit distance (e.g., greater than 3, greater than 4, greater than 5, greater than 6, greater than 7, greater than 8, greater than 9, greater than 10, greater than 11, or greater than 12). In some embodiments, the fixed edit distance of the gRNAs in the gRNA library is greater than 6. In some embodiments, the fixed edit distance of the gRNAs in the gRNA library is greater than 7.
In some embodiments, a gRNA library comprises different gRNAs that are complementary to genomic target regions in different genomic loci. In some embodiments, a gRNA library comprises different gRNAs that are complementary to different genomic target regions of the same genomic locus (i.e., different regions of the same genomic locus). In some embodiments, a gRNA library comprises different gRNAs that are complementary to different genomic target regions. In some embodiments, different gRNAs of a gRNA library are complementary to different genomic target regions that are separated from one another by at least 5 nucleotides (e.g., at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 75 nucleotides, or at least 100 nucleotides). In some embodiments, a gRNA library comprises different gRNAs that are complementary to different genomic target regions that separated from one another by at least 5 nucleotides (e.g., at least 10 nucleotides, at least 20 nucleotides, at least 30 nucleotides, at least 40 nucleotides, at least 50 nucleotides, at least 75 nucleotides, or at least 100 nucleotides). In some embodiments, a gRNA library comprises different gRNAs that are complementary to different genomic target regions that separated from one another by at least 30 nucleotides.
In some embodiments, a method comprises transfecting a gRNA library into cells (e.g., using a method of transfecting as described herein). In some embodiments, transfecting a gRNA library into the cells comprises transfecting using a lentiviral library encoding a gRNA library. A lentiviral library can comprise a plurality of lentiviral vectors that collectively encode the gRNAs of a gRNA library.
An “RNA-guided nuclease (RGN)” includes a nuclease that is capable of binding to a gRNA and that is directed to a target polynucleotide by the gRNA. In some embodiments, an RGN is a CRIS PR-associated protein (Cas protein) or a variant thereof. In some embodiments, a Cas protein is Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxll, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Casl4, C2cl0), Cas 12g, Casl2h, Casl2i, Cas 12k (C2c5), C2c4, C2c8, and C2c9. In some embodiments, a Cas protein is Cas9.
In some embodiments, an RGN comprises a nuclear localization sequence. A “nuclear localization sequence (NLS)” includes an amino acid sequence that, when located in a protein, directs the protein to be imported into the nucleus of a cell e.g., as described in Lu et al. Cell Commun. Signal 19, 60 (2021). In some embodiments, the NLS is a monopartite or bipartite NLS. A monopartite NSL typically comprises 4-8 basic amino acids where 4 of the basic amino acids are positively charged. In some embodiments, a monopartite NLS comprises a sequence motif of K(K/R)X(K/R) where X is any amino acid residue. A bipartite NSL typically comprises 2-3 clusters of positively charged amino acids that are separated by a 9-12 amino acid region, which comprises a few proline residues. In some embodiments, a bipartite NSL comprises a sequence motif of (R/K)(X)io-i2KRXK where X is any amino acid. NLS sequences are further described in Lu, J., Wu, T., Zhang, B. et al. Cell Commun Signal 19, 60 (2021). In some embodiments, a nuclear localization sequence comprises a sequence of any one of SEQ ID NOs: 39-41. In some embodiments, an RGN comprises 1, 2, 3, 4, 5, or 6 nuclear localization sequences. In some embodiments, an RGN comprises 2 nuclear localization sequences. In some embodiments, an RGN comprises 3 nuclear localization sequences. In some embodiments, the RGN comprises NLS of SEQ ID NO: 39. In some embodiments, the RGN comprises NLS of SEQ ID NO: 40. In some embodiments, the RGN comprises NLS of SEQ ID NO: 41. In some embodiments, the RGN comprises an NLS SEQ ID NO: 39, SEQ ID NO: 40, and SEQ ID NO: 41. In some embodiments, an RGN is a Cas9 protein comprising 2 nuclear localization sequences. In some embodiments, an RGN is a Cas9 protein comprising 3 nuclear localization sequences. In some embodiments, the RGN comprises from n-terminal to c-terminal an amino acid sequence comprising an NLS of SEQ ID NO: 39, a linker of SEQ ID NO: 42, a RGN (e.g., Cas9), a linker of SEQ ID NO: 43, an NLS of SEQ ID NO: 40, a linker of SEQ ID NO: 44, and an NLS of SEQ ID NO: 41. In some embodiments, a Cas9 sequence comprises a sequence having at least 80%, at least 85%, at least 90%, or at least 95% identity to the amino acid sequence of SEQ ID NO: 1. In some embodiments, a Cas9 sequence comprises the amino acid sequence of SEQ ID NO: 1.
In some embodiments, a method comprises transfecting DNA encoding an RGN into cells. In some embodiments, a method comprises transfecting RNA encoding an RGN into cells. In some embodiments, a method comprises using cells that express an RGN. In some embodiments, a method comprises nucleofecting an RGN into the cells.
A “double- stranded oligodeoxynucleotide (DSO)” includes a double- stranded DNA that is orthologous to the genome of the cells in which the DSO is being integrated (i.e., the strands of the DSO are not present in the genome of the cell and are not complementary to sequences present in the genome of the cell). In some embodiments, each strand of a DSO comprises a unique PCR priming sequence. A unique PCR priming sequence is not found in the genome of the host cell. In some embodiments, the DSO comprises two PCR primer binding sites, one on each strand). In some embodiments, a DSO comprises a restriction enzyme recognition site, optionally a relatively rare site in the genome of the cell. In some embodiments, a DSO is about 15-75 base pairs (bp) long. In some embodiments, a DSO is 15-50 bp long, 50-75 bp long, 30-35 bp long, 60-65 bp long, or 50-65 bp long, 15-50 bp long, 20-40 bp long or 30-35 bp long. In some embodiments, a DSO is 32 bp long or 34 bp long. In some embodiments, a DSO comprises a polynucleotide having at least 95% identity to SEQ ID NO: 2, 3, 45 or 46. In some embodiments, a DSO comprises a polynucleotide of SEQ ID NO: 2, 3, 45 or 46. In some embodiments, a DSO comprises a first strand of SEQ ID NO: 2 or SEQ ID NO: 45 and a second strand of SEQ ID NO: 3 or SEQ ID NO: 46. In some embodiments, a DSO is expressed by the cells. In some embodiments, a DSO is integrated into the genome of a cell during GUIDE-SEQ (e.g. Pooled GUIDE-seq-2) via non-homologous end joining (NHEJ) e.g., as described in Tsai, Shengdar Q., et al. Nature biotechnology 33.2 (2015): 187-197.
In some embodiments, a DSO is chemically modified. In some embodiments, the 5' end of a DSO is phosphorylated. In some embodiments, a phosphorothioate linkage is present on both 3' ends and both 5' ends of a DSO. In some embodiments, a DSO is blunt-ended. In some embodiments, a DSO comprises randomized 1, 2, 3, 4 or more nucleotide overhangs. A DSO can also include one or more additional modifications, such as those known in the art or described in PCT/US2011/060493. For example, in some embodiments a DSO is biotinylated. In some embodiments, a biotinylated version of a DSO is used as a substrate for integration into the DSB site of the genome. Biotin can be anywhere internal to the DSO (e.g., using a modified thymidine residue (biotin-dT), or biotin azide). An “integrated DSO” refers to a DSO that has been inserted into a piece of DNA (e.g., a genomic locus).
In some embodiments, a method comprises transfecting a DSO into the cells. In some embodiments, a method comprises transfecting about 50-500 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 50-400 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 50-300 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 50-200 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 100-500 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 100-400 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 100-300 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting about 100-200 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting 100-200 pmol of DSO per 0.5-1 million cells. In some embodiments, a method comprises transfecting 100-200 pmol of DSO per about 0.6 million cells. In some embodiments, a method comprises transfecting at least 50 pmol (e.g., least 50 pmol, least 75 pmol, least 100 pmol, least 125 pmol, least 150 pmol, least 175 pmol, least 200 pmol, least 250 pmol, least 300 pmol, least 400 pmol, or at least 500 pmol) of DSO per about 0.5-1 million cells. A “genomic DSB” refers to a double strand break (DSB) in genomic DNA. A genomic DSB is “tagged by an integrated DSO” when the DSO is inserted into the genomic strand break. For example, the genomic DNA comprises a sequence of:
5’-ATCTG-3’
3’-TAGAC-5’
After the genomic DSB the genomic DNA can comprise a sequence of: 5’-ATC TG-3’
3’ -TAG AC-5’
The DSO can comprise a sequence of:
5' -CCTTAATT-
3’ -GGAATTAA-5'
The genomic DSB tagged by the integrated DSO would have a sequence of: 5’-ATCCC7TAAT7TG-3’ (SEQ ID NO: 69) 3’ -TAGGGAA7TAAAC-5’ (SEQ ID NO: 70)
Detecting Genomic Loci
In some embodiments, a method further comprises (b) detecting genomic loci of the cells that comprise the integrated DSO. Detecting can be performed using any suitable means. In some embodiments, detecting comprises sequencing the DNA of the cells. Sequencing can be performed using any suitable means including sanger sequencing and/or next-generation sequencing (e.g., ILLUMINA sequencing, PACBIO sequencing, nanopore sequencing, SOLID sequencing, ELEMENT sequencing, or ION TORRENT sequencing). In some embodiments, detecting genomic loci of the cells comprises performing single cell sequencing. In some embodiments, performing single cell sequence comprises sequencing a gRNA sequence of the cells and one or more genomic loci of the cell comprising an integrated DSO.
In some embodiments, detecting comprises preparing the cells for next-generation sequencing. In some embodiments, preparing the cells for next-generation sequencing comprises extracting DNA from the cells, fragmenting the DNA, and/or adding sequencing adapter to the DNA (e.g., adding sequencing adapters to the fragmented DNA). DNA of the cells can be extracted from the cells in any suitable way (e.g., by chemical or mechanical cell lysis). DNA of the cells can be fragmented in any suitable manner. In some embodiments, fragmenting the DNA comprises mechanically shearing DNA of the cells. In some embodiments, fragmenting the DNA does not comprise mechanically shearing DNA of the cells. In some embodiments, fragmenting the DNA comprises enzymatically fragmenting the DNA (e.g., using restriction enzymes and/or nicking enzymes). In some embodiments, adding sequencing adapters to the DNA comprises ligating sequencing adapters. In some embodiments, the type of adapters ligated to the DNA will be determined based on the type of sequencing being used to detect. For example, i5 and i7 primers can be used when performing next-generation sequencing.
In some embodiments, fragmenting the DNA and adding the sequencing adapters is performed using tagmentation. In some embodiments, tagmentation comprises fragmenting and ligating adapter (e.g., a next-generation sequencing adapter sequence, e.g., an i5 or i7 sequencing adapter) to the DNA using a bead-linked Transposome, e.g., as described in Bruinsma S et al. BMC Genomics. 2018 Oct 1 ; 19( 1):722. In some embodiments, tagmentation is used to ligate a sequencing adapter of any one of SEQ ID NOs: 4-12 to a DNA fragment (e.g., a DNA fragment comprising a DSO integration).
In some embodiments, detecting comprises amplifying genomic loci of the cells that comprise the integrated DSO and then sequencing the amplicons. In some embodiments, amplifying is performed after tagmentation. In some embodiments, amplifying comprises amplifying using polymerase chain reaction. In some embodiments, amplifying comprises amplifying using linear amplification. In some embodiments, amplifying comprises amplifying using a pair of primers where one primer of the pair is complementary to the DSO. In some embodiments, amplifying comprises amplifying using a first primer that is complementary to an adapter ligated to a DNA (e.g., an i5 or i7 primer) and a second primer that is complementary to the DSO. In some embodiments, the second primer comprises an overhang region comprising an adapter sequence (e.g., and i5 or i7 adapter). In some embodiments, tagmentation is used to ligate a first sequencing adapter to the fragmented DNA (e.g., an i5 adapter) and the second primer, during amplification, adds a second sequencing adapter to the fragmented DNA (e.g., an i7 adapter). In some embodiments, the second primer is complementary to a region of the DSO that is at least 3 base pairs (e.g., at least 4 base pairs, at least 5 base pairs, at least 6 base pairs, at least 7 base pairs, or at least 8 base pairs) from the 3’ end of the DSO. This at least 3 base pairs can allow for detection of mis-priming during PCR (e.g., off-target binding of the second primer) when these base pairs do not precede the primer sequence during sequencing. A previous method of GUIDE-Seq did not use the primer design (e.g., see FIG. 11) and thus could not detect mis-priming events; and therefore, results of the previous GUIDE-SEQ method had more false DSO integration sites due to mis-priming events. In some embodiments, the first primer comprises the nucleic acid sequence of SEQ ID NO: 13. In some embodiments, the second primer comprises the nucleic acid sequence of any one of SEQ ID NOs: 14-37 or 47-58.
In some embodiments, amplifying comprises amplifying the DNA of the cells comprising using a pair of primers respectively complementary to a region of the DSO that is at least three nucleotides from the 3’ and 5’ terminus of the DSO. In some embodiments, amplifying comprises amplifying the DNA of the cells using of a pair of primers respectively complementary to a region of the DSO that is at least four nucleotides from the 3’ and 5’ terminus of the DSO. In some embodiments, amplifying comprises performing only one PCR reaction. In some embodiments, the method comprises amplifying using an i7 plus (+) strand primer and/or an i7 minus (-) strand primer. In some embodiments, the i7 + strand primer and the i7 - strand primer have similar melting temperatures (e.g., within 5 °C, within 4 °C, within 3 °C, within 2 °C, or within 1 °C of one another). Similar melting temperatures may be achieved by modifying the i7 + strand primer and/or the i7 - strand primer length (e.g., by increasing the length of the primer with the lower melting temperature).
Associating a gRNA with a Genomic Loci comprising a DSO integration
In some embodiments, a method comprises associating a gRNA with a DSO integration (e.g., a DSO integration detected in a genomic locus). In some embodiments, the method comprises associating one or more of the different gRNAs of the gRNA library with one or more genomic loci comprising a DSO integration. In some embodiments, associating a gRNA with a DSO integration comprises identifying a gRNA of a gRNA library that has the highest nucleotide precent identity to the genomic locus comprising the DSO. Highest nucleotide percent identity of a gRNA to a genomic locus can be determined in any suitable way. In some embodiments, nucleotide percent identity between a gRNA and a gene locus (e.g., a gene locus comprising a DSO integration) can be determined using a Hamming distance or Levenshtein distance. In some embodiments, associating a gRNA with a DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is within 50 base pairs on either side of the integrated DSO. In some embodiments, associating a gRNA with a genomic locus comprising a DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is within 25 base pairs on either side of the integrated DSO. In some embodiments, associating a gRNA with a genomic locus comprising DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus and the region is within sufficient proximity to a protospacer adjacent motif (e.g., prior to DSO integration) to induce a RGN directed DSB. In some embodiments, the region (e.g., cut site) is within 10 nucleotides (e.g., within 7 nucleotides, within 6 nucleotides, within 5 nucleotides, within 4 nucleotides, within 3 nucleotides or within 2 nucleotides) of the RGN (e.g., Cas9) PAM site. In some embodiments, the region (e.g., cut site) is within 10 nucleotides (e.g., within 7 nucleotides, within 6 nucleotides, within 5 nucleotides, within 4 nucleotides, within 3 nucleotides or within 5 nucleotides) of the RGN PAM site and downstream of the RGN PAM site. In some embodiments, the region (e.g., cut site) is within 10 nucleotides (e.g., within 7 nucleotides, within 6 nucleotides, within 5 nucleotides, within 4 nucleotides, within 3 nucleotides or within 5 nucleotides) of the RGN PAM site and upstream of the RGN (e.g., Cas9) PAM site. The skilled person will understand that when targeting double stranded polynucleotides (e.g., double stranded DNA), upstream and downstream are defined relative to the coding strand.
In some embodiments, the region (e.g., cut site) is within 6 nucleotides of the RGN (e.g., Cas9) PAM site. In some embodiments, the region (e.g., cut site) is within 6 nucleotides downstream of the RGN (e.g., Cas9) PAM site. In some embodiments, the region (e.g., cut site) is within 6 nucleotides upstream of the RGN (e.g., Cas9) PAM site. The skilled person will understand that when targeting double stranded polynucleotides (e.g., double stranded DNA), upstream and downstream are defined relative to the coding strand.
In some embodiments, associating a gRNA with a genomic locus comprising DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is adjacent to the integrated DSO and the region is within sufficient proximity to a protospacer adjacent motif (prior to DSO integration) to induce an RGN directed DSB. In some embodiments, associating a gRNA with a genomic locus comprising a DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is within 50 base pairs (e.g., within 25 base pairs) on either side of the integrated DSO. In some embodiments, a method comprises associating a plurality of gRNAs with a plurality of a genomic locus comprising a DSO integration. In some embodiments, associating a gRNA with a DSO integration comprises identifying the gRNA of a gRNA library that has the highest nucleotide precent identity to a region of the genomic locus that is within the Cas RGN cute range from the PAM site on either side of the integrated DSO.
In some embodiments, multiple guide RNAs can have the highest percent identity to a given genomic locus. In some embodiments, if multiple guide RNAs have the highest percent identity to a given genomic locus then all of the multiple gRNAs are associated with the given genomic locus. In some embodiments, if multiple guide RNAs have the highest percent identity to a given genomic locus then none of the multiple guide RNA are associated with the genomic locus. In some embodiments, the method comprises, before associating, removing genomic loci comprising DSO integrations that have a mismatch number to a given genomic locus of less than 5 (e.g., less than 4, less than 3, or less than 2) to at least two different gRNAs of the gRNA library having different sequences. “Mismatch number” includes a description of the complementarity between a guide RNA homology region and a target polynucleotide sequence (e.g., a target genomic locus). Mismatch number includes the number of non-complementary base pairs between a guide RNA homology region and a target polynucleotide sequence (e.g., an off-target polynucleotide sequence. For example, if a given gRNA homology region comprises 20 nucleotides and 19 of those nucleotides are complementary to the target polynucleotide then the mismatch number between the given gRNA and the target polynucleotide is 1. If a given gRNA homology region comprises 20 nucleotides, and 17 of those nucleotides are complementary to the target polynucleotide then the mismatch number between the given gRNA and the target polynucleotide is 3.
In embodiments, where detecting comprises detecting using single cell sequencing, associating comprises determining a gRNA sequence and one or more genomic loci comprising a DSO integration in a given cell that was sequenced using single cell sequencing. In some embodiments, where detecting comprises detecting using single cell sequencing, transfecting a gRNA library into the cells comprises transfecting with an MOI less than or equal to 1 whereby there is normally 1 gRNA identified per cell. In some embodiments, if more than one gRNA is identified in a cell, each gRNA of the cell can be associated with corresponding genomic loci and DSO integration based on sequence identity between the genomic loci and the gRNA as described herein (e.g., the gRNA with the highest nucleotide percent identity to the genomic loci is associated with the genomic loci).
In some embodiments, associating a gRNA with a genomic locus comprises sequencing read support from two molecularly distinct read classes (e.g., bidirectional and/or from both primers). Identifying a Highly Active and/or Highly Specific gRNA
In some embodiments, identifying a highly active and/or highly specific gRNA from a gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus of the genomic loci with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
In some embodiments, identifying a highly active gRNA from a gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on- target genomic locus in the cells is higher than a first threshold value.
In some embodiments, identifying a highly specific gRNA from a gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
In some embodiments, determining a number of on-target genomic locus comprising an DSO integration in the cells that is associated with a given gRNA comprises determining using next-generation sequencing data of genomic material of the cells. In some embodiments, determining the number of on-target genomic locus comprising the DSO integration comprises determining a number of sequencing reads covering the on-target genomic locus comprising the DSO integration. In some embodiments, the genomic locus is only sequenced when it comprises a DSO integration (e.g., when the DSO integrations are specifically amplified for sequencing). In some embodiments, the method comprises determining a relative number of on-target genomic locus comprising the DSO integration by normalizing the number of sequencing reads covering the on-target genomic locus by a relative amount of the given gRNA in the gRNA library (e.g., the gRNA library before transfection or shortly after transfection). In some embodiments, the normalization can be used to compare results between different guide RNAs of the guide RNA library. For example, a first gRNA of the gRNA library can comprise a first gRNA at a relative concentration of 3% and a second gRNA at a relative concentration of 6%. Genomic loci comprising DSO integrations associated with the second gRNA can appear twice as frequent and those associated with the first gRNA because the second gRNA was more abundant and thus more frequently transfected than the first gRNA. Thus, normalizing by the relative amount of a gRNA in a gRNA library can allow for direct comparison of the activity and specificity of different gRNAs in the library. In some embodiments, the relative number of on-target genomic locus comprising a DSO integration is determined by dividing the number of sequencing reads covering the genomic locus comprising the DSO integration and the frequency of the given gRNA in the gRNA library. In some embodiments, the relative number of on-target genomic locus comprising the DSO integration identified during sequencing is indicative of the activity (e.g., a relative activity) of the given guide RNA compared to the other guide RNAs in a gRNA library.
In some embodiments, the method comprises determining a number of off-target genomic loci comprising the DSO integration in the cells that are associated with the given gRNA. In some embodiments, determining a number of off-target genomic loci comprising the DSO integration comprises determining the number of off-target genomic loci (e.g., 1, 2, 3, 4, 5, 6 or more off-target genomic loci) identified when detecting e.g., as described herein. In some embodiments, determining a number of off-target genomic loci comprising the DSO integration comprises determining a number of sequencing reads covering the off-target genomic loci comprising the DSO integration. In some embodiments, determining the number of sequencing reads covering the off-target genomic loci comprises determining a relative number of sequencing reads covering the off-target genomic loci. In some embodiments, determining the relative number of sequencing reads comprises determining using the number of sequencing reads covering the off-target genomic loci and the total number of sequencing reads of genomic loci that are associated with the gRNA (i.e., the total number of reads covering the on-target and off-target genomic loci). In some embodiments, determining the relative number of sequencing reads comprises dividing the number of sequencing reads covering off-target genomic loci associated with a given guide RNA by the total number of sequencing reads covering off-target and on-target genomic loci associated with the given guide RNA.
In some embodiments, the method comprises identifying a given gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value. In some embodiments, the method comprises identifying a given gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value. In some embodiments, the first threshold value is the average relative number of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the median relative number of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 75 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 90 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 95 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 98 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, the first threshold value is the 99 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library. In some embodiments, identifying a given gRNA as a highly active gRNA comprising identifying the given gRNA as highly active when it has the highest relative activity of the guide RNAs of the gRNA library.
In some embodiments, the first threshold value is the average relative number of on- target genomic locus associated with each different guide RNAs of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the median relative number of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the 75 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the 90 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the 95 percentile of the relative numbers of on- target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the 98 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, the first threshold value is the 99 percentile of the relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library that target the same gene locus as the given guide RNA. In some embodiments, identifying a given gRNA as a highly active gRNA comprising identifying the given gRNA as highly active when it has the highest relative activity of the guide RNAs of the gRNA library that target the same gene locus as the given guide RNA.
In some embodiments, the method comprises identifying the gRNA as a highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value. In some embodiments, determining the relative amount comprises determining using (i) a number of sequencing reads covering the on-target genomic loci comprising the DSO integration, and (ii) a number of sequencing reads covering the off-target genomic loci comprising the DSO integration. In some embodiments, the relative amount is determined by dividing (i) by (i) + (ii) (this is also referred to as relative specificity). In some embodiments, determining the second threshold value comprises determining a relative amount for each guide RNA in a gRNA library to produce a plurality of relative amounts. In some embodiments, determining the second threshold value comprises determining a relative amount for each guide RNA in a gRNA library that targets the same genomic locus as the given gRNA to produce a plurality of relative amounts. In some embodiments, the second threshold value is the average of the plurality of relative amounts. In some embodiments, the second threshold value is the median of the plurality of relative amounts. In some embodiments, the second threshold value is the 75th percentile of the plurality of relative amounts. In some embodiments, the second threshold value is the 90th percentile of the plurality of relative amounts. In some embodiments, the second threshold value is the 95th percentile of the plurality of relative amounts. In some embodiments, the second threshold value is the 98th percentile of the plurality of relative amounts. In some embodiments, the second threshold value is the 99th percentile of the plurality of relative amounts. In some embodiments, identifying a highly active and/or highly specific gRNA from a gRNA library comprises associating one or more of the different gRNAs of the gRNA library using single cell sequencing, wherein the genomic locus comprising a DSO integration with the highest percent identity to the gRNA in a given cell is an on-target genomic locus, and wherein other genomic loci comprising a DSO integration in the given cell are off-target genomic loci.
In some embodiments, identifying a highly active and/or highly specific gRNA from the gRNA library comprises: determining (i) the number of cells comprising an integrated DSO in the on-target genomic locus and (ii) the number of cells comprising an integrated DSO in the off-target genomic loci; and identifying the gRNA as a highly active gRNA if the number of cells comprising the integrated DSO in the on-target genomic locus is higher than a first threshold value, and/or identifying the gRNA as a highly specific gRNA if the number of off- target genomic loci in the cells is lower than a second threshold value.
In embodiments, a method comprises determining (i) the number of cells comprising an integrated DSO in the genomic locus of interest and (ii) the number of other genomic loci in the cells targeted by the gRNA. In some embodiments, determining the number of cells comprising the integrated DSO in the genomic locus of interest comprises determining a relative number of cells (e.g., relative to the total number of cells) that comprise an integrated DSO in the genomic locus of interest. In some embodiments, a method comprises determining the number of cells that comprise the integrated DSO in the genomic locus of interest and the total number of cells. In some embodiments, a method comprises determining a ratio of the number of cells that comprise the integrated DSO in the genomic locus of interest and the total number of cells. In some embodiments, determining the number of cells comprises determining, using sequencing data produced by sequencing the cells, (a) the number of reads comprising the integrated DSO in the genomic locus of interest and (a) a total number of reads covering the genomic locus of interest. In some embodiments, a method further comprises determining a ratio between (a) and (b) to determine the relative number of cells comprising an integrated DSO in the genomic locus.
In some embodiments, a method comprises identifying the gRNA as a highly active gRNA if the number of cells comprising the integrated DSO in the genomic locus of interest is higher than a first threshold value. In some embodiments, the number of cells is a relative number of cells (e.g., a fraction or percentage of cells comprising the integrated DSO in the genomic locus). In some embodiments, the first threshold value is 15% of cells, 20% of cells, 25% of cells, 30% of cells, 40% of cells, 50% of cells, 60% of cells, 70% of cells, 80% of cells, 90% of cells, 95% of cells, 98% of cells, or 99% of cells comprise the integrated DSO in the genomic locus of interest (or a fractional equivalent thereof). In some embodiments, determining the number of cells comprising the integrated DSO in the genomic locus of interest comprises determining if a relative number of next-generation sequencing reads of the genomic locus that comprise the integrated DSO is higher than a first threshold value. In some embodiments, the first threshold value is 15% of reads, 20% of reads, 25% of reads, 30% of reads, 40% of reads, 50% of reads, 60% of reads, 70% of reads, 80% of reads, 90% of reads, 95% of reads, 98% of reads, or 99% of reads of the genomic locus of interest comprise the integrated DSO.
In some embodiments, a method comprises determining the number of other genomic loci in the cells targeted by the gRNA. A genomic locus can be “targeted” by a gRNA when the gRNA (in conjunction with an RGN) introduces a DSB into the genomic locus (e.g., the gRNA is associated with the genomic locus). In some embodiments, the number of other genomic loci is a relative number (e.g., relative to the total number of genomic loci associated with the gRNA). In some embodiments, determining the number of other genomic loci in the cells targeted by the gRNA comprises determining a number of genomic loci in the cells that are associated with the gRNA and comprise an integrated DSO, optionally excluding the genomic locus of interest.
In some embodiments, a method comprises identifying the gRNA as a highly specific gRNA if the number of other genomic loci in the cells is lower than a second threshold value. In some embodiments, the second threshold value is 1, 2, 3, 4, or 5 of other genomic loci. In some embodiments, the second threshold value is 1 other genomic loci. In some embodiments, the number of other genomic loci in the cells is a relative number. For example, (a) a number of sequencing reads covering the other genomic loci in the cells targeted by gRNA relative to (b) a number of sequencing reads covering all the genomic loci in the cells targeted by gRNA. In some embodiments, the second threshold value is 1%, 2%, 3%, 4%, 5%, 10%, 25%, or 50% of reads of covering the genomic loci targeted by gRNA cover the other genomic loci.
In some embodiments, a method comprises determining if a plurality of gRNAs (e.g., gRNAs of a gRNA library) are highly active and/or highly specific for a corresponding plurality of genomic loci of interest.
In some embodiments, this disclosure provides a method for identifying a gRNA for editing a gene of interest. For example, identifying a gRNA that is highly specific and highly active for a gene locus of a gene of interest. “Editing a gene of interest” includes changing the chemical structure of the nucleotides encoding the gene. For example, editing a gene of interest can comprise introducing a mutation into the gene of interest. In some embodiments, the mutation is an insertion, deletion, single nucleotide polymorphism, a frameshift, or any combination thereof. A gene of interest can be any gene. A gene of interest can include a gene comprising a mutation that is associated with a disease or disorder. In some embodiments, editing the gene of interest comprises correcting a mutation that is associated with the disease or disorder. In some embodiments, editing the gene of interest comprises making a mutation in the gene of interest that decreases or eliminates the function of the protein encoded by the gene. For example, editing a gene of interest can comprise introducing a frameshift mutation into the gene of interest that eliminates the function of a protein encoded by the gene of interest. A gene of interest can include a wildtype gene. In some embodiments, editing the gene of interest comprises introducing a mutation into a wildtype gene.
In some embodiments, a method comprises (a) transfecting cells comprising an RGN and a DSO with a gRNA library comprising different gRNAs targeting different genomic regions of a gene of interest. In some embodiments, transfecting the cells produces cells comprising genomic DSBs tagged by an integrated DSO in the gene of interest. In some embodiments, a method further comprises detecting genomic loci in the gene of interest that comprise the integrated DSO. In some embodiments, a method further comprises identifying from a gRNA library a gRNA to be highly active and/or highly specific for the gene of interest using the genomic loci detected (e.g., as described herein).
Identifying Orthologous Cas Enzyme Editing System
In some aspects, this disclosure provides a method of identifying a Cas enzyme ortholog (e.g., for use with a gRNA targeting a gene of interest). In other words, this method can be used to screen the activity and/or specificity of multiple Cas enzyme orthologs on one or more genomic loci to determine which Cas enzyme orthologs have highest activity and/or specificity for the one or more genomic loci. This method can be performed for multiple gRNAs targeting the same genomic locus and/or multiple gRNAs targeting a plurality of genomic loci. In some embodiments, Cas enzyme orthologs include RGNs from different organisms. In some embodiments, the Cas enzyme ortholog is any Cas enzyme ortholog including a Cas enzyme ortholog described herein. In some embodiments, a method comprises identifying a Cas enzyme ortholog that has an improved property relative to a different Cas enzyme ortholog, e.g., increased activity and/or specificity for a given genomic loci. In some embodiments, a method comprises comparing the activity and/or specificity of two or more Cas enzyme orthologs. In some embodiments, the method comprises: (a) transfecting first cells with a first Cas enzyme ortholog, a double-stranded oligodeoxynucleotide (DSO), and a first gRNA library comprising different gRNAs to produce first cells comprising genomic DSBs tagged by an integrated DSO; and (b) transfecting second cells with a second Cas enzyme ortholog, a doublestranded oligodeoxynucleotide (DSO), and a second gRNA library comprising different gRNAs to produce second cells comprising genomic DSBs tagged by an integrated DSO. The transfection steps can be performed using any suitable method including those described herein in the section entitled “Transfecting Cells with a Guide RNA Library.” In some embodiments, the first cells are transfected with a different DSO than the second cells. Different DSOs include DSOs that have different sequences. In some embodiments, the different DSOs differ by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more base pairs. In some embodiments, the different DSO differ by such an amount that they are distinguishable by sequencing (e.g., nextgeneration sequencing).
The first cells can be any cells including the cells described herein. The second cells can be any cells including the cells described herein. In some embodiments, the first cells and the second cells are the same type of cells. In some embodiments, the first cells and the second cells are the different types of cells. In some embodiments, the first cells and the second cells are T cells.
In some embodiments, the first Cas enzyme ortholog is any one of Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxl l, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Cas 14, C2cl0), Cas 12g, Casl2h, Casl2i, Cas 12k (C2c5), C2c4, C2c8, and C2c9. The guide RNAs of the first gRNA library are typically compatible with the first Cas enzyme.
In some embodiments, the second Cas enzyme ortholog is any one of Cas3, Cas8a, Cas5, Cas8b, Cas8c, CaslOd, Csel, Cse2, Csyl, Csy2, Csy3, GSU0054, CaslO, Csm2, Cmr5, CaslO, Csxl l, CsxlO, Csfl, Cas9, Csn2, Cas4, Casl2, Casl2a (Cpfl), Casl2b (C2cl), Casl2c (C2c3), Casl2d (CasY), Casl2e (CasX), Casl2f (Cas 14, C2cl0), Cas 12g, Casl2h, Casl2i, Cas 12k (C2c5), C2c4, C2c8, and C2c9. The guide RNAs of the second gRNA library are typically compatible with the second Cas enzyme.
In some embodiments, the method comprises: (c) combining, after (a) and (b), the first cells and the second cells. Combining can be performed in any suitable manner. In some embodiments, combining comprises combining the first cells and second cells in a petri dish or cell culture. In some embodiments, combining further comprises culturing the combined first cells and second cells. In some embodiments, combining comprises extracting DNA from the first cells, extracting DNA from the second cells, and combining the extracted DNA from the first cells and the extracted DNA from the second cells.
In some embodiments, the method comprises (d) detecting genomic loci of the first cells that comprise the integrated DSO and detecting genomic loci of the second cells that comprise the integrated DSO. Detecting the genomic loci of the first cells that comprise the integrated DSO can be performed using any suitable means including those described herein in the section entitled, “Detecting Genomic Loci.” Detecting the genomic loci of the second cells that comprise the integrated DSO can be performed using any suitable means including those described herein in the section entitled, “Detecting Genomic Loci.” In some embodiments, detecting the genomic loci further comprises associating one or more genomic loci with the first Cas enzyme ortholog or the second Cas enzyme ortholog (e.g., based on the sequence a guide RNA and/or the sequence of the DSO (if different DSOs are used with the first Cas enzyme ortholog and the second Cas enzyme ortholog)). In some embodiments, the method further comprises associating one or more gRNAs of the first gRNA library with one or more genomic loci associated with the first Cas enzyme ortholog (e.g., using a method described herein). In some embodiments, the method further comprises associating one or more gRNAs of the second gRNA library with one or more genomic loci associated with the second Cas enzyme ortholog (e.g., using a method described herein).
In some embodiments, the method comprises (e) identifying, using one or more genomic loci detected in (d), the Cas enzyme ortholog having the highest activity and/or specificity for the one or more genomic loci. In some embodiments, activity (e.g., relative activity) of a given guide RNA can be determined using any suitable means including those described herein in the section entitled, “Identifying a Highly Active and/or Highly Specific gRNA. In some embodiments, specificity (e.g., relative specificity) of a given guide RNA can be determined using any suitable means including those described herein in the section entitled, “Identifying a Highly Active and/or Highly Specific gRNA.
In some embodiments, this method can be performed with 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 50 or more Cas enzyme orthologs by adding a transfecting step for each Cas enzyme ortholog and combining the transfected cells prior to detection. In some embodiments, this method can be performed with at least 2 (e.g., at least 3, at least 5, at least 10, at least 25, or at least 50) Cas enzyme orthologs by adding a transfecting step for each Cas enzyme ortholog and combining the transfected cells prior to detection.
Identifying a Cas Enzyme Variant Editing System
In some embodiments, this disclosure provides a method of identifying a Cas enzyme variant gene editing system. In some embodiments, the method comprises comparing the activity and/or specificity of two or more different Cas enzyme variants (e.g., Cas9 enzyme variants) on the same library of guide RNAs. The library of guide RNAs can be any suitable library of guide RNAs including those described herein. In some embodiments, the method comprises identifying a Cas enzyme variant that has improved properties (e.g., improved activity and/or specificity) with a given guide RNA compared to a different Cas enzyme variant.
In some embodiments, the method comprises (a) transfecting cells with a candidate Cas enzyme variant, a double- stranded oligodeoxynucleotide (DSO), and a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO. The candidate Cas enzyme variant can be any Cas enzyme variant. Transfecting can be performed using any suitable method including those described herein in the section entitled, “Transfecting Cells with a Guide RNA library.” In some embodiments, the method comprises performing a first transfection with a first candidate Cas enzyme variant and a second transfection with a second Cas enzyme variant. In some embodiments, the method comprises combining the transfected cells from the first transfection and the second transfection (e.g., combining in any suitable way including those described herein).
In some embodiments, the method further comprises (b) detecting genomic loci of the cells (e.g., the combined cells) that comprise the integrated DSO. Detecting can be performed in any suitable way including those described herein in the section entitled, “Detecting Genomic Loci.”
In some embodiments, a method further comprises (c) identifying, from a gRNA library, one or more gRNAs as highly active and/or highly specific for a genomic loci detected in (b) (e.g., as described herein). In some embodiments, a method comprises repeating (a)-(c) for a plurality of Cas enzyme variant editing systems to obtain a plurality of results. In some embodiments, a method comprises analyzing the results to identify one or more gRNAs as highly active and/or highly specific for a given genomic locus with a given Cas enzyme variant editing systems. In some embodiments, the method further comprises identifying the Cas enzyme variant gene editing system and gRNA with the corresponding highest activity and/or highest specificity for the given genomic locus. In some embodiments, a method comprises repeating step (c) for multiple genetic loci. In some embodiment, the method comprises repeating steps (a)-(c) for a plurality of Cas enzymes variant editing system and/or repeating step (a) for a plurality of Cas enzyme variant editing systems, combining the transfected cells and the proceeding with steps (b)-(c).
EXAMPLES Example 1: Developing Pooled-GUIDE-seq-2 GUIDE-seq-2 is a cell-based method for defining the genome- wide on- and ‘off-target’ activity of genome editors. It is a streamlined version of GUIDE-seq, which is based on the principle of integrating a short end-protected DNA tag into sites of nuclease-induced doublestranded breaks. GUIDE-seq-2 uses a Tn5 -tagmentation library preparation method, and enables selective sequencing of genomic DNA that has incorporated a GUIDE-seq tag. It streamlines library preparation by eliminating a requirement for nested PCR, and also enables primer misamplification detection.
Scalability of GUIDE-seq-2 has been greatly increased through developing Pooled- GUIDE-seq-2, which allows efficient profiling of cellular Cas9 activities of numerous sgRNAs (e.g., hundreds to thousands) in a high-throughput manner. This method utilizes pooled lentivirus CRISPR sgRNA libraries that contain many different sgRNAs targeting various location of the chromosome together with double- stranded oligodeoxynucleotide (DSO) and Cas9 nucleofection followed by next-generation sequencing of the DSO integration sites.
As a proof of concept, activated primary human T-cells were selected for use with Pooled-GUIDE-seq-2, as they are clinically relevant for genome editing applications such as CAR-T therapy. The first challenge was to obtain good editing efficiency and DSO integration in these T cells. Efficient DSO tag integration is important for GUIDE-seq-2 assay, as tag integration is proportional to GUIDE-seq-2 read counts (reads counts are used in determining gRNA activity and specificity). Initially, delivering guides using a dual plasmid transposase system yielded low editing and DSO integration rates; however, switching to lentiviral-mediated gRNA expression system yielded better results. Lentivirus delivery was further improved by comparing lentivirus transduction enhancer, timing of transduction post T-cell activation, Cas9- NLS variants, and DSO concentration (FIGs. 1A-1C). This resulted in transducing T-cells 24 hours post activation using LENTIBOOST®, using 3xNLS-Cas9, and 100 - 200 pmol of DSO per 0.6 million cells. GUIDE-seq-2 read counts correlated very well between RNP and lentivirus delivery of sgRNA (FIG. ID).
A new DSO integration amplification method was also developed. Previous GUIDE-seq methods used nested PCR to amplify a region of a genomic loci comprising a DSO integration. Like many PCR methods, the previous method was prone to mis-priming events during PCR amplification that in turn resulted in false positive identifications of DSO integrations in GUIDE-seq (FIG. 11). A new method was developed that allows for identification of mis- priming events during data analysis, which in turn allows for these mis-priming events to be removed instead of biasing results. This new method replaced the nested PCR with a single PCR step and redesigned primers such that the primer was not complementary the 4 base pairs at a 3’ end of the DSO. In this new method, when the primer is elongated during PCR, the first four base pairs after the primer sequence should be the last four base pairs on a 3’ end of a DSO strand if the primers have bound to the correct location. However, if the first four base pairs after the primer sequence are not the last four base pairs on a 3’ end of a DSO strand then this indicates a mis-priming event, and that sequence can be excluded from further analysis.
To identify the genome- wide editing profile of each guide from the pool, the guides were intentionally designed to be distinct from each other (FIG. 2B). Up to 5,000 guides were selected from across the genome to have an edit distance of greater than seven (FIG. 2C) (e.g., a Levenshtein distance of greater than seven between all guides in the library of 5,000). GUIDE- seq-2 NGS reads were mapped to the genome, and the genomic sequence was extracted ±25 bp from the mapped peak. This sequence was then subjected to similarity analysis to re-identify the closest match to the list of guide sequences in the pool. This allowed identification of the on- or off-target sequence, genomic location, frequency of edit, and guide sequence that created the edit and mismatch information (FIG. 2D).
Initially, the U6 promoter was utilized to express the sgRNA, which artificially constrained the target sequence selection to start with “G”. To eliminate this 5’ starting G requirement, a lentivirus vector that replaced the U6 promoter was generated with a single human cysteine tRNA (hCtRNA) (FIGs. 3A-3C). To avoid hCtRNA processing by cellular nucleases of the lentivirus packaging cells, the tRNA sequence was cloned into the minus strand to enable lentivirus production (FIG. 3A). Editing efficiency from 6 different protospacer sequences was compared under either the U6 promoter or hCtRNA delivered by lentivirus (FIG. 3B). hCtRNA driven gRNA showed similar editing efficiency. Of note, the U6 backbone contained a “G” after the promoter which added a “G” in front of all protospacers. A new lentivirus library of 4435 guides without any G constraints was designed and cloned into the hCtRNA lentivirus vector (FIG. 3C).
To avoid under- sampling of a pooled sgRNA edit, the number of cells processed was increased from -10,000 cells to 80 million cells per sample. To process this amount, 50 mL tubes were used for bulk tagmentation instead of 96-well plate format and GUIDE-seq-2 PCR was increased from 1 well to 192 wells.
To establish proof-of-concept, an sgRNA library size of 500 driven by the U6 promoter was used on T-cells transduced by lentivirus (FIGs. 4A-4G). Using the Pooled-GUIDE-seq-2 analysis pipeline, the editing profile of all 500 guides was identified (FIG. 4A). For example, a guide with 5 identified GUIDE-seq-2 sites has an on-target edit, typically with the most read counts, and 4 off-target sites with different mismatch numbers compared to the on-target sequence (FIG. 4B). Pooled-GUIDE-seq-2 was reproducible between two experiments (roundl and round2, Pearson correlation r=0.95) (FIG. 4C). Next, 20 guides were sampled with different on-target editing efficiency to compare results from Pooled-GUIDE-seq-2 and the original GUIDE-seq-2 where one target is edited at a time (FIG. 4D). It was observed that often off- target sites with a mismatch of 6 or 7 were not found in single GUIDE-seq-2 results (FIG. 4E), suggesting that some results may be mis-association artifacts. Using edit distance distributions to estimate likelihood of mispairing may minimize gRNA misassociation in Pooled-GUIDE-seq-2. Pooled-GUIDE-seq-2 accurately ranked the on-target editing efficiency well and filtering out mismatch >5 showed good correlation of guide specificity (FIG. 4F). Testing different filter threshold showed that some guides tested in the pooled library showed no on-target activity (FIG. 4G). In general, Pooled-GUIDE-seq-2 was highly reproducible with mismatch <4 showing the best correlation (FIGs. 5A-5D). FIG. 8 also shows the activity and specificity of the 500 gRNAs.
From an expanded library size of 4501 guides, 3965 guides were identified with mismatch <4 filter. Preliminary sequence feature analysis was performed using two different library sizes (FIGs. 6A-6B). The 500 sgRNA library data was analyzed without any filtering, and the 4501 sgRNA library data was analyzed with <4 mismatch filter. Both showed NGG PAM specificity, and showed tolerability of G>A at the 7th and 11th position of the protospacer from both datasets (FIG. 6A). Sequence data showed a preference of higher GC content for highly active and specific guides and higher T content for the worst performing guides (FIG. 6B). This could be due to early termination of the guide transcription by polIII with polyT resulting in truncated guides.
Pooled-GUIDE-SEQ-2 can be performed on a single gene of interest to identify a gRNA that has high specificity (e.g., highly specific) and high activity (e.g., highly active) for that gene (e.g., for use in a therapeutic) or can be performed genome wide (see FIGs. 7 and 9). Pooled- GUIDE-SEQ-2 can also be used to measure the effects of using different Cas9 orthologs with a gRNA library in order to a identify a Cas9 ortholog-gRNA combination that work well in a given application (e.g., knocking out a target gene) (FIG. 10).
Pooled-GUIDE-SEQ-2 allows for measuring the activity and specificity of thousands of gRNAs in a single experiments. Performing previous methods, like Guide-SEQ-2, on even a few hundred gRNAs takes many months of laboratory time. In contrast, Pooled-GUIDE-SEQ-2 can be performed in a just a few weeks on thousands of gRNAs.
Experimental methods sgRNA design
Criteria used to select protospacer sequence of guides were the following. Selected sequences from a commonly expressed exon from each gene, edit distance >7, position distance >30, Repeat mask overlap <10. Maximal gRNA number per gene <5.
Lentivirus library preparation
Pooled oligos were ordered with protospacers flanked with PCR handle, then PCR amplified to produce Gibson overhangs. The amplified oligos were then inserted into BsmBI digested Lentiguide-Puro by Gibson assembly and further amplified using electrocompetent E. coli. Purified plasmids were then transfected into HEK293T cells with packaging vectors to produce lentivirus. Harvested lentivirus were checked for their transduction efficiency by taking the ratio of puromycin selected / non selected cells in activated T-cells.
T-cell culture
Frozen primary human T-cells were thawed and activated with CD3/CD28 beads (Trans Act) in X-VIVO™ 15, 10% human serum albumin, 10 ng/mL IL-7, and 10 ng/mL IL- 15. Cells were cultured at 37C, 5%CO2, 100% humidity for 24 hours.
Pooled-GUIDE- seq-2
Twenty-four hours after T-cell activation, the cells were transduced with Lentivirus (0.8 MOI) using Lentiboost® (transduction enhancer). After another 24 hours, puromycin was added to start selection. After 72 hours of selection, the cells were nucleofected with 75 pmol Cas9, 100 pmol DSO per 0.6 million cells using Lonza P3 Primary Nucleofector® kit. Three days post nucleofection, cells were harvested and gDNA was extracted from 80 million cells using Puregene® tissue kit. Collected gDNA was then subjected to tagmentation using 8 uL of assembled Tn5 per 100 ng gDNA input. Specifically, 100 ng of gDNA was tagmented with a custom Tn5-transposome containing Illumina P5 adapter and an 8-nt i5 barcode followed by a 10-nt UMI (one example AATGATACGGCGACCACCGAGATCTACACGTAAGGAGNNNNNNNNNNTCGTCGGCAGCGTCAGA TGTGTATAAGAGACAG ( SEQ ID NO : 9 ) ) to an average length of 600 bp and purified using SPRI- guanidine magnetic beads.
Purified gDNA fragments were then amplified using DSO specific primers for each strand. i5 oligo sequences for tagmentation
There are 8 different i5 index options. One of them will be annealed with Tn5-A bottom (1st line) and assembled with Tn5. This will be used for tagmenting the gDNA.
GUIDE-seq-2 primer sequences Below is the common i5 primer sequence used in both (+) and (-) strand PCR. This anneals to the oligo (in above table) added to the gDNA during tagmentation. Below are the (+) and (-) strand primers with 12 different i7 indexes. One is chosen based on the availability and a need to distinguish different sgRNA pool experiments.
However, the same index was matched for (+) and (-). For example, N701_+ and N701_- will be used if amplification is performed with N701. The different strands are amplified in separate wells and after quantification, the same concentration from each PCR product is combined for sequencing.
DSO containing amplicons were then purified with SPRI beads and was eluted with IDTE. Amplicon range of 250 - 400 bp were purified by size selection gels and were quantified with KAPA library quantification kit. Equal volume of 2 nM samples from each strand were combined and then diluted to 550 pM. 15% PhiX was added and the samples were sequenced using Nextseq 2000 P3-300 cycle cartridge with paired-end, Readl: Index : Index2 : Read2 = 146 : 8 : 18 : 146. The integration of the guide library into the genome was analyzed by amplifying gDNA between the integrated U6 (or hCtRNA) promoter and the sgRNA scaffold and sequenced with single-end, Readl: Indexl = 80 : 8 using Miseq 100-cycle cartridge. Data was analyzed by a Pooled-GUIDE-seq-2 pipeline briefly explained in FIGs. 2A-2D. Integrated coverage data was further normalized by the total reads. Normalized coverage data was then used to calculate normalized activity of each Pooled-GUIDE-seq-2 identified sites. [Normalized activity] = [Readcounts] / [normalized coverage] . On- and off-target readcounts from each guide were calculated and were used to determine the specificity to the on-target of each guide. [Specificity] = [on-target readcount] / [total readcounts] . For each guide, the number of Pooled- GUIDE-seq-2 identified sites were determined by the number of different locations a single guide edited in the genome.
Example 2: Pooled GUIDE-seq-2 improvements New GUlDE-seq-2 i7 (+) primer sequences
The i7_+ primer (i7 primer targeting the plus strand) sequence was elongated 3 nucleotides to match the melting temperature Tm of the i7_- primers (FIG. 12A), which increased the amplification efficiency (FIG. 12B) and the balance of the final GUIDE-seq-2 mapped regions (FIG. 12C). Using the new primers, the GUIDE-seq-2 identified sites were reproducible (FIG. 12D) and it correlated well to previous results (FIG. 12E). The new sequence provided better work-flow productivity and sequencing quality.
* indicates a phosphorothioate bond
Updated stringency in associating a gRNA with a DSO integration
To more accurately associate the gRNA with the DSO integration site, the following were implemented the following in the pipeline.
• Mis -priming events were removed by detecting dsODNs that did not include the full dsODN sequence. For example, if dsODN sequence was GTTTAATTGAGTTGTCATATGTTAATAACGGTAT (SEQ ID NO: 2), the (+) primers anneal to 5’-ATACCGTTATTAACATATGACA -3 ’(SEQ ID NO: 61) (reverse complement), so for the (+) strand, if the sequencing result did not have the rest of the dsODN sequence 5’-ACTCAATTAAAC-3’ (SEQ ID NO: 62) after the (+) primer sequence, it would be considered as mispriming of the primer and will be removed. It is the similar for the (-) primer, as it anneals to 5’- GTTTAATTGAGTTGTCATATGTTAATAAC-3’ (SEQ ID NO: 63) the requirement is to have 5’-GGTAT-3’ sequence.
• Filtered with Cas9 cutting distance: requires dsODN integration site to be within the Cas9 cut range (6 nucleotides in this example) from the PAM site.
• Added filter for read distribution: requires read support from two molecularly distinct read classes (i.e., support for forward and reverse reads (bidirectional) or in reads from each i7 primer (both primers)).
Increased stringency in lentivirus MOI
To ensure single guide integration into one cell, the stringency of transduction efficiency was increased during the lentiviral guide pool transduction from 80% to 30% (equivalent to MOI=0.35, formula below). This results in approximately 25% of cells receiving a single viral particle and almost 5% receiving multiple viral particles. This increased stringency was important for clinical guide nomination. X = MOI
T = transduction efficiency measured by puromycin selection k = number of viral particles per cell <formula> = -ln(l-T)
P(k) = e-XXk/k!
Example 3: Exemplary Pooled-GUIDE-seq-2 improvements
Several Pooled-GUIDE-seq-2 experiments were performed using the “Updated stringency in associating a gRNA with a DSO integration” described in Example 2. First, the number of pooled guides was increased to almost 10-times from the proof of concept study (FIG. 13 A, 4,501 guides). Second, two libraries were made with different strategies to include guides that did not start with a “G”. One was switching the promoter from U6 to tRNA (FIG. 13B, 4,435 guides), and third another was a widely used strategy by adding a “G” in front of the guide sequences that did not start with a “G” while keeping the U6 promoter for expression (FIG. 13C, 552 guides). These results show the versatility and scalability of the Pooled-GUIDE- seq-2 approach for measuring the activity and/or specificity of thousands of guide RNAs in a single experiment.

Claims

CLAIMS What is claimed is:
1. A method for identifying one or more highly active and/or highly specific guide RNA (gRNA), comprising:
(a) transfecting cells comprising an RNA-guided nuclease (RGN) and a double- stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO;
(b) detecting genomic loci of the cells that comprise the integrated DSO; and
(c) identifying one or more highly active and/or highly specific gRNA from the gRNA library using the genomic loci detected in (b).
2. The method of claim 1, wherein an edit distance between any two different gRNAs is greater than a fixed edit distance.
3. The method of claim 2, wherein the fixed edit distance is 5, 6, or 7.
4. The method of claim any one of the preceding claims, wherein genomic target regions in the cells are separated from each other by at least 30 nucleotide base pairs.
5. The method of claim any one of the preceding claims, wherein the gRNA library comprises at least 100, at least 1000, or at least 5000 different gRNAs, optionally all targeting different sequences within a gene.
6. The method of any one of the preceding claims, wherein detecting of the genomic loci is performed without intentionally shearing DNA of the cells.
7. The method of any one of the preceding claims, wherein detecting of the genomic loci comprises extracting DNA from the cells and processing the DNA using DNA tagmentation.
8. The method of any one of the preceding claims, wherein detecting of the genomic loci of the cells comprises amplifying and/or sequencing DNA of the cells, optionally using next generation sequence (NGS) technology.
9. The method of claim 8, wherein amplifying the DNA of the cells comprising use of a pair of primers respectively complementary to a region of the DSO that is at least three nucleotides from the 3’ and 5’ terminus of the DSO.
10. The method of claim 8 or 9, wherein amplifying the DNA of the cells comprises performing a polymerase chain reaction (PCR), optionally no more than a single PCR.
11. The method of any one of the preceding claims, wherein the gRNA library comprises a lentiviral library encoding the different gRNAs, optionally wherein the different gRNAs are encoded on minus strands of lentiviruses of the lentiviral library, and optionally wherein the cells are transduced at multiplicity of infection (MOI) that is less than or equal to 1.
12. The method of any one of the preceding claims, wherein the RGN is a Cas nuclease, optionally a Cas9 nuclease.
13. The method of claim 12, wherein the Cas nuclease comprises multiple nuclear localization sequences, optionally three nuclear localization sequences.
14. The method of any one of the preceding claims, wherein transfecting the cells with the DSOs comprises transfecting the cells with about 50 pmol to about 500 pmol of DSO per 0.5 million cells to about 1 million cells, optionally about 100-200 pmol of DSO per 0.6 million cells.
15. The method of any one of the preceding claims, wherein each of at least a subset of the different gRNAs of the gRNA library is operably linked to a U6 promoter and/or human cysteine tRNA (hCtRNA).
16. The method of any one of the preceding claims, further comprising associating one or more of the different gRNAs of the gRNA library with one or more genomic loci comprising a DSO integration.
17. The method of claim 16, wherein associating one or more of the different gRNAs of the gRNA library comprises, for each different gRNA, associating the different gRNA with one or more genomic loci comprising an integrated DSO if the homology region of the different gRNA has the highest nucleotide percent identity to the one or more genomic loci compared to the homology regions of the other different gRNAs of the gRNA library.
18. The method of claim 17, wherein determining highest nucleotide percent identity comprises determining highest nucleotide percent identity to a region of the one or more genomic loci that is within 25 base pairs of the integrated DSO.
19. The method of any one of the preceding claims, wherein identifying a highly active and/or highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest nucleotide percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off-target genomic loci; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
20. The method of claim 19, wherein determining the number of on-target genomic loci comprises determining a relative number using (i) a number of sequencing reads covering the on-target genomic loci comprising the DSO integration and (ii) a frequency of the gRNA in the gRNA library.
21. The method of claim 19 or claim 20, wherein determining the relative amount of on target and off-target genomic loci comprises determining using (i) a number of sequencing reads covering the on-target genomic loci comprising the DSO integration, and (ii) a number of sequencing reads covering the off-target genomic loci comprising the DSO integration.
22. The method of any one of claims 19-21, wherein the first threshold value is the 90th percentile of relative numbers of on-target genomic locus associated with each different guide RNA of the gRNA library.
23. The method of any one of claims 19-22, wherein the second threshold value is the 90th percentile of relative amounts of on target and off-target genomic loci for each guide RNA in the gRNA library.
24. The method of claim 16 or claim 17, wherein associating one or more of the different gRNAs of the gRNA library comprises associating using single cell sequencing, wherein the genomic locus comprising a DSO integration with the highest percent identity to the gRNA in a given cell is an on-target genomic locus, and wherein other genomic loci comprising a DSO integration in the given cell are off-target genomic loci.
25. The method of claim 24, wherein identifying a highly active and/or highly specific gRNA from the gRNA library comprises: determining (i) the number of cells comprising an integrated DSO in the on-target genomic locus and (ii) the number of cells comprising an integrated DSO in the off-target genomic loci; and identifying the gRNA as a highly active gRNA if the number of cells comprising the integrated DSO in the on-target genomic locus is higher than a first threshold value, and/or identifying the gRNA as a highly specific gRNA if the number of off-target genomic loci in the cells is lower than a second threshold value.
26. A method for identifying a guide RNA (gRNA) for editing a gene of interest, the method comprising:
(a) transfecting cells comprising an RNA-guided nuclease (RGN) and a double- stranded oligodeoxynucleotide (DSO) with a gRNA library comprising different gRNAs targeting different genomic regions of a gene of interest to produce cells comprising genomic DSBs tagged by an integrated DSO in the gene of interest;
(b) detecting genomic loci in the gene of interest that comprise the integrated DSO; and
(c) identifying from the gRNA library a gRNA that is highly active and/or highly specific for the gene of interest using the genomic loci detected in (b).
27. A method for identifying a Cas enzyme variant gene editing system, the method comprising:
(a) transfecting cells with a candidate Cas enzyme variant, a double-stranded oligodeoxynucleotide (DSO), and a gRNA library comprising different gRNAs to produce cells comprising genomic DSBs tagged by an integrated DSO;
(b) detecting genomic loci of the cells that comprise the integrated DSO; and
(c) identifying from the gRNA library one or more gRNAs as highly active and/or highly specific for a genomic loci detected in (b), thereby identifying a Cas enzyme variant gene editing system comprising a Cas enzyme variant and one or more gRNAs.
28. A method for identifying a Cas enzyme ortholog gene editing system, the method comprising:
(a) transfecting first cells with a first Cas enzyme ortholog, a double-stranded oligodeoxynucleotide (DSO), and a first gRNA library comprising different gRNAs to produce first cells comprising genomic DSBs tagged by an integrated DSO;
(b) transfecting second cells with a second Cas enzyme ortholog, a double- stranded oligodeoxynucleotide (DSO), and a second gRNA library comprising different gRNAs to produce second cells comprising genomic DSBs tagged by an integrated DSO;
(c) combining, after (a) and (b), the first cells and the second cells;
(d) detecting genomic loci of the first cells that comprise the integrated DSO and detecting genomic loci of the second cells that comprise the integrated DSO; and
(e) identifying, using one or more genomic loci detected in (d), the Cas enzyme ortholog having the highest activity and/or specificity for the one or more genomic loci.
29. A method for identifying one or more highly active and/or highly specific guide RNA (gRNA), comprising: (a) transfecting cells comprising a Cas nuclease and a double- stranded oligodeoxynucleotide (DSO) with a lentiviral library encoding at least 100 different gRNAs to produce cells comprising genomic DSBs tagged by integrated DSOs, wherein an edit distance between any two different gRNAs is greater than 5;
(b) extracting DNA from the cells and processing the DNA using DNA tagmentation to produce a DNA library;
(c) amplifying DNA of the DNA library using a polymerase chain reaction (PCR), wherein the PCR comprises a pair of primers respectively complementary to a region of the DSO that is at least three nucleotides from the 3’ and 5’ terminus of the DSO;
(d) sequencing amplified DNA of (c) to detect genomic loci of the cells that comprise the integrated DSOs; and
(e) associating the at least 100 different gRNAs with genomic loci; and
(f) identifying one or more highly active and/or highly specific gRNA from the gRNA library.
30. The method of claim 29, wherein identifying a highly active and/or highly specific gRNA from the gRNA library comprises: associating a gRNA of the gRNA library with one or more genomic loci comprising a DSO integration, wherein the genomic locus with the highest percent identity to the gRNA is an on-target genomic locus, and wherein other genomic loci of the associated genomic loci are off- target genomic loci; determining (i) a number of on-target genomic locus comprising the DSO integration in the cells, and (ii) a number of off-target genomic loci comprising the DSO integration in the cells; and identifying the gRNA as a highly active gRNA if the number of on-target genomic locus in the cells is higher than a first threshold value, and/or identifying the gRNA as highly specific gRNA if a relative amount of on target and off-target genomic loci in the cells is higher than a second threshold value.
31. The method of claim 29 or 30, wherein genomic target regions in the cells are separated from each other by at least 30 nucleotide base pairs.
32. The method of any one of claims 29-31, wherein sequencing of the amplified DNA is performed using next generation sequencing.
33. The method of any one of claims 29-32, wherein no more than a single PCR is performed.
34. The method of claim 33, wherein the Cas nuclease is a Cas9 nuclease, optionally comprising multiple nuclear localization sequences, optionally three nuclear localization sequences.
35. The method of any one of claims 29-34, wherein transfecting the cells with the DSO comprises transfecting the cells with about 50 pmol to about 500 pmol of DSO per 0.5 million cells to about 1 million cells, optionally about 100-200 pmol of DSO per 0.6 million cells.
36. The method of any one of claims 29-35, wherein each of at least a subset of the different gRNAs of the library is operably linked to a U6 promoter and/or human cysteine tRNA (hCtRNA).
37. The method of any one of the preceding claims, wherein the cells comprise mammalian cells.
38. The method of claim 37, wherein the mammalian cells comprise human cells.
PCT/US2025/033041 2024-06-11 2025-06-10 High-throughput unbiased identification of double-stranded dna breaks Pending WO2025259693A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US202463658667P 2024-06-11 2024-06-11
US63/658,667 2024-06-11

Publications (1)

Publication Number Publication Date
WO2025259693A1 true WO2025259693A1 (en) 2025-12-18

Family

ID=98051532

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/US2025/033041 Pending WO2025259693A1 (en) 2024-06-11 2025-06-10 High-throughput unbiased identification of double-stranded dna breaks

Country Status (1)

Country Link
WO (1) WO2025259693A1 (en)

Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200270632A1 (en) * 2017-09-15 2020-08-27 The Board Of Trustees Of The Leland Stanford Junior University Multiplex production and barcoding of genetically engineered cells

Patent Citations (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200270632A1 (en) * 2017-09-15 2020-08-27 The Board Of Trustees Of The Leland Stanford Junior University Multiplex production and barcoding of genetically engineered cells

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
SCHMIERER BERNHARD, BOTLA SANDEEP K, ZHANG JILIN, TURUNEN MIKKO, KIVIOJA TEEMU, TAIPALE JUSSI: "CRISPR/Cas9 screening using unique molecular identifiers", MOLECULAR SYSTEMS BIOLOGY, vol. 13, no. 10, 1 October 2017 (2017-10-01), GB , pages 945, XP055886148, ISSN: 1744-4292, DOI: 10.15252/msb.20177834 *

Similar Documents

Publication Publication Date Title
Matreyek et al. An improved platform for functional assessment of large protein libraries in mammalian cells
US20230193381A1 (en) Compositions and methods for accurately identifying mutations
US11913017B2 (en) Efficient genetic screening method
US20220267759A1 (en) Methods and compositions for scalable pooled rna screens with single cell chromatin accessibility profiling
KR102598819B1 (en) Genomewide unbiased identification of dsbs evaluated by sequencing (guide-seq)
US20240200130A1 (en) Highly sensitive in vitro assays to define substrate preferences and sites of nucleic-acid binding, modifying, and cleaving agents
EP3642334A1 (en) Nucleic acid-guided nucleases
JP2018532419A (en) CRISPR-Cas sgRNA library
JP2024522171A (en) CRISPR-Transposon System for DNA Modification
Talkish et al. Rapidly evolving protointrons in Saccharomyces genomes revealed by a hungry spliceosome
JP2026048971A (en) A method for specifying nucleazeon/off-target editing positions, called &#34;CTL-seq&#34; (CRISPR Tag Linear-seq).
WO2007142608A1 (en) Nucleic acid concatenation
US20250243511A1 (en) Crispr-associated transposon systems and methods of using same
WO2024119461A1 (en) Compositions and methods for detecting target cleavage sites of crispr/cas nucleases and dna translocation
US20250354297A1 (en) Shielded small nucleotides for intracellular barcoding
US20240209447A1 (en) Compressive molecular probes for genomic editing and tracking
WO2023060539A1 (en) Compositions and methods for detecting target cleavage sites of crispr/cas nucleases and dna translocation
Salnikov et al. Here and there: the double-side transgene localization
WO2021058145A1 (en) Phage t7 promoters for boosting in vitro transcription
US20230287396A1 (en) Methods and compositions of nucleic acid enrichment
US20260117280A1 (en) Targeted genomic sequencing in single cells
US20250320483A1 (en) Systems and methods for gene insertions
WO2026064531A1 (en) Single-cell edit capture sequencing
WO2025101943A1 (en) Nucleic acid-guided dna synthesis
Byrne Building a Better Transcriptome

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25822078

Country of ref document: EP

Kind code of ref document: A1