WO2024197857A1 - 一种筛选向导rna的方法 - Google Patents
一种筛选向导rna的方法 Download PDFInfo
- Publication number
- WO2024197857A1 WO2024197857A1 PCT/CN2023/085605 CN2023085605W WO2024197857A1 WO 2024197857 A1 WO2024197857 A1 WO 2024197857A1 CN 2023085605 W CN2023085605 W CN 2023085605W WO 2024197857 A1 WO2024197857 A1 WO 2024197857A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- rna
- editing
- guide rna
- gene
- vector
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- A—HUMAN NECESSITIES
- A61—MEDICAL OR VETERINARY SCIENCE; HYGIENE
- A61K—PREPARATIONS FOR MEDICAL, DENTAL OR TOILETRY PURPOSES
- A61K48/00—Medicinal preparations containing genetic material which is inserted into cells of the living body to treat genetic diseases; Gene therapy
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/10—Processes for the isolation, preparation or purification of DNA or RNA
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/11—DNA or RNA fragments; Modified forms thereof; Non-coding nucleic acids having a biological activity
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/11—DNA or RNA fragments; Modified forms thereof; Non-coding nucleic acids having a biological activity
- C12N15/113—Non-coding nucleic acids modulating the expression of genes, e.g. antisense oligonucleotides; Antisense DNA or RNA; Triplex- forming oligonucleotides; Catalytic nucleic acids, e.g. ribozymes; Nucleic acids used in co-suppression or gene silencing
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
Definitions
- the invention belongs to the field of biomedicine and relates to a method for screening guide RNA.
- CRISPR Clustered regularly interspaced short palindromic repeats
- the system requires the simultaneous expression of Cas protein and guide RNA (gRNA) and the provision of homologous repair DNA sequences.
- gRNA guide RNA
- CRISPR/Cas protein cuts genomic DNA to cause DSB (double-strand break), and off-target effects can cause genomic instability.
- the CRISPR/Cas system may also trigger cellular innate immune responses. These have greatly limited its clinical application.
- A-to-I editing at the RNA level occurs under the mediation of ADAR family proteins.
- the I base will be recognized as the G base during translation, so the G-to-A mutation in the genome can be repaired at the RNA level through ADAR editing.
- ADAR family proteins There are three main subtypes of ADAR family proteins, ADAR1, ADAR2 and ADAR3, of which ADAR3 does not have deaminase activity. Editing events are mainly completed by ADAR1 and ADAR2, and the broad spectrum expression of ADAR1 protein in human tissues can ensure its application in a variety of tissues or organs.
- A-to-I editing of specific sites at the RNA level has been developed using overexpressed ADAR family proteins or endogenously expressed ADAR family proteins combined with guide RNA.
- ADAR family proteins prefer to edit double-stranded RNA, but there are also multiple mutations in the editing region to form a certain secondary structure, which leads to the need to screen gRNA when designing gRNA, and no more efficient screening methods have been reported.
- the present invention provides a method for screening guide RNA for editing RNA, comprising the following steps: (1) each guide RNA to be screened is connected to a linker and is connected to the target RNA on a nucleic acid chain that can form the same transcript, thereby forming a plurality of the nucleic acid chains; (2) the nucleic acid chains are transferred into cells expressing ADAR proteins; and (3) the RNA in the cells is extracted and measured to determine the editing level of the target RNA, and the editing level of the corresponding guide RNA is determined by the barcode sequence.
- the linker is located between the target RNA and the guide RNA.
- the barcode sequence is contained in the linker.
- the nucleic acid chain is connected with a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, and a guide RNA nucleic acid fragment in sequence from upstream to downstream.
- the linker sequence acts as a tag in the present invention.
- the linker can preferably be located between the target RNA and the guide RNA, in which case the linker also acts as a Separating the target RNA and guide RNA to a certain extent is conducive to the folding of the two into secondary structures.
- this connection method is also conducive to application in second-generation sequencing, thereby achieving efficient and high-throughput detection.
- the sequencing can be cut into two parts: target RNA + linker and linker + guide RNA. Through the labeling of the linker, the editing level of the target RNA under the action of the corresponding guide RNA can be obtained.
- the transcript is linked to a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, a guide RNA nucleic acid fragment and a polyA tail nucleic acid fragment in sequence from upstream to downstream.
- the reporter gene is selected from the group consisting of yfp gene, gus gene, rfp gene, gfp gene, cfp gene, bfp gene or kanamycin resistance gene, but is not limited thereto.
- the length of the linker is 26 nt-2121 nt.
- the linker includes SEQ ID NO:549 or SEQ ID NO:557.
- a long linker sequence is added to the screening method in order to simulate the formation of double strands between intermolecular RNAs for editing.
- a mixed library of gRNAs of different lengths in the case of short linkers can also be screened for gRNAs with different editing efficiencies using the methods of the present invention.
- step (3) includes establishing a targeted sequencing library for next-generation sequencing.
- contacting the nucleic acid strand with the ADAR protein is achieved by introducing the nucleic acid strand into a cell expressing the ADAR protein.
- the cell expresses an ADAR protein.
- the ADAR protein comprises ADAR1 or ADAR2.
- the guide RNA is capable of recruiting ADAR protein in the cell to edit the target RNA.
- the cell overexpresses an ADAR protein.
- the editing is A-to-I editing.
- the guide RNA is a guide RNA extended upstream and downstream with base A as the center.
- the guide RNA is a guide RNA with a length of 30nt to 151nt obtained by extending 15nt to 75nt upstream and downstream from the A base.
- the guide RNA is a 151nt guide RNA that is extended 75nt upstream and downstream from the A base.
- the guide RNA is a guide RNA of 28 nt to 89 nt extending upstream and downstream from the A base.
- the screening method comprises adding a 20-200 nt sequence upstream and downstream of the editing site to the downstream of the reporter gene and inserting it into a vector to obtain a reporter gene vector, and then transferring the reporter gene vector into cells.
- the screening method is a high throughput screening method.
- the present invention provides a method for high-throughput screening of guide RNA (gRNA) sequences that can recruit ADAR proteins and perform A-to-I editing on specific sites, filling the gaps in the prior art.
- gRNA guide RNA
- the present invention provides a gRNA screening method established by the above method for any RNA editing system based on ADAR.
- the target RNA sequence is longer than the guide RNA sequence.
- the vector comprises a plasmid or a viral vector.
- the vector is a plasmid or viral vector for expression in higher eukaryotic cells or prokaryotic cells.
- the vector backbone is a pcDNA3.1 vector.
- the screening method comprises: connecting a plurality of different guide RNAs to the target RNA to form a transcription system, and transferring the system into cells via the vector.
- the guide RNA of interest is a guide RNA that targets a site for efficient editing.
- the method further comprises: after obtaining the editing level of the guide RNA from step (3), selecting guide RNAs with high editing levels, synthesizing single-stranded DNA with the same sequence, and then designing error-prone PCR primers to randomly introduce mutations into the targeted region of the guide RNA, and then repeating steps (1) to (3).
- the present invention provides a method for screening guide RNA for editing RNA, comprising: (a) designing a guide RNA library; (b) inserting a gene fragment to be edited and a reporter gene into a vector to obtain a reporter gene vector; (c) inserting the guide RNA library into the reporter gene vector to obtain a plasmid library to be screened; (d) transferring the plasmid library to be screened into cells, extracting RNA from the cells after transfection, and sequencing; (e) using a bioinformatics assay to obtain a plasmid library to be screened; Information analysis of the editing level of the guide RNA; wherein the order of (a) and (b) can be swapped.
- step (b) comprises: adding the sequence upstream and downstream of the editing site to the downstream of the reporter gene and inserting it into the vector to obtain a reporter gene vector. In some embodiments, step (b) comprises: adding the sequence 20-200 nt upstream and downstream of the editing site to the downstream of the reporter gene and inserting it into the vector to obtain a reporter gene vector.
- step (b) includes: amplifying the full length of the fluorescent gene and the gene fragment to be edited by PCR, adding homologous sequences for homologous recombination at both ends or introducing restriction sites, and performing NheI and HindIII double restriction digestion on the empty vector and recovering the linear vector, and inserting the multiple fragments amplified by PCR into the linear vector to obtain a reporter gene vector.
- step (c) comprises: amplifying the guide RNA in step (a) and adding homology arms or restriction sites to obtain a guide RNA sequence pool, and linearizing the reporter gene vector in step (b) by double restriction digestion with BamHI and EcoRI; then inserting the guide RNA sequence pool into the linearized reporter gene vector, and obtaining a plasmid library to be screened after transformation.
- the screening method further comprises: (f) selecting a guide RNA with high editing efficiency from the guide RNA to directly synthesize a single-stranded DNA with the same sequence, designing error-prone PCR primers to randomly introduce mutations into the guide RNA targeting region, and repeating steps (b) to (e).
- the present invention provides a method for high-throughput screening of guide RNA sequences that can recruit ADAR proteins and perform A-to-I editing on specific sites, comprising the following steps: S1. Designing and synthesizing a linker sequence for the gRNA library to be screened. Specifically, one end of the gRNA library is a fixed sequence required for chip synthesis, and the other end is a linker sequence of about 20 nt that can distinguish different libraries; S2. Design and construction of a screening vector. Specifically, the present invention designs a strategy of adding a target RNA and a gRNA sequence downstream of the green fluorescent protein GFP coding region. The three are implemented in the same transcript. The three can be directly connected or can be directly connected.
- the gRNAs are spaced by linker sequences of different lengths; S3. Construction of the gRNA plasmid library to be screened, specifically, PCR amplification of the mixed library by linker sequences corresponding to different libraries and further insertion into the vector system of S2; S4. Transfection of the plasmid library into specific cells by transfection; S5. RNA extraction and construction of targeted sequencing library and next-generation sequencing; S6. Bioinformatics analysis of the editing efficiency of each gRNA corresponding to the target RNA; S7. Select gRNA sequences with high editing levels, synthesize single-stranded DNA with the same sequence, perform error-prone PCR to construct the plasmid library and perform the second round of screening according to the methods of S3 to S6.
- the present invention provides a method for high-throughput screening of guide RNA sequences that can recruit ADAR proteins and perform A-to-I editing on specific sites.
- the present invention fills the gap in high-throughput screening of gRNA sequences for target RNA editing based on ADAR by developing experimental and bioinformatics methods, and completes the screening through a series of preferred experimental design methods.
- the method described in the present invention includes 1) designing the target RNA and gRNA to the same transcript by simulating different regions of the gene.
- the introduction of the fluorescent gene facilitates the observation of the transfection efficiency and can be replaced by other fluorescent genes; 2) constructing a targeted sequencing library and an analysis method that can accurately match each gRNA with its editing efficiency; 3) randomly introducing mutations based on the screened gRNA sequences with high editing levels to further screen out gRNAs with higher editing levels.
- the present invention provides a guide RNA obtained by any of the screening methods.
- the present invention provides a guide RNA, the sequence of which is as shown in any one of SEQ ID NO: 557, SEQ ID NO: 559, SEQ ID NO: 561, SEQ ID NO: 563, SEQ ID NO: 565, SEQ ID NO: 567, SEQ ID NO: 569, SEQ ID NO: 571, SEQ ID NO: 573, SEQ ID NO: 575, SEQ ID NO: 577, SEQ ID NO: 579, SEQ ID NO: 583 to SEQ ID NO: 602, SEQ ID NO: 603 to SEQ ID NO: 622, or SEQ ID NO: 623 to SEQ ID NO: 642.
- the present invention provides a method for obtaining the guide RNA or the use of the obtained guide RNA in preparing a drug for treating a disease.
- the drug is used to site-specifically edit nucleotides in a target gene in a eukaryotic cell to achieve the purpose of treating a disease.
- the disease is a disease caused by a mutation in the GAPDH gene, ATP7B gene, or FGFR gene.
- the GAPDH gene specific site motif is AAC, AAU, AAG, UAA or UAG, but is not limited thereto.
- the ATP7B gene specific site motif is UAG, but is not limited thereto.
- the FGFR gene specific site motif is CAG, but is not limited thereto.
- the disease includes, but is not limited to, Wilson's disease, achondroplasia, thanatoskeletal dysplasia, or dwarfism.
- the use is through the action of an RNA editing entity naturally present in the cell and capable of editing the nucleotide sequence. use.
- the present invention provides an engineered RNA, the structure of which includes: a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, and a guide RNA nucleic acid fragment are sequentially connected from upstream to downstream.
- the present invention provides an engineered RNA, whose structure includes: a reporter gene nucleic acid fragment, a target RNA nucleic acid fragment, a linker nucleic acid fragment, a guide RNA nucleic acid fragment and a polyA tail nucleic acid fragment are sequentially connected from upstream to downstream.
- the reporter gene is selected from the group consisting of yfp gene, gus gene, rfp gene, gfp gene, cfp gene, bfp gene, or kanamycin resistance gene.
- the linker has a length of 26 nt to 2121 nt.
- the linker includes SEQ ID NO:549 or SEQ ID NO:557.
- the guide RNA is a guide RNA extended upstream and downstream with base A as the center.
- the guide RNA is a 151nt guide RNA that is extended 75nt upstream and downstream from the A base.
- the present invention provides an engineered RNA structure library comprising the plurality of different engineered RNAs.
- the different engineered RNAs are obtained by connecting the target RNA to different guide RNAs via different linkers.
- the target RNA nucleic acid fragments contained in the engineered RNA in the library are the same, and the guide RNA nucleic acid fragment sequences are different.
- the present invention provides a vector comprising the engineered RNA or the engineered RNA structure library.
- the vector is obtained by adding the target RNA sequence downstream of the reporter gene sequence and inserting it into the vector backbone to obtain a reporter gene vector, and then inserting the amplified guide RNA library into the reporter gene vector.
- the present invention provides a cell comprising the engineered RNA or the engineered RNA structure library or the vector.
- FIG1 is a schematic diagram of a high-throughput screening method.
- Figure 1 shows the overall process of high-throughput screening of guide RNA, including the design of sequences to be screened; construction of the synthesized sequences to be screened into screening vectors; introduction into cells for editing; RNA extraction, construction of second-generation sequencing libraries and high-throughput sequencing; analysis of sequencing data using bioinformatics methods to obtain the editing efficiency of each gRNA; selection of gRNAs with high editing efficiency for verification; selection of multiple gRNAs with high editing efficiency for random introduction of error prone mutations and repeating the screening process to obtain gRNAs with higher editing efficiency.
- FIG. 2 is a schematic diagram of the transcript information of the screening vector.
- Figure 2 shows the specific transcript composition information of the screening vector.
- GFP GFP
- the transcript simulates the expression of normal genes
- GFP is the coding region
- the target sequence to be edited and the linker and gRNA simulate the 3'UTR region, where there will be a linker sequence of a preset length in the target RNA (Reporter) and gRNA, which can separate the two to simulate the formation of a secondary structure between molecules, and on the other hand, adding a Barcode (used to distinguish different gRNAs) sequence in the Linker region can distinguish different gRNAs and corresponding editing levels in the subsequent analysis of high-throughput sequencing data.
- a Barcode used to distinguish different gRNAs
- Figure 3 Schematic diagram of screening vector transcripts for RNA editing.
- Figure 3 shows that after the screening vector expresses RNA, it will self-fold to form a double-stranded RNA region of Reporter and gRNA, which will be recognized and bound by the ADAR protein and edit the target site.
- Figure 4 shows the results of gRNA screening at different sites of the GAPDH gene.
- Figure 4a The result of editing site 1 with an AAC motif.
- Figure 4b The result of editing site 2 with an AAU motif.
- Figure 4c The result of editing site 3 with an AAG motif.
- Figure 4d The result of editing site 4 with an UAA motif.
- Figure 4 shows the editing effect of each gRNA obtained by high-throughput screening of four editing sites with different motifs selected on the GAPDH gene, where the horizontal and vertical axes represent two biological replicates.
- the black dots in the figure are gRNAs with C bases at the editing site and complementary pairs at other positions, and the non-black dots are sequences obtained by introducing different mutations on the basis of the sequence by other gRNAs.
- the editing level corresponding to each site is obtained by high-throughput sequencing.
- Figure 5 is a comparison of gRNA editing and off-target editing at different sites of the GAPDH gene.
- Figure 5 shows a heat map comparison of the editing of the target site and the off-target editing of the upstream and downstream of the gRNA at different sites of the GAPDH gene.
- the horizontal axis 0 indicates that the target site is edited
- the vertical axis represents the editing level of each gRNA at different A bases, and they are arranged downward according to the highest editing sequence of the target site. Blue to red indicates the changing trend of the editing level of the A base site from low to high.
- FIG6 shows the editing and off-target editing efficiencies of gRNAs of different lengths targeting site 787 of the human GAPDH gene.
- Figure 6 shows the heat map of the on-target editing and upstream and downstream off-target editing of all gRNAs at position 787 of the GAPDH gene.
- the horizontal axis between -1 and 1 is the target site, -1 to -25 is the interval upstream of the editing site, 1 to 25 is the interval downstream of the editing site, and the vertical axis is the editing level of each gRNA at different A bases, arranged downward according to the highest editing sequence of the target site, and blue to red indicates the trend of the editing level of the A base site from low to high.
- Figure 7 shows that the gRNA screened by high-throughput screening was verified by low-throughput method.
- the horizontal axis represents the editing level of each gRNA measured in the high-throughput screening method
- the vertical axis represents the editing level measured by synthesizing the modified RNA of each gRNA separately and transfecting it into cells.
- Figure 8 shows the original results of first-generation sequencing of the editing level of a single gRNA on the GAPDH gene.
- Figure 8 shows the original peak graph of the editing level obtained by first-generation sequencing after a single modified gRNA was transfected into cells.
- different bases correspond to peaks of different colors
- a base corresponds to green peaks
- G base corresponds to black peaks.
- peaks of different colors appear at a site, it means that there are different base compositions at the site.
- the calculation method of the editing level is the peak area of G base divided by the sum of the peak areas of A base and G base at the site.
- FIG. 9 shows the gRNA screening results for the human ATP7B gene.
- Figure 9 shows the editing effect of each gRNA obtained by high-throughput screening of the editing sites selected on the ATP7B gene, where the horizontal and vertical axes represent two biological replicates.
- the black dots in the figure are gRNAs with C bases at the editing site and complementary pairs at other positions, and the non-black dots are sequences obtained by introducing different mutations on the basis of this sequence by other gRNAs.
- the editing level corresponding to each point is obtained by high-throughput sequencing.
- len41 indicates the gRNA with different mutations designed corresponding to the target editing site and the length of 41nt upstream and downstream
- len31 indicates the gRNA with different mutations designed corresponding to the length of 31nt upstream and downstream of the target editing site.
- Figure 10 shows the gRNA targeted editing and off-target editing efficiency of the human ATP7B gene.
- Figure 10 shows the heat map of all gRNAs editing the target site and the upstream and downstream off-target editing of the ATP7B gene.
- the horizontal axis 0 is the target site, 0 to -20 is the interval upstream of the editing site, 0 to 20 is the interval downstream of the editing site, and the vertical axis is the editing level of each gRNA at different A bases, arranged downward according to the highest editing sequence of the target site, and blue to red indicates the trend of the editing level of the A base site from low to high.
- FIG. 11 shows the second round screening results of gRNAs with the UAG motif of the ATP7B gene.
- Figure 11 shows that the top 20 sequences with high editing efficiency of different lengths obtained by the first round of screening of the ATP7B gene were subjected to error prone introduction of random mutations and the second round of screening. After the screening vector was transfected into the cells, the newly generated gRNA and its corresponding editing efficiency were measured by high-throughput sequencing.
- the horizontal and vertical axes in the figure represent two biological replicates.
- len41 represents the result of the second round of screening by selecting the gRNA of the 41nt length group
- len31 represents the result of the second round of screening by selecting the gRNA of the 31nt length group.
- Figure 12 Results of gRNA screening for human FGFR gene.
- Figure 12 shows the editing effect of each gRNA obtained by high-throughput screening of the editing sites selected on the FGFR gene, where the horizontal and vertical axes represent two biological replicates.
- the black dots in the figure are gRNAs with C bases at the editing site and complementary pairs at other positions, and the non-black dots are sequences obtained by introducing different mutations on the basis of this sequence by other gRNAs.
- the editing level corresponding to each point is obtained by high-throughput sequencing.
- len41 indicates the gRNA with different mutations designed corresponding to the target editing site and the length of 41nt upstream and downstream
- len31 indicates the gRNA with different mutations designed corresponding to the length of 31nt upstream and downstream of the target editing site.
- Figure 13 shows the gRNA targeted editing and off-target editing efficiency of the human FGFR gene.
- Figure 13 shows a heat map of all gRNAs editing the FGFR gene at the target site and off-target editing upstream and downstream.
- the horizontal axis 0 is the target site, 0 to -20 is the interval upstream of the editing site, 0 to 20 is the interval downstream of the editing site, and the vertical axis is the editing level of each gRNA at different A bases, arranged downward according to the highest editing sequence of the target site. Blue to red indicates the trend of the editing level of the A base site from low to high.
- FIG. 14 shows the second round screening results of gRNAs with the CAG motif of human FGFR gene.
- Figure 14 shows that the top 20 sequences with high editing efficiency of 41nt length obtained by the first round of screening of FGFR gene were subjected to error prone introduction of random mutations and the second round of screening. After the screening vector was transfected into cells, the newly generated gRNA and its corresponding editing efficiency were measured by high-throughput sequencing method.
- the horizontal and vertical axes in the figure represent two biological replicates.
- len41 represents the result of the second round of screening by selecting the gRNA of 41nt length group.
- Engineered polynucleotides or “engineered guide RNAs” can be used interchangeably with circular guide RNAs.
- Engineered polynucleotides can include recombinant polynucleotides of DNA or RNA or hybrid DNA/RNA constructs.
- Engineered polynucleotides can produce guide RNAs, more specifically circular guide RNAs.
- mutation may refer to a change in a nucleic acid sequence encoding a protein relative to the consensus sequence of the protein.
- a “missense” mutation causes a codon to be replaced by another codon; a “nonsense” mutation changes a codon from a codon encoding a specific amino acid to a stop codon.
- Nonsense mutations generally result in truncated translation of a protein.
- a “silent” mutation is a mutation that has no effect on the resulting protein.
- point mutation may refer to a mutation that affects only one nucleotide in a gene sequence.
- “Splice site mutations” are those mutations present in pre-mRNA (before processing to remove introns) that cause mistranslation of a protein due to incorrect depiction of a splice site and generally result in truncation. Mutations may include single nucleotide variations (SNVs). Mutations may include sequence variants, sequence variations, sequence changes, or allelic variants. Reference DNA sequences may be obtained from reference databases. Mutations may affect function. Mutations may not affect function. Mutations may occur in one or more nucleotides at the DNA level, in one or more nucleotides at the ribonucleic acid (RNA) level, in one or more amino acids at the protein level, or any combination thereof.
- SNVs single nucleotide variations
- Reference DNA sequences may be obtained from reference databases. Mutations may affect function. Mutations may not affect function. Mutations may occur in one or more nucleotides at the DNA level, in one or more nucleotides at the ribonucleic acid (RNA
- Reference sequences can be obtained from databases such as the NCBI Reference Sequence Database (RefSEQ ID NO:) database.
- Specific changes that may constitute mutations may include substitutions, deletions, insertions, inversions or transitions in one or more nucleotides or one or more amino acids. Mutations may be point mutations.
- RNA editing refers to the co-transcriptional or post-transcriptional modification process that introduces changes in the sequence of genomic encoded RNA, resulting in RNA mutations.
- Editing of adenosine to inosine (A-to-I) in double-stranded RNA (dsRNA), catalyzed by the adenosine deaminase acting on RNA (ADAR) family of enzymes is a common type of RNA editing in mammals.
- dsRNA double-stranded RNA
- ADAR adenosine deaminase acting on RNA
- ADAR1 and ADAR2 ADARs catalyze all currently known A-to-I editing sites.
- ADAR3 has no known deaminase activity. Inosine (I) mimics guanosine (G), so ADAR proteins introduce virtual A to G substitutions in transcripts. This change can lead to specific amino acid substitutions, alternative splicing, microRNA-mediated gene silencing, or changes in transcript localization and stability.
- RNA sequencing methods known in the art can be used to detect modifications of RNA sequences.
- RNA editing can cause changes in the amino acid sequence encoded by the mRNA. For example, the point Mutations are introduced into mRNA, or innate or acquired point mutations in mRNA are restored to produce wild-type gene products because "A" is converted to "I".
- Amino acid sequencing by methods known in the art can be used to find any changes in amino acid residues in the encoded protein.
- Modification of stop codons can be determined by assessing the presence of functional, elongated, truncated, full-length and/or wild-type proteins. For example, when the target adenosine is located in the UGA, UAG or UAA stop codon, modification of the target A (UGA or UAG) or multiple A (UAA) can create a read-through mutation and/or an extended protein, or a truncated protein encoded by the target RNA can be restored to produce a functional, full-length and/or wild-type protein.
- Editing of the target RNA can also generate abnormal splicing sites and/or alternative splicing sites in the target RNA, resulting in extended, truncated or misfolded proteins, or the encoded abnormal splicing or alternative splicing sites in the target RNA are restored to produce functional, correctly folded, full-length and/or wild-type proteins.
- the present application contemplates editing of innate and acquired genetic changes, such as missense mutations, premature stop codons, abnormal splicing, or alternative splicing sites encoded by target RNA.
- the function of the protein encoded by the target RNA is evaluated using known methods, and it can be found whether RNA editing has achieved the desired effect.
- the identification of deamination to inosine can provide an assessment of whether a functional protein exists, or whether the RNA associated with disease or drug resistance caused by the mutated adenosine has been restored or partially restored.
- the deamination of adenosine (A) to inosine (I) may introduce point mutations in the resulting protein, the identification of deamination to inosine can find out the functional indications of disease causes or disease-related factors.
- the target RNA is a regulatory RNA.
- the target RNA to be edited is a ribosomal RNA, a transfer RNA, a long non-coding RNA or a small RNA (e.g., miRNA, pri-miRNA, pre-miRNA, piRNA, siRNA, snoRNA, snRNA, exRNA or scaRNA).
- the deamination of the target adenosine includes, for example, ribosomal RNA, transfer RNA, a long non-coding RNA or a small RNA (e.g., miRNA), including changes in three-dimensional structure and/or loss of function or gain of function.
- the deamination of target A in the target RNA changes the expression level of one or more downstream molecules (e.g., proteins, RNA and/or metabolites) of the target RNA.
- the change in the expression level of the downstream molecule can be an increase or decrease in the expression level.
- RNA editing entity refers to a biological molecule that can cause chemical modification of nucleotides to change nucleotides into different nucleotides.
- RNA editing entities can be recruited to specific sites in polynucleotides to cause changes in the nucleic acid sequence at the desired site.
- RNA editing entities include APOBEC proteins (e.g., APOBEC1, APOBEC2, APOBEC3A, APOBEC3B, APOBEC3C, APOBEC3E, APOBEC3F, APOBEC3G, APOBEC3H, or APOBEC4 proteins) or ADAR proteins (e.g., ADAR1, ADAR2, or ADAR3 proteins).
- target RNA refers to an RNA sequence that is designed to recruit a deaminase to have complete complementarity or substantial complementarity therewith, and the target sequence and dRNA are hybridized to form a double-stranded RNA (dsRNA) region containing a target adenosine, which recruits an adenosine deaminase (ADAR) acting on RNA, which deaminates the target adenosine.
- dsRNA double-stranded RNA
- ADAR adenosine deaminase
- ADAR is naturally present in a host cell, such as a eukaryotic cell (preferably a mammalian cell, more preferably a human cell).
- the ADAR is introduced into a host cell.
- barcode sequence refers to a unique nucleotide sequence used to identify and/or track the source of a polynucleotide in a reaction.
- the size and composition of the barcode sequence can vary greatly.
- the length of the barcode sequence can range from 4 to 36 nucleotides, or 6 to 30 nucleotides, or 8 to 20 nucleotides.
- vector refers to a nucleic acid molecule capable of transporting another nucleic acid connected thereto.
- the vector includes, but is not limited to, single-stranded, double-stranded or partially double-stranded nucleic acid molecules; nucleic acid molecules (e.g., circular) containing one or more free ends, without free ends; nucleic acid molecules containing DNA, RNA or both; and other various polynucleotides known in the art.
- plasmid refers to a circular double-stranded DNA loop, in which additional DNA segments can be inserted, for example, by standard molecular cloning techniques.
- Certain vectors are capable of autonomous replication in the host cell into which they are introduced (e.g., bacterial vectors and additional mammalian vectors having a bacterial origin of replication). Other vectors (e.g., non-additional mammalian vectors) are integrated into the genome of the host cell after being introduced into the host cell, thereby replicating together with the host genome.
- certain vectors are capable of directing the transcription or expression of the encoding nucleotide sequence to which they are operably connected. Such vectors are referred to herein as "expression vectors”.
- the editing entity is usually proteinaceous in nature, such as the ADAR enzymes found in metazoans including mammals.
- the editing entity may also include a complex of a nucleic acid and a protein or peptide, such as a ribonucleoprotein.
- the editing enzyme may include only nucleic acids or consist of nucleic acids, such as a ribozyme. All of these editing entities are included in the present invention, as long as they are recruited by the oligonucleotide construct according to the present invention.
- the editing entity is an enzyme, more preferably an adenosine deaminase or a cytidine deaminase, and even more preferably an adenosine deaminase.
- Y is preferably cytidine or uridine, most preferably cytidine.
- ADAR hADAR1 and hADAR2
- any subtypes thereof such as hADAR1p110 and p150.
- the target RNA can be any cellular or viral RNA sequence, but is more typically a protein-coding pre-mRNA or mRNA.
- the CDS of human ADAR1 (the sequence of human ADAR1 CDS is shown in SEQ: ID NO: 2) was amplified by PCR and homology arms were added.
- the pcDNA3.1 vector (the sequence is shown in SEQ ID NO: 1) was linearized by double restriction enzyme digestion with NheI and EcoRI, and the inserted fragment was connected to the vector by homologous recombination.
- the full length of the fluorescent gene GFP and the gene fragment to be edited were amplified by PCR, and homologous sequences for homologous recombination were added at both ends.
- the pcDNA3.1 empty vector was double-digested with NheI and HindIII and the linear vector was recovered.
- the vector and the inserted fragment were homologously recombined through multi-fragment recombination to complete the construction of the reporter gene vector.
- the gRNA mixed sequence pool was amplified by PCR and homology arms were added at the same time.
- the reporter gene vector was linearized by double restriction digestion with BamHI and EcoRI and the gRNA sequence pool was inserted into the reporter gene vector by homologous recombination. After transformation and plasmid plasmid extraction, all single clones were collected together to finally obtain the gRNA sequence pool plasmid library.
- the required cell culture plate is set, and a certain amount of culture medium and cells are added to the corresponding culture plate according to the cell passaging method.
- the cells reach 70% confluence after 24 hours of growth, the cell transfection experiment can be carried out.
- the ADAR1 expression vector was transfected into HEK293 cells. After 48 hours, the screening drug G418 was added, and the solution was changed every two days. When there were no living cells in the non-transfected group, the screening of HEK293 cell lines stably transfected with ADAR1 was completed.
- Total cellular RNA can be extracted using a conventional RNA extraction kit to ensure that there is no obvious degradation of RNA.
- a conventional reverse transcription kit is used for reverse transcription reaction.
- sequencing adapter sequences need to be added to both ends of the amplification primers, and the number of cycles in the first round of amplification is controlled to be less than 20 cycles.
- Agarose gel electrophoresis is used to verify the first round of PCR products and recover the library fragments of the corresponding size.
- the second round of PCR reaction is performed.
- the second round of primers have sequencing adapter sequences and specific base barcode sequences.
- the second round of PCR is generally controlled within 10 cycles.
- the PCR product is subjected to agarose gel electrophoresis again to confirm the size of the library fragments and the gel is cut to recover the purified library.
- the targeted sequencing library is sequenced by second-generation sequencing.
- Agilent's second-generation error-prone PCR kit was used to amplify specific regions and randomly introduce mutations. Samples were prepared according to the instructions, with a template addition amount of 0.1 ng and 30 amplification cycles.
- the original data was processed by removing the linkers using cutadapt software, and then the original read length after removing the linkers was aligned to the reference sequence.
- the editing of the reporter gene was matched one by one with each gRNA according to the linker sequence, and then the editing level of the reporter gene was calculated to obtain the editing level of each gRNA on the target site.
- thiophosphorylation modification (hereinafter replaced by *) is selected, and in some embodiments, 2' position LNA modification (hereinafter replaced by L) is selected.
- LNA modification 2' position LNA modification
- This example takes the gRNA screening method for editing multiple A base sites of the human endogenous gene GAPDH (NM_002046.7) as an example to further illustrate the present invention. Specifically, the method comprises the following steps:
- a 151nt gRNA was designed with a specific A base as the center and extended 75nt upstream and downstream. Mutations of different numbers and positions were introduced into the gRNA to obtain a library to be screened.
- This embodiment includes A bases at four sites, and the triplet motifs at which they are located are AAC, AAU, AAG and UAA, respectively. Multiple gRNA libraries of the four sites were screened.
- This article lists the library sequence information of the AAC motifs in the four sites as an example, as shown in SEQ ID NO: 1 to SEQ ID NO: 540.
- Sequence synthesis can be distinguished by designing adapter sequences, which will not cause crossover between different groups in the downstream amplification step.
- the editing site information and adapter sequences corresponding to the four libraries are shown in Table 1 below.
- the 100bp sequence upstream and downstream of the editing site (a total of 201 bases) was added to the downstream of the GFP gene and recombined into the pcDNA3.1 vector, and then the linker 1 sequence (sequence such as SEQ ID NO: 549) was inserted downstream.
- the linker 1 used in this embodiment is 2121bp long, and a reporter gene plasmid vector is obtained, wherein the linker 1 contains a 10bp NNNYRNNNYR (Barcode sequence, wherein N represents one of the four bases A, C, G, and T, Y represents C or T base, and R represents A or G base) degenerate bases are used to distinguish the editing efficiency of each gRNA.
- the amplified gRNA library is recombined into the reporter gene plasmid vector to obtain the screening vector as shown in Figure 2.
- the RNA editing mode of the transcript expressed by the vector is shown in Figure 3.
- library construction primers are shown in Table 2
- the target site editing and off-target editing levels of each gRNA were analyzed, as shown in Figures 4 and 5.
- a long linker 1 sequence is added to simulate the formation of double strands between intermolecular RNA for editing.
- Two libraries are constructed during library construction. One is the regional library of Reporter and Linker to construct the correspondence between the barcode sequence and the editing level on each Linker.
- the first round of primers used are cDNA-NGS-round1 primer pairs, and the second round of primers are NGS-round2 primer pairs; the other is the region of Linker and gRNA library. Because the interval between Linker and gRNA exceeds 2000bp, it is impossible to directly construct a library for sequencing. Therefore, the full-length sequences of Linker and gRNA are first cyclized, and after cyclization, the first round of amplification is performed with the plasmid-NGS-round1 primer pair, and then the NGS-round2 primer pair is used to complete the construction of the entire library.
- This library can establish the correspondence between the barcode on the Linker and the gRNA sequence, and the editing efficiency corresponding to each gRNA can be obtained by comparing the barcode sequences of the two libraries.
- Editing site 1 is an AAC motif ( Figure 4a), and the black dots represent the editing level of the gRNA (the editing site base is a C base) extending from the editing site to both ends for a total of 151nt in two repeated experiments. The remaining yellow dots are the editing levels of different gRNAs after the introduction of mutations in two repeated experiments.
- the gRNA for the AAC motif has a uniform gRNA distribution within 50% of the editing efficiency, and only a few gRNAs can reach the editing efficiency range of 50% to 75%. It can be seen from the heat map (Figure 5a) that the gRNA designed for the AAC motif in this embodiment has obvious off-target editing at multiple sites, but the editing level of the target site is higher than that of the off-target site as a whole.
- the AAU motif gRNA of editing site 2 (Figure 4b) showed an overall editing effect of less than 40%, and the editing efficiency of the gRNAs extending upstream and downstream from the editing site, except for the AC mismatch at the editing site, was also less than 30%.
- the gRNA editing efficiency of the AAG motif (Figure 4c) is similar to that of the AAC motif. Unlike the other three editing sites, the overall editing level of the gRNA of the UAA motif ( Figure 4d) is between 25% and 75%, but the A base adjacent to the gRNA of the UAA motif has relatively serious off-target editing, and only the gRNA screened by the AAG motif has low overall off-target editing ( Figure 5).
- the gRNA designed in this embodiment is about 151nt in length.
- the gRNAs corresponding to different sites A have different editing abilities, indicating that gRNA sequences with different editing levels can be screened out by the present invention.
- This example aims to verify that short gRNA and short linkers can be screened for editing specific sites on the human GAPDH gene.
- the A base at site 787 on the GAPDH gene was selected, and the motif where it is located is the TAG motif. 1073 gRNAs (length range 28nt-89nt) with different lengths and mutations at different sites were designed for this site.
- a short linker was selected to verify the screening system, in which the sequence of linker 2 is NNNNNNNNTGGGTTGAGGGTAGTGAG (sequence number is SEQ ID NO: 556; N is any one of the bases A, T, C, G, and 8 Ns are the barcode sequence of the linker).
- the specific implementation scheme is as follows: first, directly synthesize the gRNA library, directly connect the linker 2 to the gRNA sequence to synthesize the gRNA library with the barcode sequence; secondly, construct a reporter vector, and recombinantly insert the sequence of 201bp extending upstream and downstream with the editing site as the center into the pcDNA3.1 vector with the GFP sequence; thirdly, insert the gRNA library of step 1 into the reporter vector to complete the construction of the screening vector; fourthly, transfect the screening vector into HEK293 cells overexpressing ADAR1 protein for 48 hours, extract RNA, sequence the targeted library, and analyze the target site editing and off-target editing levels of each gRNA, as shown in Figure 6.
- the results show that in this embodiment, in the case of a short linker, a mixed library of gRNAs of different lengths can also be screened for gRNAs with different editing efficiencies by this method.
- the high-throughput screening established based on the method of the present invention is to transfect a mixed gRNA library into cells, and there may be a certain degree of interference between gRNAs. Therefore, this embodiment verifies a single gRNA on the basis of high-throughput screening.
- This embodiment selects gRNA sequences with editing levels ranging from 3.93% to 21.66% from high-throughput screening.
- the sequence information is GAPDH-TAG-S1 to GAPDH-TAG-S12 in Table 3, and the chemically modified gRNA sequences correspond to GAPDH-TAG-M1 to GAPDH-TAG-M12.
- the chemically modified single RNAs are transfected into cells to verify the editing level (as shown in Figure 8 and Table 3).
- the stability of the modified gRNA is higher than that of the gRNA expressed by the plasmid, so the editing level is higher than that of the high-throughput screening.
- the modification strategy of the chemically modified gRNA is to modify the bases with phosphorothioate and add 3 LNA modifications at both ends.
- Analysis of the editing levels screened by high-throughput screening and the editing levels of single modified gRNAs shows that the two show a strong correlation, with a correlation coefficient of 0.83 (as shown in Figure 7), confirming the accuracy of the method of the present invention.
- the editing levels of GAPDH-TAG-S1 to GAPDH-TAG-S13 in Table 3 are the gRNA editing levels measured by the high-throughput gRNA mixed system; the editing levels of GAPDH-TAG-M1 to GAPDH-TAG-M13 are the editing levels measured by the single gRNA system.
- This example aims to verify that short gRNAs for editing specific sites can be screened on the human ATP7B (NM_000053.4) gene.
- ATP7B gene mutation is the main cause of Wilson's disease.
- the A base at site 1532 on the ATP7B gene is selected.
- the motif of the editing site after mutation is the UAG motif.
- 1011 gRNAs of different lengths and mutations at different sites are designed for this site.
- This example selects gRNAs with lengths of 31nt and 41nt, including the editing site and its upstream and downstream. Introducing mutations or deletions or adding protrusions for these two lengths will cause the actual length to float a little, with a floating range of 27nt-59nt.
- connector 2 is selected to verify the screening system.
- the specific implementation plan is as follows: first, directly synthesize gRNA libraries with lengths of 31nt and 41nt, and directly connect linker 2 to the gRNA sequence to synthesize a gRNA library with a barcode sequence; secondly, construct a reporter vector, and recombinantly insert a 201bp sequence extending upstream and downstream with the editing site as the center and the GFP sequence into the pcDNA3.1 vector; the third step is to insert the gRNA library of step one into the reporter vector to complete the construction of the screening vector; the fourth step is to transfect the screening vector into HEK293 cells overexpressing ADAR1 protein, extract RNA and sequence the targeted library 48 hours later, and analyze the target site editing and off-target editing levels of each gRNA, and the results are shown in Figures 9 and 10.
- sequences with high editing efficiency and long sequencing reads were selected from the two groups of results of high-throughput screening for the next round of error prone verification of the editing effect after the introduction of mutations.
- the selected sequences from the 41nt group are shown in SEQ ID NO: 583 to SEQ ID NO: 602, and the selected sequences from the 31nt group are shown in SEQ ID NO: 603 to SEQ ID NO: 622.
- Error prone PCR primer sequences were added to both ends of all selected sequences, with TCGAAGGTCGCTTAGACGGC (SEQ ID NO: 581) at the 5' end and CGGTCGTGAGTGCAGACGTG (SEQ ID NO: 582) at the 3' end to directly synthesize single-stranded sequences.
- TCGAAGGTCGCTTAGACGGC SEQ ID NO: 581
- CGGTCGTGAGTGCAGACGTG SEQ ID NO: 582
- the editing efficiency of the 31nt group was improved from less than 20% in the first round to 20% to 80%, and was mainly distributed within 40%.
- the editing ratio of 40% to 80% was relatively small, and the editing efficiency of the 41nt group was 20% to 50%.
- After the second round of screening it was obviously seen that the editing ratio of 40% to 90% was relatively high.
- This example aims to verify that short gRNAs for editing specific sites can be screened on the human FGFR (NM_015850.4) gene.
- FGFR gene mutations are the cause of diseases such as achondroplasia, lethal bone dysplasia, and dwarfism.
- the G358R site on the FGFR gene produces a mutation from G to A, so the motif of the editing site after mutation is the CAG motif.
- 989 gRNAs of different lengths and mutations at different sites are designed for this site.
- This example selects gRNAs with lengths of 31nt and 41nt including the editing site and its upstream and downstream. The introduction of mutations or deletions or the addition of protrusions for these two lengths will cause the actual length to fluctuate a little, with a floating range of 23nt-62nt.
- the screening system was verified by using linker 2.
- the specific implementation plan is as follows: first, directly synthesize 31nt and 41nt gRNA libraries, directly connect linker 2 to the gRNA sequence to synthesize the gRNA library with the barcode sequence; secondly, construct a reporter vector, and recombinantly insert the 201bp sequence extending upstream and downstream with the editing site as the center into the pcDNA3.1 vector with the GFP sequence; thirdly, insert the gRNA library of step 1 into the reporter vector to complete the construction of the screening vector; fourthly, transfect the screening vector into HEK293 cells overexpressing ADAR1 protein, extract RNA and sequence the targeted library 48 hours later, and analyze the target site editing and off-target editing levels of each gRNA, as shown in Figures 12 and 13.
Landscapes
- Health & Medical Sciences (AREA)
- Life Sciences & Earth Sciences (AREA)
- Genetics & Genomics (AREA)
- Engineering & Computer Science (AREA)
- Biomedical Technology (AREA)
- Chemical & Material Sciences (AREA)
- Biotechnology (AREA)
- Organic Chemistry (AREA)
- Zoology (AREA)
- General Engineering & Computer Science (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Wood Science & Technology (AREA)
- Molecular Biology (AREA)
- General Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Microbiology (AREA)
- Biochemistry (AREA)
- Plant Pathology (AREA)
- Physics & Mathematics (AREA)
- Crystallography & Structural Chemistry (AREA)
- Medicinal Chemistry (AREA)
- Pharmacology & Pharmacy (AREA)
- Epidemiology (AREA)
- Animal Behavior & Ethology (AREA)
- Public Health (AREA)
- Veterinary Medicine (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
一种编辑RNA的向导RNA的筛选方法,包括:(1)将所述向导RNA通过连接子分别与靶RNA连接在同一个转录本上;(2)将所述的转录本通过载体转入细胞中;(3)提取所述细胞的RNA,分析所述向导RNA对靶位点的编辑水平,筛选出目的向导RNA。所述筛选方法将底物序列和向导RNA序列设计为同一个转录本,直接在细胞内完成靶RNA编辑,通过高通量测序将每个向导RNA的序列和其对应的靶位点的编辑效率测序出来,从而快速筛选出高编辑效率的向导RNA。
Description
本发明属于生物医药领域,涉及一种筛选向导RNA的方法。
近年来随着三代测序技术的快速发展,被誉为生命科学“登月计划”的人类基因组测序最后一块拼图已经完成,新增绘制了近2亿个碱基对的完整无间隙基因组图谱,这将意味着越来越多与疾病相关的基因突变位点可被更彻底的挖掘出来。在此之前已经发现超过3000多个基因突变位点与疾病的发生发展有密切关联,而很多突变位点却很难成为合适的成药靶点,除此之外同一个疾病可能对应多个突变位点,且在人群中分布不均,让传统的大分子小分子药物研发变得更加艰难。此时,基因疗法慢慢展现出其巨大的临床应用价值,直接对突变位点进行修复达到疾病治疗的目的。
CRISPR(Clustered regularly interspaced short palindromic repeats)最早发现于细菌免疫系统中,可结合cas蛋白抵抗外源病毒的入侵。自2012年被报道该系统也可在哺乳动物细胞中进行基因组的剪切及修复后对CRISP/Cas的研究如同雨后春笋,其编辑的特性机制、可应用的物种组织等被一一攻克,目前CRISPR/Cas技术已经广泛应用到生物医学及农业学研究中。CRISPR/Cas系统可被设计在基因组中插入完整的基因、敲低基因的表达、敲除基因或者对单个突变位点进行修复等功能。但该系统也有其致命的缺陷,该系统需要同时表达Cas蛋白及向导RNA(gRNA)及提供同源修复DNA序列,且在多篇研究报道中都检测出其在DNA水平或RNA水平有明显的脱靶效应,CRISPR/Cas蛋白切割基因组DNA造成DSB(double-strand break),脱靶会造成基因组的不稳定性,此外,CRISPR/Cas系统还可能会引发细胞固有免疫反应,这些都极大的限制了其在临床上的应用。
RNA水平的A-to-I编辑则是在ADAR家族蛋白的介导下发生,I碱基在翻译的时候会被识别为G碱基,故在基因组上G-to-A的突变就可以通过ADAR的编辑在RNA水平进行修复。ADAR家族蛋白主要有三个亚型,ADAR1,ADAR2和ADAR3,其中ADAR3不具有脱氨酶活性,编辑事件主要由ADAR1和ADAR2完成,且ADAR1蛋白在人体组织中广谱表达可确保其应用到多种组织或器官中。目前已经开发出多种利用过表达ADAR家族蛋白或者内源表达的ADAR家族蛋白结合向导RNA在RNA水平进行特定位点的A-to-I编辑。从临床应用的角度分析最具应用价值的方法则是利用内源表达的ADAR家族蛋白对突变位点进行修复,因为只需对gRNA进行设计,而目前比较成熟的技术则是如下两种:第一种gRNA由两段组成,一段是54nt长度的带有结构特性的ADAR底物加上另外一段可定制的靶向序列(RESTORE方法),该方法优化版本可减少脱靶编辑(CLUSTER方法);第二种gRNA只有一段与靶标序列互补的长度大于100nt的靶向序列(LEAPER方法)。ADAR家族蛋白偏好于编辑双链RNA,但在编辑区域范围内同时也会存在多个突变形成一定的二级结构,这就导致在设计gRNA的时候需对gRNA进行筛选,而目前还未有报道出比较高效的筛选方法。
发明内容
一些实施方案中,本发明提供了一种编辑RNA的向导RNA的筛选方法,包括步骤;(1)待筛选的每条向导RNA连有连接子,并与靶RNA连接在能够形成同一个转录本的核酸链上,形成多个所述的核酸链;(2)将所述的核酸链转入表达ADAR蛋白的细胞中;(3)提取和测定所述细胞中的RNA,确定靶RNA的被编辑水平,通过条形码序列确定对应的向导RNA的编辑水平。
一些实施方案中,所述的连接子位于靶RNA和向导RNA之间。
一些实施方案中,所述的条形码序列包含于所述连接子中。
一些实施方案中,所述核酸链从上游到下游依次连接有报告基因核酸片段、靶RNA核酸片段、连接子核酸片段、向导RNA核酸片段。
连接子序列在本发明中起到标签的作用。连接子优选地可以位于靶RNA和向导RNA之间,此时连接子同时起到在
一定程度隔开靶RNA和向导RNA的作用,有利于两者折叠成二级结构。并且,此种连接方式还有利于在二代测序中应用,从而实现高效高通量的检测。为了适应二代测序的序列长度限制,测序时可以切割成靶RNA+连接子、连接子+向导RNA两部分测序结果,通过连接子的标签作用,可获悉相应向导RNA作用下靶RNA的编辑水平。
一些实施方案中,所述转录本从上游到下游依次连接有报告基因核酸片段、靶RNA核酸片段、连接子核酸片段、向导RNA核酸片段和polyA尾核酸片段。
一些实施方案中,所述报告基因选自yfp基因、gus基因、rfp基因、gfp基因、cfp基因、bfp基因或卡那霉素抗性基因,但又不限于此。
一些实施方案中,所述连接子的长度为26nt-2121nt。
一些实施方案中,所述连接子包括SEQ ID NO:549或SEQ ID NO:557。
一些实施方案中,筛选方法中加入的是长连接子序列,目的是为了模拟分子间RNA之间形成双链进行编辑。
一些实施方案中,在短连接子的情况下不同长度的gRNA混合的文库也可以通过本发明的方法筛选出不同编辑效率的gRNA。
一些实施方案中,步骤(3)包括建立靶向测序文库进行二代测序。
一些实施方案中,所述核酸链接触ADAR蛋白通过将所述核酸链转入表达ADAR蛋白的细胞中实现。
一些实施方案中,所述细胞表达ADAR蛋白。
一些实施方案中,所述ADAR蛋白包括ADAR1或ADAR2。
一些实施方案中,所述向导RNA能够在所述细胞中招募ADAR蛋白对靶RNA进行编辑。
一些实施方案中,所述的细胞过表达ADAR蛋白。
一些实施方案中,所述编辑为A-to-I编辑。
一些实施方案中,所述向导RNA为以A碱基为中心上下游延伸得到的向导RNA。
一些实施方案中,所述向导RNA为以A碱基为中心上下游延伸15nt~75nt得到的30nt~151nt长度的向导RNA。
一些实施方案中,所述向导RNA为以A碱基为中心上下游延伸75nt得到的151nt的向导RNA。
一些实施方案中,所述向导RNA为以A碱基为中心上下游延伸得到的28nt~89nt的向导RNA。
一些实施方案中,所述筛选方法包括把编辑位点上下游20-200nt序列加入报告基因的下游并插入到载体中获得报告基因载体,再把报告基因载体转入细胞中。
一些实施方案中,所述筛选方法为高通量筛选方法。
一些实施方案中,本发明提供了一种可招募ADAR蛋白并对特定位点进行A-to-I编辑的高通量筛选向导RNA(gRNA)序列的方法,填补了现有技术的空缺。
一些实施方案中,本发明提供了上述方法对基于ADAR的任何RNA编辑系统建立的gRNA筛选方法。
一些实施方案中,所述靶RNA序列长度大于向导RNA序列。
一些实施方案中,所述载体包括质粒或病毒载体。
一些实施方案中,所述载体是用于在高等真核细胞或原核细胞中表达的质粒或病毒载体。一些实施方案中,所述载体骨架为pcDNA3.1载体。
一些实施方案中,所述筛选方法包括:将多条不同的所述向导RNA分别与所述靶RNA连接组成转录本体系,通过所述载体转入细胞中。
一些实施方案中,所述目的向导RNA为靶向位点进行高效编辑的向导RNA。
一些实施方案中,所述方法还包括:从所述步骤(3)中获得所述向导RNA的编辑水平后,挑选编辑水平高的向导RNA合成序列相同的单链DNA后设计易错PCR引物对所述向导RNA的靶向区域进行随机引入突变,再重复所述步骤(1)~(3)。
一些实施方案中,本发明提供了一种编辑RNA的向导RNA的筛选方法,包括:(a)向导RNA文库的设计;(b)把需要编辑的基因片段与报告基因插入到载体中获得报告基因载体;(c)把所述向导RNA文库插入所述报告基因载体获得待筛选质粒文库;(d)将所述待筛选质粒文库转入到细胞,转染后,提取所述细胞中的RNA,测序;(e)使用生物信
息学分析所述向导RNA的编辑水平;其中(a)和(b)的顺序可以调换。
一些实施方案中,步骤(b)包括:把编辑位点上下游序列加入报告基因的下游插入到载体中获得报告基因载体。一些实施方案中,步骤(b)包括:把编辑位点上下游20-200nt序列加入报告基因的下游插入到载体中获得报告基因载体。
一些实施方案中,步骤(b)包括:将荧光基因全长以及需要编辑的基因片段通过PCR扩增出来同时两端加上同源重组的同源序列或引入酶切位点,同时对空载体进行NheI和HindIII双酶切并回收线性载体,将PCR扩增的多片段插入到线性载体中获得报告基因的载体。
一些实施方案中,步骤(c)包括:扩增步骤(a)中的所述向导RNA同时加上同源臂或酶切位点获得向导RNA序列池,步骤(b)中的报告基因载体用BamHI及EcoRI进行双酶切线性化;然后把所述的向导RNA序列池插入到线性化后的报基因载体中,转化后获得待筛选质粒文库。
一些实施方案中,所述筛选方法还包括:(f)从所述向导RNA中挑选出编辑效率高的向导RNA直接合成序列相同的单链DNA后,设计易错PCR引物对所述向导RNA靶向区域进行随机引入突变,再重复步骤(b)~(e)。
一些实施方案中,本发明提供了一种可招募ADAR蛋白并对特定位点进行A-to-I编辑的高通量筛选向导RNA序列的方法,包括如下步骤:S1.对需要筛选的gRNA文库进行接头序列设计并合成,具体地,gRNA文库一端为芯片合成所需的固定序列,另一端为可区分不同文库的20nt左右的接头序列;S2.筛选载体的设计及构建,具体地,本发明设计了在绿色荧光蛋白GFP编码区下游添加靶标RNA和gRNA序列,三者在同一条转录本的策略进行实施,三者之间可以直接连接也可以通过不同长度的接头序列进行间隔;S3.待筛选的gRNA质粒文库的构建,具体地,通过对应不同文库的接头序列对混合的文库进行PCR扩增并进一步插入到S2的载体系统中;S4.通过转染的方式把质粒文库转染到特定的细胞中;S5.RNA提取及靶向测序文库构建并进行二代测序;S6.生物信息学方法分析每条gRNA对应靶标RNA的编辑效率;S7.挑选高编辑水平的gRNA序列合成序列相同的单链DNA并进行易错PCR构建质粒文库并按照S3至S6的方法进行第二轮的筛选。
一些实施方案中,本发明提供了一种可招募ADAR蛋白并对特定位点进行A-to-I编辑的高通量筛选向导RNA序列的方法。本发明通过开发实验及生物信息学方法填补了高通量筛选基于ADAR的靶标RNA编辑的gRNA序列方法上的空缺,通过一系列优选的实验设计方法完成筛选。
一些实施方案中,本发明所述的方法包括1)采用模拟基因不同区域的方式把靶标RNA及gRNA设计到同一个转录本上,荧光基因的引入便于观测转染效率且可替换成其他荧光基因;2)构建靶向测序文库及分析方法可以精准的把每条gRNA与其编辑效率一一对应;3)在筛选出的高编辑水平的gRNA序列基础上进行随机引入突变进一步筛选出更高编辑水平的gRNA。
一些实施方案中,本发明提供了任一所述的筛选方法获得的向导RNA。
一些实施方案中,本发明提供了一种向导RNA,所述向导RNA的序列如SEQ ID NO:557、SEQ ID NO:559、SEQ ID NO:561、SEQ ID NO:563、SEQ ID NO:565、SEQ ID NO:567、SEQ ID NO:569、SEQ ID NO:571、SEQ ID NO:573、SEQ ID NO:575、SEQ ID NO:577、SEQ ID NO:579、SEQ ID NO:583~SEQ ID NO:602、SEQ ID NO:603~SEQ ID NO:622、或SEQ ID NO:623~SEQ ID NO:642任一所示。
一些实施方案中,本发明提供了获得所述的向导RNA的方法或者获得的向导RNA在制备治疗疾病的药物中的用途。
一些实施方案中,所述药物用于在真核细胞中定点编辑靶基因中的核苷酸从而达到治疗疾病的目的。
一些实施方案中,所述疾病为包括GAPDH基因、ATP7B基因或FGFR基因突变引起的疾病。
一些实施方案中,所述GAPDH基因特定位点基序为AAC、AAU、AAG、UAA或UAG,但又不限于此。
一些实施方案中,所述ATP7B基因特定位点基序为UAG,但又不限于此。
一些实施方案中,所述FGFR基因特定位点基序为CAG,但又不限于此。
一些实施方案中,所述疾病包括肝豆状核变性、软骨发育不全、致死性骨发育不良或矮小症,但又限于此。
一些实施方案中,所述用途是通过天然存在于所述细胞中并且能够进行所述核苷酸的编辑的RNA编辑实体的作
用。
一些实施方案中,本发明提供了一种工程化RNA,其结构包括:从上游到下游依次连接有报告基因核酸片段、靶RNA核酸片段、连接子核酸片段、向导RNA核酸片段。
一些实施方案中,本发明提供了一种工程化RNA,其结构包括:从上游到下游依次连接有报告基因核酸片段、靶RNA核酸片段、连接子核酸片段、向导RNA核酸片段和polyA尾核酸片段。
一些实施方案中,所述报告基因选自yfp基因、gus基因、rfp基因、gfp基因、cfp基因、bfp基因、或卡那霉素抗性基因。
一些实施方案中,所述连接子的长度为26nt到2121nt。
一些实施方案中,所述连接子包括SEQ ID NO:549或SEQ ID NO:557。
一些实施方案中,所述向导RNA为以A碱基为中心上下游延伸得到的向导RNA。
一些实施方案中,所述向导RNA为以A碱基为中心上下游延伸75nt得到的151nt的向导RNA。
一些实施方案中,本发明提供了一种工程化RNA结构文库,包含所述的多个不同的工程化RNA。
一些实施方案中,所述不同的工程化RNA是通过将靶RNA通过不同的连接子与不同的向导RNA连接获得的。
一些实施方案中,所述工程化RNA结构文库中,所述文库中,所述的工程化RNA所含有的靶RNA核酸片段是相同的,向导RNA核酸片段序列是不同的。
一些实施方案中,本发明提供了一种载体,包含所述的工程化RNA或所述的工程化RNA结构文库。
一些实施方案中,所述载体通过把靶RNA序列加入到报告基因序列的下游插入到载体骨架中,获得报告基因载体,再把扩增好的向导RNA文库插入入报告基因载体中获得的。
一些实施方案中,本发明提供了一种细胞,包含所述的工程化RNA或所述的工程化RNA结构文库或所述的载体。
图1为高通量筛选方法示意图。
图1显示高通量筛选向导RNA的整体流程,待筛选序列的设计;合成好的待筛选序列构建到筛选载体;导入到细胞中进行编辑;RNA提取、二代测序文库构建及高通量测序;利用生物信息学方法对测序数据进行分析获取每条gRNA的编辑效率;挑选高编辑效率的gRNA进行验证;从高编辑gRNA中挑选多条进行Error prone随机引入突变再次重复筛选流程得到更高编辑效率的gRNA。
图2为筛选载体转录本信息示意图。
图2显示筛选载体具体的转录本组成信息,以GFP作为示例构建一个可以同时检测到编辑水平和gRNA的转录本,该转录本模拟正常基因的表达,GFP为编码区,待编辑的靶向序列及连接子及gRNA模拟3’UTR区域,其中靶RNA(Reporter)和gRNA中会有一段可预设长度的连接子(Linker)序列,一方面可以间隔开二者模拟分子间形成二级结构,另一方面在Linker区域添加Barcode(用于区分不同的gRNA)序列可以在后续对高通量测序数据分析时对不同的gRNA及对应的编辑水平做区分。
图3筛选载体转录本进行RNA编辑模式图。
图3显示筛选载体表达出RNA后会进行自我折叠形成Reporter和gRNA的双链RNA区域,该区域会被ADAR蛋白识别结合并对靶位点进行编辑。
图4为GAPDH基因不同位点gRNA筛选结果。(图4a)编辑位点1为AAC基序的结果。(图4b)编辑位点2为AAU基序的结果。(图4c)编辑位点3为AAG基序的结果。(图4d)编辑位点4为UAA基序的结果。
图4显示GAPDH基因上选择的四个不同基序的编辑位点进行高通量筛选得到的每条gRNA的编辑效果,其中横纵坐标表示的为两次生物学重复。图中黑色的点为编辑位点对位为C碱基其他位置为互补配对的gRNA,非黑色的点则为其他的gRNA在该序列的基础上引入不同的突变等得到的序列,每个位点对应的编辑水平均为高通量测序获得。
图5为GAPDH基因不同位点gRNA编辑及脱靶编辑比较。
图5显示GAPDH基因不同位点gRNA在靶位点的编辑及上下游的脱靶编辑热图比较。其中横坐标的0表示靶位点被编辑
的A碱基的位置,0~-25则为编辑位点上游的区间,0~25为编辑位点下游的区间,纵坐标为每一条gRNA在不同A碱基上的编辑水平的情况,按照靶位点最高编辑序列往下排列,蓝色到红色表示A碱基位点的编辑水平从低到高的变化趋势。
图6为不同长度gRNA靶向人GAPDH基因的第787位点的编辑及脱靶编辑效率。
图6显示GAPDH基因的第787位所有gRNA在靶位点编辑及上下游脱靶编辑热图。其中横坐标的-1与1之间为靶位点,-1~-25为编辑位点上游的区间,1~25为编辑位点下游的区间,纵坐标为每一条gRNA在不同A碱基上的编辑水平情况,按照靶位点最高编辑序列往下排列,蓝色到红色表示A碱基位点的编辑水平从低到高的变化趋势。
图7高通量筛选gRNA验证分析。
图7显示高通量筛选出的gRNA用低通量的方式进行验证。横坐标表示每条gRNA在高通量筛选的方法中测量得到的编辑水平,纵坐标为每条gRNA单独合成修饰后的RNA并转染到细胞内测量得到的编辑水平。相关性系数R2值越高表示横纵坐标相关性越高。
图8为单条gRNA在GAPDH基因上的编辑水平一代测序原始结果。
图8显示单条修饰后的gRNA转染到细胞中通过一代测序获得的编辑水平的原始峰图。如图所示,不同的碱基对应不同颜色的峰,A碱基对应绿色的峰,G碱基对应黑色的峰,当一个位点出现不同颜色的峰表示该位点存在不同的碱基组成,在本结果指示A碱基的编辑水平,编辑水平的计算方法为G碱基的峰面积除以该位点A碱基和G碱基的峰面积之和。
图9为人ATP7B基因gRNA筛选结果。
图9显示ATP7B基因上选择的编辑位点进行高通量筛选得到的每条gRNA的编辑效果,其中横纵坐标表示的为两次生物学重复。图中黑色的点为编辑位点对位为C碱基其他位置为互补配对的gRNA,非黑色的点则为其他的gRNA在该序列的基础上引入不同的突变等得到的序列,每个点对应的编辑水平均为高通量测序获得,len41表示靶编辑位点及上下游41nt长度对应设计出的带有不同突变的gRNA,len31表示靶编辑位点上下游31nt长度对应设计出的带有不同突变的gRNA。
图10为人ATP7B基因gRNA靶向编辑及脱靶编辑效率。
图10显示ATP7B基因所有gRNA在靶位点编辑及上下游脱靶编辑热图。其中横坐标的0位置为靶位点,0~-20为编辑位点上游的区间,0~20为编辑位点下游的区间,纵坐标为每一条gRNA在不同A碱基上的编辑水平情况,按照靶位点最高编辑序列往下排列,蓝色到红色表示A碱基位点的编辑水平从低到高的变化趋势。
图11为ATP7B基因UAG基序优选gRNA的第二轮筛选结果。
图11显示ATP7B基因通过第一轮筛选得到的不同长度的高编辑效率前20条序列进行Error prone引入随机突变并进行第二轮筛选,筛选载体转染细胞后用高通量测序方法测量新产生的gRNA和其对应的编辑效率,图中横纵坐标表示的为两次生物学重复。len41表示选择41nt长度组的gRNA进行第二轮筛选的结果,len31表示选择31nt长度组的gRNA进行第二轮筛选的结果。
图12人FGFR基因gRNA筛选结果。
图12显示FGFR基因上选择的编辑位点进行高通量筛选得到的每条gRNA的编辑效果,其中横纵坐标表示的为两次生物学重复。图中黑色的点为编辑位点对位为C碱基其他位置为互补配对的gRNA,非黑色的点则为其他的gRNA在该序列的基础上引入不同的突变等得到的序列,每个点对应的编辑水平均为高通量测序获得,len41表示靶编辑位点及上下游41nt长度对应设计出的带有不同突变的gRNA,len31表示靶编辑位点上下游31nt长度对应设计出的带有不同突变的gRNA。
图13为人FGFR基因gRNA靶向编辑及脱靶编辑效率。
图13显示FGFR基因所有gRNA在靶位点编辑及上下游脱靶编辑热图。其中横坐标的0位置为靶位点,0~-20为编辑位点上游的区间,0~20为编辑位点下游的区间,纵坐标为每一条gRNA在不同A碱基上的编辑水平情况,按照靶位点最高编辑序列往下排列,蓝色到红色表示A碱基位点的编辑水平从低到高的变化趋势。
图14为人FGFR基因CAG基序优选gRNA的第二轮筛选结果。
图14显示FGFR基因通过第一轮筛选得到的41nt长度高编辑效率前20条序列进行Error prone引入随机突变并进行第二轮筛选,筛选载体转染细胞后用高通量测序方法测量新产生的gRNA和其对应的编辑效率,图中横纵坐标表示的为两次生物学重复。len41表示选择41nt长度组的gRNA进行第二轮筛选的结果。
以下通过具体的实施例进一步说明本发明的技术方案,具体实施例不代表对本发明保护范围的限制。其他人根据本发明理念所做出的一些非本质的修改和调整仍属于本发明的保护范围。
除非特别说明,以下实施例所用试剂和材料均为市场上可以购买获得的。
本文中向导RNA高通量筛选方法流程如图1所示。
如本文和所附权利要求书所用,单数形式“一个/一种”、“一个/一种”和“所述”包括复数指代物,除非上下文另外清楚地规定。因此,例如,提及“一种方法”包括多个此类方法,并且提及“所述片段”包括提及一个或多个片段及其本领域技术人员已知的等同物,等等。
此外,除非另外说明,否则“或”的使用意指“和/或”。类似地,“包含”、“包括”是可互换的,并不旨在是限制性的。
应进一步理解,在使用术语“包含”描述各个实施方案的情况下,本领域技术人员将理解,在一些特定情形下,可以可替代地使用语言“基本上由……组成”或“由……组成”来描述实施方案。
除非另外定义,否则本文所用的所有技术和科学术语都具有与本公开文本所属领域的普通技术人员通常所理解的相同的含义。尽管许多方法和试剂与本文所述的那些类似或等同,但本文公开了示例性方法和材料。
应当理解,本公开文本不限于本文所述的特定方法、方案和试剂等,并且本身可以变化。本文所用的术语仅用于描述特定实施方案或方面的目的,并不旨在限制本公开文本的范围。
“工程化的多核苷酸”或“工程化的向导RNA”可与环状向导RNA互换地使用。经工程化的多核苷酸可以包括DNA或RNA或杂交DNA/RNA构建体的重组多核苷酸。经工程化的多核苷酸可以产生向导RNA,更具体地可以产生环状向导RNA。
如本文所用,术语“突变”可以指代编码蛋白质的核酸序列相对于所述蛋白质的共有序列的改变。“错义”突变导致一个密码子取代为另一个密码子;“无义”突变将密码子从编码特定氨基酸的密码子改变为终止密码子。无义突变通常导致蛋白质的截短翻译。“沉默”突变是对所得蛋白质没有影响的突变。如本文所用,术语“点突变”可以指代只影响基因序列中的一个核苷酸的突变。“剪接位点突变”是存在于前体mRNA(在加工以除去内含子之前)中的那些突变,其由于剪接位点的不正确描绘而导致蛋白质的错译并且通常导致截短。突变可以包括单核苷酸变异(SNV)。突变可以包括序列变体、序列变异、序列改变或等位基因变体。参考DNA序列可以从参考数据库获得。突变可能影响功能。突变可能不影响功能。突变可以在DNA水平上在一个或多个核苷酸中发生,在核糖核酸(RNA)水平上在一个或多个核苷酸中发生,在蛋白质水平上在一个或多个氨基酸中发生,或其任何组合。参考序列可以从诸如NCBI参考序列数据库(RefSEQ ID NO:)数据库等的数据库获得。可以构成突变的特定改变可以包括在一个或多个核苷酸或者一个或多个氨基酸中的取代、缺失、插入、倒位或转换。突变可以是点突变。
术语“RNA编辑”是指,在基因组编码的RNA序列中引入变化,从而导致RNA突变的共转录或转录后修饰过程。双链RNA(dsRNA)中的腺苷编辑为肌苷(A-到-I),由作用于RNA(ADAR)酶家族的腺苷脱氨酶催化,是哺乳动物中常见的RNA编辑类型。在脊椎动物中,先前已对三种ADAR蛋白ADAR1、ADAR2和ADAR3的家族进行了表征。ADAR1和ADAR2(ADAR)催化所有当前已知的A-到-I编辑位点。ADAR3没有已知的脱氨酶活性。肌苷(I)模拟鸟苷(G),因此ADAR蛋白在转录物中引入了虚拟的A到G取代。这种变化可以导致特定的氨基酸取代、可变剪接、微RNA介导的基因沉默或转录物定位和稳定性的变化。
在脱氨基之后,根据靶标RNA中的靶标腺苷的位置,可以使用不同方法确定靶标RNA和/或靶标RNA编码的蛋白质的修饰。例如,为了确定在靶标RNA中是否已将“A”编辑为“I”,可以使用本领域已知的RNA测序方法来检测RNA序列的修饰。当靶标腺苷位于mRNA的编码区中时,RNA编辑可以引起mRNA编码的氨基酸序列的改变。例如,可以将点
突变引入mRNA,或者mRNA中的先天的或获得性的点突变被回复以产生野生型基因产物,因为“A”转化为了“I”。通过本领域已知的方法进行氨基酸测序可用于发现编码蛋白质中氨基酸残基的任何变化。终止密码子的修饰可以通过评估功能性的、伸长的、截短的、全长的和/或野生型的蛋白质的存在来确定。例如,当靶标腺苷位于UGA,UAG或UAA终止密码子中时,靶标A(UGA或UAG)或多个A(UAA)的修饰可创造出通读突变和/或延长的蛋白质,或者,由靶标RNA编码的截短蛋白质可以被回复以产生有功能的、全长的和/或野生型的蛋白质。靶标RNA的编辑还可以在靶标RNA中产生异常剪接位点和/或可变的剪接位点,从而导致延长的,截短的或错误折叠的蛋白质,或者靶标RNA中经编码的异常剪接或可变剪接位点经回复以产生有功能的、正确折叠的、全长的和/或野生型的蛋白质。在一些实施方案中,本申请考虑编辑先天的和获得性的基因变化,例如,错义突变,过早的终止密码子,异常剪接或由靶标RNA编码的可变剪接位点。使用已知方法评估由靶标RNA编码的蛋白质的功能,可以发现RNA编辑是否实现了期望的效果。因为腺苷(A)对肌苷(I)的脱氨基可以纠正编码蛋白质的突变RNA中的靶标位置处的突变A,所以对脱氨基为肌苷的鉴定可以提供功能蛋白质是否存在的评估,或者突变的腺苷引起的、与疾病或药物抗性相关的RNA是否已回复或部分回复的评估。类似地,因为腺苷(A)对肌苷(I)的脱氨基作用可能在所得蛋白质中引入点突变,所以对脱氨基为肌苷的鉴定可以找出疾病原因或疾病相关因素的功能性指征。
在一些实施方案中,靶标RNA是调节RNA。在一些实施方案中,待编辑的靶标RNA是核糖体RNA,转移RNA,长非编码RNA或小RNA(例如,miRNA,pri-miRNA,pre-miRNA,piRNA,siRNA,snoRNA,snRNA,exRNA或scaRNA)。靶标腺苷的脱氨基作用包括例如核糖体RNA,转移RNA,长非编码RNA或小RNA(例如,miRNA),包括了三维结构和/或功能损失或功能获得的变化。在一些实施方案中,靶标RNA中靶标A的脱氨基作用改变了靶标RNA的一种或多种下游分子(例如,蛋白质,RNA和/或代谢物)的表达水平。所述下游分子的表达水平的变化可以是表达水平的增加或减少。
术语“RNA编辑实体”是指可以引起核苷酸的化学修饰以将核苷酸改变为不同核苷酸的生物分子。在一些实施方案中,可以将RNA编辑实体募集到多核苷酸中的特定位点以引起所需位点处核酸序列的改变。RNA编辑实体的例子包括APOBEC蛋白(例如,APOBEC1、APOBEC2、APOBEC3A、APOBEC3B、APOBEC3C、APOBEC3E、APOBEC3F、APOBEC3G、APOBEC3H或APOBEC4蛋白)或ADAR蛋白(例如,ADAR1、ADAR2或ADAR3蛋白)。
在本申请的上下文中,“靶RNA”是指将脱氨酶招募RNA序列设计为与其具有完全互补性或基本互补性的RNA序列,并且靶标序列与dRNA之间杂交形成含有靶标腺苷的双链RNA(dsRNA)区域,其招募作用于RNA的腺苷脱氨酶(ADAR),该酶使靶标腺苷脱氨基。在一些实施方案中,ADAR天然地存在于宿主细胞中,例如真核细胞(优选哺乳动物细胞,更优选人类细胞)。在一些实施方案中,将所述ADAR引入宿主细胞中。
如本文所用的术语“条形码序列”是指将唯一性核苷酸序列用于在反应中鉴定和/或跟踪多核苷酸的来源。条形码序列的大小和组成可以差异巨大。在特定实施方案中,条形码序列的长度范围可为4至36个核苷酸、或6至30个核苷酸或8至20个核苷酸。
术语“载体”是指能够转运与其连接的另一种核酸的核酸分子。所述载体包括但不限于单链,双链或部分双链的核酸分子;包含一个或多个游离末端的,没有游离末端的(例如环状)核酸分子;包含DNA,RNA或两者的核酸分子;以及本领域已知的其他多种多核苷酸。一种类型的载体是“质粒”,它是指环状的双链DNA环,其中可以例如通过标准分子克隆技术插入额外的DNA区段。某些载体能够在引入它们的宿主细胞中自主复制(例如,具有细菌复制起点的细菌载体和附加型哺乳动物载体)。其他载体(例如,非附加型哺乳动物载体)在引入宿主细胞后整合到宿主细胞的基因组中,由此与宿主基因组一起复制。此外,某些载体能够指导它们可操作地连接的编码核苷酸序列的转录或表达。此类载体在本文中称为“表达载体”。
编辑实体通常在性质上是蛋白质的,比如在后生动物包括哺乳动物中发现的ADAR酶。编辑实体还可以包括核酸和蛋白质或肽的复合物,比如核糖核蛋白。编辑酶可以仅包括核酸或由核酸组成,比如核酶。所有这些编辑实体都包括在本发明中,只要它们是由根据本发明的寡核苷酸构建体募集的。优选地,编辑实体是酶,更优选为腺苷脱氨酶或胞苷脱氨酶,还更优选为腺苷脱氨酶。当编辑实体是腺苷脱氨酶时,Y优选为胞苷或尿苷、最优选胞苷。最感兴趣的之一是人ADAR,hADAR1和hADAR2,包括其任何亚型,比如hADAR1p110和p150。
靶RNA可以是任何细胞的或病毒的RNA序列,但更通常是具有蛋白质编码功能的mRNA前体或mRNA。
材料和方法
1、人源ADAR1表达质粒构建
利用PCR将人源ADAR1的CDS(人源ADAR1 CDS序列如SEQ:ID NO:2所示)扩增出来并加上同源臂,利用NheI和EcoRI双酶切把pcDNA3.1载体(序列如SEQ ID NO:1所示)线性化并用同源重组的方法把插入片段连接到载体。
2、报告基因质粒构建
首先将荧光基因GFP全长及需要编辑的基因片段通过PCR扩增出来同时两端加上同源重组的同源序列,同时对pcDNA3.1空载进行NheI和HindIII双酶切并回收线性载体,通过多片段重组对载体及插入片段进行同源重组完成报告基因的载体构建。
3、gRNA序列池质粒文库构建
通过PCR把gRNA混合序列池扩增出来同时加上同源臂,把报告基因载体用BamHI及EcoRI进行双酶切线性化并用同源重组的方式把gRNA序列池插入到报基因载体,转化涂板后把所有单克隆收集到一起进行质粒提取,最终得到gRNA序列池质粒文库。
4、细胞培养
4.1细胞传代实验步骤:
(1)紫外照射超净工作台消毒30分钟。
(2)从培养箱中取出待传代的细胞,先用PBS清洗三遍,要轻轻地加入轻轻地吸弃,防止细胞脱落,再加入可淹没细胞表面量的0.25%胰酶,待细胞变圆后加入含有血清的培养基终止消化。
(3)把消化后的细胞悬液转移到15ml离心管中,室温1000rpm离心4分钟。
(4)小心吸弃上清液,加入新鲜的培养基并把细胞吹打成单悬浮。
(5)取20ul细胞悬液加入到96孔细胞培养板并再加入20ul台盼蓝染色液,吹打混匀后吸取20ul加入到细胞计数板并在显微镜下进行细胞计数。
(6)根据接种的细胞培养板大小计算特定细胞数所需的细胞悬液体积,并加入对应体积的新鲜细胞培养基,放入细胞培养箱培养。
4.2细胞转染实验步骤:
(1)根据实验需求细胞量设定所需要的细胞培养板,根据细胞传代的方法在对应的培养板中加入一定的培养基及细胞量,在细胞生长24小时可达到70%汇合度时则可进行细胞转染实验。
(2)紫外照射超净工作台,消毒30分钟。
(3)根据转染试剂lipofectamine 3000提供的说明书配制转染试剂混合液,室温静置5分钟。
(4)取出细胞并把转染液小心加入到细胞中,轻轻混合均匀后放入培养箱。
4.3 ADAR1过表达HEK293细胞株构建
按照上述的细胞传代及转染步骤把ADAR1表达载体转染到HEK293细胞,48小时后加入筛选药物G418,每两天换一次液,待未转染组已无活细胞完成稳转ADAR1的HEK293细胞株筛选。
4.4靶向测序文库构建及质检
细胞总RNA可以用常规RNA提取试剂盒进行提取,确保RNA无明显降解,用常规逆转录试剂盒进行逆转录反应,在第一轮PCR扩增时需在扩增引物两端加上测序接头序列,且控制第一轮扩增的循环数低于20个循环,琼脂糖凝胶电泳对第一轮PCR产物进行验证并回收对应大小的文库片段,纯化后再进行第二轮PCR反应,第二轮引物带有测序接头序列及特异性的碱基条形码序列,第二轮PCR一般控制在10个循环以内。再次对PCR产物进行琼脂糖凝胶电泳确认文库片段大小并切胶回收纯化文库。最后对靶向测序文库进行二代测序。
4.4易错PCR(Error prone PCR)扩增
利用Agilent公司的二代易错PCR试剂盒对特定区域进行扩增并随机引入突变,根据说明书进行样品配制,其中模板加入量为0.1ng,扩增循环数为30个循环。
4.5靶向测序数据分析
首先对原始数据用cutadapt软件进行去接头处理,再把去接头后的原始读长比对到参考序列,根据连接子序列把报告基因的编辑与每条gRNA进行一一对应,再对报告基因的编辑水平进行计算,从而得出每条gRNA对靶位点的编辑水平。
4.6化学修饰后的RNA合成
修饰后的RNA合成主要由领域内的常规固相合成完成,其中碱基之间会有硫代磷酸化的骨架修饰和碱基上的2’位修饰。在一些实施方案中选择了硫代磷酸化修饰(下文用*代替),在一些实施方案中选择了2’位LNA修饰(下文用L代替)。硫代磷酸化修饰和LNA修饰的结构式分别如下所示:
硫代磷酸化修饰
LNA修饰
实施例1在HEK293细胞中筛选GAPDH基因上的不同A碱基位点的gRNA序列
本实施例以编辑人内源基因GAPDH(NM_002046.7)的多个A碱基位点的gRNA筛选方法为例,进一步说明本发明,具体地,所述方法包括如下步骤:
1.1对人源GAPDH基因特定A碱基进行gRNA文库设计及接头序列设计
以特定A碱基为中心上下游延伸75nt设计一条151nt的gRNA,在此gRNA上引入不同数量及位置的突变从而得到待筛选的文库,本实施例包含四个位点的A碱基,所在的三联基序分别为AAC,AAU,AAG和UAA,对该四个位点的多条gRNA文库进行筛选,本文举例列出了该四个位点中的AAC基序的文库序列信息,如SEQ ID NO:1~SEQ ID NO:540所示。
序列合成可以通过设计接头序列加以区分,该接头序列在下游扩增步骤中不会造成不同组别间的交叉,四个文库对应的编辑位点信息及接头序列如下表1所示。
表1编辑位点信息及接头序列
1.2筛选所需载体的设计及构建
如上述材料方法内容所述,在本实施例中,把编辑位点上下游100bp序列(共201个碱基)加入到GFP基因的下游重组到pcDNA3.1载体中,再在下游插入连接子1序列(序列如SEQ ID NO:549),在本实施例中使用的连接子1长度为2121bp,获得报告基因质粒载体,其中连接子1包含一段10bp的NNNYRNNNYR(Barcode序列,其中N表示A、C、G、T四种碱基中的一种,Y表示C或T碱基,R表示A或G碱基)简并碱基用于把每个gRNA的编辑效率区分开。得到插入了连接子1的载体后再把扩增好的gRNA文库重组进报告基因质粒载体中,获得筛选载体如图2所示。所述载体表达出的转录本进行RNA编辑模式如图3所示。
1.3接着将获得的筛选载体转入到过表达ADAR1的HEK293细胞中,48小时后提取RNA及靶向建库测序(建库引物如表2所示)并分析每条gRNA的靶位点编辑及脱靶编辑水平情况,如图4、图5所示。在本实施例中加入了长的连接子1序列,目的是为了模拟分子间RNA之间形成双链进行编辑,在建库时构建两种文库,一种是Reporter与Linker的区域文库构建每个Linker上的barcode序列与编辑水平的对应关系,所用的第一轮引物为cDNA-NGS-round1引物对,第二轮引物为NGS-round2引物对;另一种是Linker与gRNA库的区域,因为Linker与gRNA中间间隔超过2000bp无法直接建库测序,故先对Linker与gRNA的全长序列进行环化,环化后用plasmid-NGS-round1引物对进行第一轮扩增,再用NGS-round2引物对完成整个文库构建,此文库可建立Linker上的barcode与gRNA序列的对应关系,通过两个文库的barcode序列比对则可得出每条gRNA对应的编辑效率。
结果显示,通过本发明的实验及数据分析方法可以筛选出每条gRNA对应的编辑效率,且生物学重复比较一致。编辑位点1为AAC基序(图4a),黑色的点表示从编辑位点往两端延伸共151nt的gRNA(编辑位点对位碱基为C碱基)在两次重复实验中的编辑水平,其余黄色的点为引入突变后的不同gRNA在两次重复实验中的编辑水平,针对AAC基序的gRNA在50%以内的编辑效率有均匀的gRNA分布,在50%到75%的编辑效率区间仅有少些gRNA能够达到。通过热图(图5a)分析可知在本实施例设计的针对AAC基序的gRNA在多个位点有明显的脱靶编辑但靶位点的编辑水平整体高于脱靶位点。
编辑位点2的AAU基序(图4b)gRNA整体表现出低于40%的编辑效果,且从编辑位点往上下游延伸的除了编辑位点处是AC的错配,其他都是互补配对的gRNA编辑效率也低于30%。
AAG基序(图4c)的gRNA编辑效率与AAC基序类似,与其他三个编辑位点不同的是UAA基序(图4d)的gRNA整体编辑水平位于25%到75%之间,但UAA基序的gRNA邻近的A碱基有比较严重的脱靶编辑,只有AAG基序筛选的gRNA整体脱靶编辑较低(图5)。本实施例设计的gRNA长度为151nt左右,此外,不同的位点A对应的gRNA具有不同的编辑能力,表明通过本发明可筛选出编辑水平不同的gRNA序列。
表2建库测序所用引物序列
实施例2在HEK293细胞中筛选GAPDH基因特定位点基序为UAG的gRNA序列
本实施例旨在验证在人源GAPDH基因上可以筛选出针对特定位点进行编辑的短gRNA及短连接子。选取GAPDH基因上787位点A碱基,所在的基序为TAG基序,针对该位点设计了1073条包含不同长度及不同位点突变的gRNA(长度范围28nt-89nt)。在本实施例选择用短连接子验证筛选体系,其中连接子2的序列为NNNNNNNNTGGGTTGAGGGTAGTGAG(序列号为SEQ ID NO:556;N为A,T,C,G碱基中的任一种,8个N为连接子的barcode序列)。
具体的实施方案如下,首先直接合成gRNA文库,把连接子2与gRNA序列直接相连合成完成带有barcode序列的gRNA库;其次构建reporter载体,以编辑位点为中心上下游延伸共201bp的序列与GFP序列重组插入到pcDNA3.1载体;第三步把步骤一的gRNA文库插入到reporter载体完成筛选载体的构建;第四步把筛选载体转染到ADAR1蛋白过表达的HEK293细胞48小时后提取RNA及靶向建库测序并分析每条gRNA的靶位点编辑及脱靶编辑水平情况,结果如图6所示。结果表明在本实施例中在短连接子的情况下不同长度的gRNA混合的文库也可以通过此方法筛选出不同编辑效率的gRNA。
基于本发明的方法建立的高通量筛选是向细胞中转染混合的gRNA文库,可能会存在一定程度的gRNA之间的干扰情况,故本实施例在高通量筛选的基础上进行单个gRNA的验证,本实施例挑选了高通量筛选出编辑水平从3.93%到21.66%的gRNA序列,序列信息如表3的GAPDH-TAG-S1至GAPDH-TAG-S12,化学修饰后的gRNA序列对应为GAPDH-TAG-M1至GAPDH-TAG-M12。针对高通量筛选出来的序列则用化学修饰后的单个RNA分别转染进入细胞验证编辑水平情况(如图8及表3所示),修饰后的gRNA稳定性高于质粒表达出的gRNA,故编辑水平均高于高通量筛选的编辑水平,化学修饰的gRNA修饰策略为碱基之间用硫代磷酸修饰,两端各加3个LNA修饰。分析高通量筛选出的编辑水平与单个修饰gRNA的编辑水平可知二者呈现较强的相关性,相关系数达到0.83(如图7所示),证实了本发明的方法的准确性。其中表3中的GAPDH-TAG-S1至GAPDH-TAG-S13的编辑水平为高通量的gRNA混合体系测定出来的gRNA编辑水平;GAPDH-TAG-M1至GAPDH-TAG-M13的编辑水平为单个gRNA体系测定出来的编辑水平。
表3高通量筛选验证序列信息表
实施例3在HEK293细胞中筛选ATP7B基因特定位点基序为UAG的gRNA序列
本实施例旨在验证在人源ATP7B(NM_000053.4)基因上可以筛选出针对特定位点进行编辑的短gRNA。ATP7B基因突变是引起肝豆状核变性疾病的主要原因,在本实施例选取ATP7B基因上1532位点A碱基,在该位点上游有一个碱基的突变C碱基突变为T碱基,故突变后编辑位点所在的基序为UAG基序,针对该位点设计了1011条不同长度及不同位点突变的gRNA,本实施例选取包含编辑位点及其上下游在内的长度为31nt和41nt的gRNA,针对这两个长度引入突变或者缺失或者增加了凸起等会使得实际长度有一点的浮动,浮动范围为27nt-59nt。在本实施例选择用连接子2验证筛选体系。
具体的实施方案如下,首先直接合成长度为31nt和41nt的gRNA文库,把连接子2与gRNA序列直接相连合成完成带有barcode序列的gRNA库;其次构建reporter载体,以编辑位点为中心上下游延伸共201bp的序列与GFP序列重组插入到pcDNA3.1载体;第三步把步骤一的gRNA文库插入到reporter载体完成筛选载体的构建;第四步把筛选载体转染到ADAR1蛋白过表达的HEK293细胞48小时后提取RNA及靶向建库测序并分析每条gRNA的靶位点编辑及脱靶编辑水平情况,结果如图9和图10所示。
结果表明,通过本实施例可以从1011条不同长度的gRNA文库中筛选出高编辑效率的gRNA序列,整体上41nt的gRNA编辑能力高于31nt的gRNA(图9),其中黑色的点表示除了编辑位点处有一个碱基的错配外,其他碱基都是互补配对的gRNA。
把所有不同长度的gRNA排列在一起比较靶位点的编辑效率及脱靶效率(图10)可知,在ATP7B 1532位点的A碱基设计的gRNA有比较弱的脱靶编辑。整体比较来看基于41nt的靶序列设计的gRNA大部分序列的编辑水平低于20%,基于31nt靶序列设计的gRNA全部低于20%的编辑。
本实施例从高通量筛选的两组结果中各挑选了20条编辑效率高且测序测的读长多的序列进行下一轮的error prone验证引入突变后的编辑效果,从41nt组的挑选出的序列如SEQ ID NO:583~SEQ ID NO:602所示,31nt组的挑选序列如SEQ ID NO:603~SEQ ID NO:622所示。在所有的挑选序列两端分别加上error prone PCR的引物序列5’端为TCGAAGGTCGCTTAGACGGC(SEQ ID NO:581),3’端为CGGTCGTGAGTGCAGACGTG(SEQ ID NO:582)直接合成单链序列,在error prone PCR扩增时分两组进行扩增,41nt组则把所有序列等质量混合后作为模板扩增,31nt组同样处理,扩增后的产物重新插入到reporter载体中按照第一轮的筛选进行第二轮筛选,结果如图11所示。
结果表明无论是41nt组还是31nt组在经过第二轮的筛选后都达到了显著的提升,31nt组从第一轮的全部低于20%的编辑效率提升到20%至80%的编辑效果,主要分布在40%的编辑效率以内,40%到80%的编辑占比较少,而41nt组20%到50%编辑效率的占比也是少的,在经过第二轮的筛选后明显可以看到40%到90%的编辑占比比较高,通过以上数据也证实了本发明的两轮筛选可以筛选出比较高的编辑效果的gRNA,尤其是第二轮error prone引入突变的方法。
实施例4在HEK293细胞中筛选FGFR基因特定位点基序为CAG的gRNA序列
本实施例旨在验证在人源FGFR(NM_015850.4)基因上可以筛选出针对特定位点进行编辑的短gRNA。FGFR基因突变是引起软骨发育不全、致死性骨发育不良及矮小症等疾病的病因,在本实施例选取FGFR基因上G358R位点产生了G到A的突变,故突变后编辑位点所在的基序为CAG基序,针对该位点设计了989条不同长度及不同位点突变的gRNA,本实施例选取包含编辑位点及其上下游在内的长度为31nt和41nt的gRNA,针对这两个长度引入突变或者缺失或者增加了凸起等会使得实际长度有一点的浮动,浮动范围为23nt-62nt。
在本实施例选择用连接子2验证筛选体系。具体的实施方案如下,首先直接合成31nt和41nt的gRNA文库,把连接子2与gRNA序列直接相连合成完成带有barcode序列的gRNA库;其次构建reporter载体,以编辑位点为中心上下游延伸共201bp的序列与GFP序列重组插入到pcDNA3.1载体;第三步把步骤一的gRNA文库插入到reporter载体完成筛选载体的构建;第四步把筛选载体转染到ADAR1蛋白过表达的HEK293细胞48小时后提取RNA及靶向建库测序并分析每条gRNA的靶位点编辑及脱靶编辑水平情况,结果如图12和图13所示。
结果表明,通过本实施例可以从989条不同长度的gRNA文库中筛选出高效编辑gRNA序列,整体上41nt的gRNA编辑能力高于31nt的gRNA(图12),其中黑色的点表示除了编辑位点处有一个碱基的错配外,其他碱基都是互补配对的gRNA,把所有不同长度的gRNA排列在一起比较靶位点的编辑效率及脱靶效率(图13)可知,在所有的gRNA中在编辑位点上游有两个位点的明显脱靶,但在编辑位点下游完全没有脱靶编辑。整体比较来看基于41nt的靶序列设计的gRNA大部分序列的编辑水平低于20%,基于31nt靶序列设计的gRNA几乎都低于10%的编辑。
鉴于本实施例31nt组整体编辑水平偏低,故仅从41nt的gRNA组高通量筛选结果挑选了20条编辑水平高且测序测的读长多的序列进行下一轮的error prone验证引入突变后的编辑效果,41nt组的挑选序列如SEQ ID NO:623~SEQ ID NO:642所示。在所有的挑选序列两端分别加上error prone PCR的引物序列5端为TCGAAGGTCGCTTAGACGGC(SEQ ID NO:581),3端为CGGTCGTGAGTGCAGACGTG(SEQ ID NO:582)直接合成单链序列,在error prone PCR扩增时把所有
序列等质量混合后作为模板扩增,扩增后的产物重新插入到reporter载体中按照第一轮的筛选进行第二轮筛选,结果如图14所示。
结果表明,在经过errror prone引入随机突变的第二轮的筛选后都达到了显著的提升,从第一轮的20%以内提升到80%的编辑效果,通过以上数据再次证实了本发明的两轮筛选可以筛选出比较高的编辑效果的gRNA,尤其是第二轮error prone引入突变的方法。
Claims (10)
- 一种编辑RNA的向导RNA的筛选方法,其特征在于,包括步骤:(1)待筛选的每条向导RNA连有连接子,并分别与靶RNA连接在能够形成一个转录本的核酸链上,形成多个所述的核酸链;所述的核酸链含有条形码序列;(2)将所述的多个核酸链转入表达ADAR蛋白的细胞中;(3)提取和测定所述细胞中的RNA,确定靶RNA的被编辑水平,通过条形码序列确定对应的向导RNA的编辑水平。
- 如权利要求1所述的筛选方法,其特征在于,所述的连接子位于靶RNA和向导RNA之间;优选地,所述的条形码序列包含于所述连接子中;所述核酸链从上游到下游依次连接有报告基因核酸片段、靶RNA核酸片段、连接子核酸片段、向导RNA核酸片段;优选地,所述向导RNA核酸片段的下游还连接有polyA尾核酸片段;优选地,所述报告基因选自yfp基因、gus基因、rfp基因、gfp基因、cfp基因、bfp基因、或卡那霉素抗性基因;优选地,所述连接子的长度为26nt~2121nt。优选地,所述连接子包括SEQ ID NO:549或SEQ ID NO:557;优选地,步骤(3)包括建立靶向测序文库进行二代测序;优选地,所述核酸链接触ADAR蛋白通过将所述核酸链转入表达ADAR蛋白的细胞中实现;优选地,所述ADAR蛋白包括ADAR1或ADAR2;优选地,所述向导RNA能够在所述细胞中招募ADAR蛋白对靶RNA进行编辑;优选地,所述的细胞过表达ADAR蛋白;优选地,所述编辑为A-to-I编辑;优选地,所述向导RNA为以A碱基为中心上下游延伸得到的向导RNA;优选地,所述向导RNA为以A碱基为中心上下游延伸15nt~75nt得到的30nt~151nt长度的向导RNA;优选地,所述向导RNA为完全互补配对或不完全互补配对的序列;优选地,所述向导RNA的长度为20nt~200nt;优选地,所述筛选方法包括把编辑位点上下游序列加入报告基因的下游并插入到载体中获得报告基因载体,再把报告基因载体转入细胞中;优选地,所述编辑位点上下游加入到报告基因的序列长度为20~200nt;优选地,所述筛选方法为高通量筛选方法;优选地,所述靶RNA序列长度大于向导RNA序列;优选地,所述载体包括质粒或病毒载体;优选地,所述载体是用于在高等真核细胞或原核细胞中表达的质粒或病毒载体;优选地,所述载体骨架为pcDNA3.1载体;优选地,所述目的向导RNA为靶向位点进行高效编辑的向导RNA;优选地,所述方法还包括:从所述步骤(3)中获得所述向导RNA的编辑水平后,挑选编辑水平高的向导RNA合成序列相同的单链DNA后设计易错PCR引物对所述向导RNA的靶向区域进行随机引入突变,再重复所述步骤(1)~(3);优选地,所述易错PCR引物的序列如SEQ ID NO:581和SEQ ID NO:582所示。
- 一种编辑RNA的向导RNA的筛选方法,其特征在于,包括:(a)向导RNA文库的设计;(b)把需要编辑的基因片段与报告基因插入到载体中获得报告基因载体;(c)把所述向导RNA文库插入进所述报告基因载体获得待筛选质粒文库;(d)将所述待筛选质粒文库转入到细胞,转染后,提取所述细胞中的RNA,测序;(e)使用生物信息学分析所述向导RNA的编辑水平;其中(a)和(b)的顺序可以调换;优选地,步骤(b)包括:把编辑位点上下游序列加入报告基因的下游插入到载体中获得报告基因载体;优选地,所述编辑位点上下游加入到报告基因的序列长度为20~200nt;优选地,步骤(b)包括:将荧光基因全长以及需要编辑的基因片段通过PCR扩增出来同时两端加上同源重组的同源序列或引入酶切位点,同时对空载体进行NheI和HindIII双酶切并回收线性载体,将多片段插入到线性载体中获得报告基因的载体;优选地,步骤(c)包括:扩增步骤(a)中的所述向导RNA同时加上同源臂或酶切位点获得向导RNA序列文库,步骤(b)中的报告基因载体用BamHI及EcoRI进行双酶切线性化;然后把所述的向导RNA序列文库插入到线性化后的报基因载体中,转化后获得待筛选质粒文库;优选地,所述筛选方法还包括:(f)从所述向导RNA中挑选出编辑效率高的向导RNA直接合成序列相同的单链DNA后,设计易错PCR引物对所述向导RNA靶向区域进行随机引入突变,再重复步骤(b)~(e)。
- 权利要求1-3任一所述的筛选方法获得的向导RNA。
- 一种向导RNA,其特征在于,所述向导RNA的序列如SEQ ID NO:557、SEQ ID NO:559、SEQ ID NO:561、SEQ ID NO:563、SEQ ID NO:565、SEQ ID NO:567、SEQ ID NO:569、SEQ ID NO:571、SEQ ID NO:573、SEQ ID NO:575、SEQ ID NO:577、SEQ ID NO:579、SEQ ID NO:583~SEQ ID NO:602、SEQ ID NO:603~SEQ ID NO:622、或SEQ ID NO:623~SEQ ID NO:642任一所示。
- 如权利要求1-3任一所述的筛选方法或权利要求4-5任一所述的向导RNA在制备治疗疾病的药物中的用途;优选地,所述药物用于在真核细胞中定点编辑靶基因中的核苷酸从而达到治疗疾病的目的;优选地,所述疾病为包括GAPDH基因、ATP7B基因或FGFR基因突变引起的疾病;优选地,所述GAPDH基因特定位点基序为AAC、AAU、AAG、UAA或UAG;优选地,所述ATP7B基因特定位点基序为UAG;优选地,所述FGFR基因特定位点基序为CAG;优选地,所述疾病包括肝豆状核变性、软骨发育不全、致死性骨发育不良或矮小症;优选地,所述用途是通过天然存在于所述细胞中并且能够进行所述核苷酸的编辑的RNA编辑实体的作用。
- 一种工程化RNA,其特征在于,包括:从上游到下游依次连接有报告基因核酸片段、靶RNA核酸片段、连接子核酸片段、向导RNA核酸片段;优选地,所述向导RNA核酸片段的下游还连接有polyA尾核酸片段;优选地,所述报告基因选自yfp基因、gus基因、rfp基因、gfp基因、cfp基因、bfp基因、或卡那霉素抗性基因;优选地,所述连接子的长度为26nt至2121nt;优选地,所述连接子包括SEQ ID NO:549或SEQ ID NO:557;优选地,所述向导RNA为以A碱基为中心上下游延伸得到的向导RNA;优选地,所述向导RNA为以A碱基为中心上下游延伸15nt~75nt得到的30nt~151nt长度的向导RNA。
- 一种工程化RNA结构文库,其特征在于,包含如权利要求7所述的多个不同的工程化RNA;优选地,所述不同的工程化RNA是通过将靶RNA通过不同的连接子与不同的向导RNA连接获得的;优选地,所述工程化RNA结构文库中,所述文库中,所述的工程化RNA所含有的靶RNA核酸片段是相同的,向导RNA核酸片段序列是不同的。
- 一种载体,其特征在于,包含权利要求7所述的工程化RNA或权利要求8所述的工程化RNA结构文库;优选地,所述载体是通过把靶RNA序列加入到报告基因序列的下游插入到载体骨架中,获得报告基因载体,再把扩增好的向导RNA文库插入报告基因载体中获得的。
- 一种细胞,其特征在于,包含权利要求7所述的工程化RNA或权利要求8所述的工程化RNA结构文库或权利要求9所述的载体。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/085605 WO2024197857A1 (zh) | 2023-03-31 | 2023-03-31 | 一种筛选向导rna的方法 |
| CN202380099852.8A CN121586774A (zh) | 2023-03-31 | 2023-03-31 | 一种筛选向导rna的方法 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2023/085605 WO2024197857A1 (zh) | 2023-03-31 | 2023-03-31 | 一种筛选向导rna的方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024197857A1 true WO2024197857A1 (zh) | 2024-10-03 |
Family
ID=92903150
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2023/085605 Ceased WO2024197857A1 (zh) | 2023-03-31 | 2023-03-31 | 一种筛选向导rna的方法 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121586774A (zh) |
| WO (1) | WO2024197857A1 (zh) |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119506359A (zh) * | 2025-01-20 | 2025-02-25 | 中国医学科学院基础医学研究所 | 新型单细胞crispr筛选文库 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111613272A (zh) * | 2020-05-21 | 2020-09-01 | 西湖大学 | 程序化框架gRNA及其应用 |
| WO2022087272A1 (en) * | 2020-10-21 | 2022-04-28 | The Board Of Trustees Of The Leland Stanford Junior University | A screening platform for adar-recruiting guide rnas |
| WO2022147573A1 (en) * | 2021-01-04 | 2022-07-07 | The Regents Of The University Of California | Programmable rna editing in vivo via recruitment of endogenous adars |
-
2023
- 2023-03-31 WO PCT/CN2023/085605 patent/WO2024197857A1/zh not_active Ceased
- 2023-03-31 CN CN202380099852.8A patent/CN121586774A/zh active Pending
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN111613272A (zh) * | 2020-05-21 | 2020-09-01 | 西湖大学 | 程序化框架gRNA及其应用 |
| WO2022087272A1 (en) * | 2020-10-21 | 2022-04-28 | The Board Of Trustees Of The Leland Stanford Junior University | A screening platform for adar-recruiting guide rnas |
| WO2022147573A1 (en) * | 2021-01-04 | 2022-07-07 | The Regents Of The University Of California | Programmable rna editing in vivo via recruitment of endogenous adars |
Cited By (1)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN119506359A (zh) * | 2025-01-20 | 2025-02-25 | 中国医学科学院基础医学研究所 | 新型单细胞crispr筛选文库 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121586774A (zh) | 2026-02-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US11306308B2 (en) | High-throughput CRISPR-based library screening | |
| Choudhuri | Bioinformatics for beginners: genes, genomes, molecular evolution, databases and analytical tools | |
| EP2705152B1 (en) | Multiplexed genetic reporter assays and compositions | |
| Raitskin et al. | Comparison of efficiency and specificity of CRISPR-associated (Cas) nucleases in plants: An expanded toolkit for precision genome engineering | |
| US12049623B2 (en) | Compositions and methods for identifying polynucleotides of interest | |
| JP2022513529A (ja) | バーコード付きガイドrna構築体を使用する効率的な遺伝子スクリーニングのための組成物及び方法 | |
| CN109072258A (zh) | 复制型转座子系统 | |
| CN108424907A (zh) | 一种高通量dna多位点精确碱基突变方法 | |
| CN119842677A (zh) | 结构导向挖掘的具有高编辑活性和无序列偏好性的胞嘧啶脱氨酶及其应用 | |
| WO2018089437A1 (en) | Compositions and methods for scarless genome editing | |
| CA3259919A1 (en) | METHODS AND COMPOSITIONS FOR NUCLEIC ACID SEQUENCE | |
| CN120173914A (zh) | 基于IscB系统的紧凑型基因组编辑器和碱基编辑器及其应用 | |
| WO2024197857A1 (zh) | 一种筛选向导rna的方法 | |
| Pizzollo et al. | Differentially active and conserved neural enhancers define two forms of adaptive noncoding evolution in humans | |
| WO2021053208A1 (en) | Methods for dna library generation to facilitate the detection and reporting of low frequency variants | |
| CN116949011A (zh) | 经分离的Cas13蛋白、基于它的基因编辑系统及其用途 | |
| CN113249362B (zh) | 经改造的胞嘧啶碱基编辑器及其应用 | |
| CN118995709A (zh) | 一种提高基因编辑效率的重组型pegRNA及其应用 | |
| CN118737260A (zh) | 一种获得adar底物候选序列结构的方法 | |
| Shi et al. | HIT-scISOseq: High-throughput and High-accuracy Single-cell Full-length Isoform Sequencing | |
| TW202603170A (zh) | 促進rna轉譯的utr | |
| WO2024261215A1 (en) | High throughput screens of translation and stability of mrna using barcoded xrna display and its variations | |
| CN121406716A (zh) | 一种核糖体病斑马鱼模型的构建方法 | |
| Hussain et al. | Recombinant DNA Technology | |
| CN121343986A (zh) | 特异性靶向鸡SATB1基因的sgRNA序列及其免疫应用 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 23929442 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 23929442 Country of ref document: EP Kind code of ref document: A1 |