EP4705449A1 - High-throughput screening and sequencing method - Google Patents
High-throughput screening and sequencing methodInfo
- Publication number
- EP4705449A1 EP4705449A1 EP24720464.7A EP24720464A EP4705449A1 EP 4705449 A1 EP4705449 A1 EP 4705449A1 EP 24720464 A EP24720464 A EP 24720464A EP 4705449 A1 EP4705449 A1 EP 4705449A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- dna
- dna modifying
- region
- modifying enzyme
- sequencing
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/10—Processes for the isolation, preparation or purification of DNA or RNA
- C12N15/1034—Isolating an individual clone by screening libraries
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/87—Introduction of foreign genetic material using processes not otherwise provided for, e.g. co-transformation
- C12N15/90—Stable introduction of foreign DNA into chromosome
- C12N15/902—Stable introduction of foreign DNA into chromosome using homologous recombination
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N9/00—Enzymes; Proenzymes; Compositions thereof; Processes for preparing, activating, inhibiting, separating or purifying enzymes
- C12N9/14—Hydrolases (3)
- C12N9/16—Hydrolases (3) acting on ester bonds (3.1)
- C12N9/22—Ribonucleases [RNase]; Deoxyribonucleases [DNase]
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/34—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving hydrolase
- C12Q1/44—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving hydrolase involving esterase
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/48—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving transferase
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12Q—MEASURING OR TESTING PROCESSES INVOLVING ENZYMES, NUCLEIC ACIDS OR MICROORGANISMS; COMPOSITIONS OR TEST PAPERS THEREFOR; PROCESSES OF PREPARING SUCH COMPOSITIONS; CONDITION-RESPONSIVE CONTROL IN MICROBIOLOGICAL OR ENZYMOLOGICAL PROCESSES
- C12Q1/00—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions
- C12Q1/68—Measuring or testing processes involving enzymes, nucleic acids or microorganisms; Compositions therefor; Processes of preparing such compositions involving nucleic acids
- C12Q1/6869—Methods for sequencing
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N2500/00—Screening for compounds of potential therapeutic value
- G01N2500/04—Screening involving studying the effect of compounds C directly on molecule A (e.g. C are potential ligands for a receptor A, or potential substrates for an enzyme A)
-
- G—PHYSICS
- G01—MEASURING; TESTING
- G01N—INVESTIGATING OR ANALYSING MATERIALS BY DETERMINING THEIR CHEMICAL OR PHYSICAL PROPERTIES
- G01N2500/00—Screening for compounds of potential therapeutic value
- G01N2500/10—Screening for compounds of potential therapeutic value involving cells
Landscapes
- Life Sciences & Earth Sciences (AREA)
- Health & Medical Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Organic Chemistry (AREA)
- Engineering & Computer Science (AREA)
- Genetics & Genomics (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Biotechnology (AREA)
- General Engineering & Computer Science (AREA)
- Molecular Biology (AREA)
- Biochemistry (AREA)
- Microbiology (AREA)
- Biomedical Technology (AREA)
- General Health & Medical Sciences (AREA)
- Proteomics, Peptides & Aminoacids (AREA)
- Physics & Mathematics (AREA)
- Biophysics (AREA)
- Immunology (AREA)
- Analytical Chemistry (AREA)
- Plant Pathology (AREA)
- Medicinal Chemistry (AREA)
- Mycology (AREA)
- Crystallography & Structural Chemistry (AREA)
- Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
Abstract
The present invention is in the field of DNA modifying enzymes and provides a high- throughput method for characterizing and sequencing multiple DNA modifying enzymes or target sites for DNA modifying enzymes. The present invention provides a method for screening and sequencing of a plurality of DNA modifying enzymes, the method comprising the steps of providing a library of expression vectors comprising a plurality of different vectors, wherein each vector comprises a first region encoding a variant of a DNA modifying enzyme, and a second region comprising one or more target sites of a DNA modifying enzyme; introducing the library of expression vectors into host cells; culturing the host cells and expressing the DNA modifying enzymes; isolating plasmid DNA from the culture of host cells; sequencing the first and the second region of the expression vectors; and determining, based on the sequencing results, whether the DNA sequence of the second region on the expression vector has been altered by the DNA modifying enzyme. The present invention further provides a method for screening and sequencing of a plurality of target sites for DNA modifying enzymes.
Description
High-throughput screening and sequencing method
The present invention relates to a high-throughput method for characterizing and sequencing multiple DNA modifying enzymes or target sites for DNA modifying enzymes.
Background of the invention
DNA modifying enzymes and in particular site-specific recombinase (SSR) systems allow precise manipulation of DNA without triggering endogenous DNA repair pathways. They possess the unique ability to fulfill both cleavage and immediate rescaling of the processed DNA in vivo.
Tyrosine-type site-specific recombinases (Y-SSRs) are widely used genome engineering tools. In addition to the Cre/loxP and the Flp/FRT systems, other recombinase systems are known in the art. For example, US 7,422,889 and US 7,915,037 disclose the so-called Dre/rox system that comprises a Dre recombinase isolated from Enterobacteria phage D6, the recognition site of which is called rox- site. Further known recombinase systems are the VCre/VloxP system isolated from Vibrio plasmid p0908, and the sCre/SloxP system (WO 2010/143606 Al; Suzuki and Nakayama, 2011). Further sitespecific DNA recombinase systems are the Nigri/nox system disclosed in EP 2877585 Bl, the Vika/vox system disclosed in EP 2690177 Bl, and the Panto/pox system disclosed in EP 3263708 Bl.
Y-SSRs are capable of exchanging DNA strands between their target DNA sequences, which can facilitate controlled excision, inversion, insertion or exchange of DNA. Because this process is very precise and not reliant on DNA repair mechanisms, it offers seamless manipulation of genomic DNA without side effects (Meinke et al. 2016). These features distinguish Y-SSRs from nuclease- based genome engineering methods, such as CRISPR-Cas systems. All nuclease-based genome editing tools rely on the DNA repair pathways of the host cell, which can ultimately lead to unintended editing results (reviewed in Anzalone et al. 2020). However, the advantage of the widely used CRISPR-Cas9 system is the speed and efficiency with which it can be programmed to specifically edit novel target sites. In this regard, Y-SSRs are lagging behind and currently require substantial time and effort to be generated with tailored specificity. This bottleneck represents a considerable hurdle to harness the full potential of designer-recombinases as a versatile genome editing tool.
Current approaches to adapt site specific recombinases to novel target sequences use directed molecular evolution. One powerful approach to evolve designer Y-SSRs utilizes a plasmid-based bacterial application called substrate-linked directed evolution (SLiDE) (Buchholz and Stewart 2001; Buchholz and Hauber 2011; Lansing et al. 2019; Hoersten et al. 2021; Lansing et al. 2022).
Substrate-linked directed evolution allows libraries of recombinases to be progressively evolved by screening random mutations in conjunction with a selection scheme for activity on new, stepwise altered target sequences. Several recombinases have already been developed this way (Buchholz and Stewart 2001; Sarkar et al. 2007; Karpinski et al. 2016; Lansing et al. 2019; Lansing et al. 2022), demonstrating the applicability of substrate-linked directed evolution. Improvements of the
approach have also been reported (Lansing et al. 2019), but the generation of a new designer Y-SSR still requires substantial resources. During rounds of mutagenesis and selection, large gene variant libraries (between about 105 and 108) are produced (Lawrence et al., 2013; Rognes et al., 2016; Vaser et al., 2017). Since the screening of libraries to identify efficient variants is conventionally a manual process, it is labour and time intensive and often requires six to twelve months of work.
These disadvantages are a direct consequence of the randomized way new generations of recombinases are created when using substrate-linked directed evolution. At the start of a new project, it is not known which sequence modifications will lead to a protein variant that shows recombination activity on the predefined target sequence. There has been some success to introduce improvements based on protein model analysis on designer-recombinases with activity on a defined target site (Abi- Ghanem et al. 2012). However, due to the complex nature of the whole enzymatic reaction, which includes DNA binding, DNA bending, and catalysis (reviewed in Meinke et al. 2016), the evolution and characterization of evolved DNA modifying enzymes remains a cumbersome and time-consuming process. Additionally, the number of variants that can be tested is limited, reducing the probability of identifying optimal variants.
It is therefore an objective of the present invention to provide a high-throughput method for characterizing and sequencing multiple DNA modifying enzymes. It is a further an objective of the present invention to provide a high-throughput method for characterizing and sequencing target sites for DNA modifying enzymes.
Summary of the invention
The objective underlying the present invention is solved by the provision of the methods according to the invention.
According to a first aspect, the present invention provides a method for screening and sequencing of a plurality of DNA modifying enzymes, the method comprising the steps of: providing a library of expression vectors comprising a plurality of different vectors, wherein each vector comprises a first region encoding a variant of a DNA modifying enzyme, and a second region comprising one or more target sites of a DNA modifying enzyme; introducing the library of expression vectors into host cells; culturing the host cells and expressing the DNA modifying enzymes; isolating plasmid DNA from the culture of host cells; sequencing the first and the second region of the expression vectors; and determining, based on the sequencing results, whether the DNA sequence of the second region on the expression vector has been altered by the DNA modifying enzyme.
According to a second aspect, the present invention provides a method for screening and sequencing of a plurality of target sites for DNA modifying enzymes, the method comprising the steps of: providing a library of expression vectors comprising a plurality of different vectors, wherein each vector comprises a first region encoding a DNA modifying enzyme, and a second region comprising variants of one or more target sites of a DNA modifying enzyme; introducing the library of expression vectors into host cells; culturing the host cells and expressing the DNA modifying enzymes; isolating plasmid DNA from the culture of host cells; sequencing the first and the second region of the expression vectors; and determining, based on the sequencing results, whether the DNA sequence of the second region on the expression vector has been altered by the DNA modifying enzyme.
According to one embodiment, in the method of the invention the first region further comprises a unique molecular identifier (UMI). According to a preferred embodiment, the unique molecular identifier is an oligonucleotide comprising at least 50 random nucleotides.
According to a preferred embodiment, the unique molecular identifier is located in the first region of the expression vector adjacent to the sequence encoding the DNA modifying enzyme.
According to another embodiment, the method further comprises the steps of: clustering the unique molecular identifiers, generating and polishing consensus sequences, determining the number of DNA modification events for each DNA modifying enzyme; and determining an activity rate for each DNA modifying enzyme.
According to a further embodiment, the first and the second region of the vector are sequenced in a single step.
According to one embodiment, the sequencing of the first and the second region of the expression vectors comprises nanopore sequencing.
According to a preferred embodiment, the DNA modifying enzyme is selected from the group consisting of a recombinase, an integrase, an adenosine base editor (ABE), a zinc -finger nuclease, a transcription activator-like effector nuclease and a Cas nuclease.
According to yet another embodiment, the DNA modifying enzyme comprises more than one subunit.
According to a further embodiment, the DNA modifying enzyme comprises at least two different subunits.
According to yet another embodiment, the first and the second region of the expression vectors are excised from the expression vector before the sequencing.
Further aspects and embodiments of the invention will become apparent from the appending claims and the following detailed description.
Description of the drawings
The invention is further illustrated by the following figures and examples without being limited thereto.
Fig. 1 shows an overview of the workflow of a particularly preferred method according to the present invention. Evolved recombinases are cloned together with a unique molecular identifier (UMI) into a vector containing loxF8 target sites. E. coli cells are transformed with the designed vectors. Transformed bacteria are cultured to express the enzymes. For screening the same enzyme variant on different target sites, plasmid DNA is isolated and the DNA modifying enzymes together with their UMIs are subcloned into additional target sites (HG1, HG2, and HG2L) and cultured as before. From all plasmids isolated from the cell cultures, the region of interest (comprising regions 1 and 2) is excised and sequenced by nanopore sequencing. For further evaluation, the UMIs are clustered, allowing consensus sequence polishing and for determining the recombination events.
Fig. 2A shows the results of a DNA editing quantification sequencing screen of loxF8 recombinases on four target sites, loxF8 (SEQ ID NO: 11), HG1 (SEQ ID NO: 12), HG2 (SEQ ID NO: 13), and HG2L (SEQ ID NO: 14). UMI-clusters containing evolved recombinases are indicated as grey spots, and D7 control clusters (SEQ ID NOs: 19 and 20) are indicated as black spots. Three selected clusters are highlighted (138 (SEQ ID NOs: 21 and 22), 181 (SEQ ID NOs: 23 and 24), and 1244 (SEQ ID NOs: 25 and 26). Fig. 2B shows median recombination rates of recombinase D7. Fig. 2C shows that 52 non-D7 clusters identified with the method of the present invention had less than 10% off-target activity on the three off-targets and more than 25% activity on the on-target.
Fig. 3A is a schematic illustration of the recombination assay (Examples 5 and 10). Recombined and non-recombined plasmids (circles at the top) are digested with restriction enzymes that excise the recombinase gene(s). Due to the difference in plasmid size between the two versions, the resulting DNA fragments containing the plasmid backbone are also different in size. This is visualized with agarose gel electrophoresis (schematic at the bottom). A mixture will contain both kinds of fragments, the strength of each band corresponds to the amount of the respective fragment, making it possible to quantify the recombination rate. Pictograms are indicated at at the right of the gel scheme to indicate to which fragment each band corresponds to. Two triangles represent nonrecombined fragments, and one triangle represents recombined fragments. Fig. 3B shows respective agarose gel analysis of the recombination assays performed for four different DNA modifying enzymes (Clusters 138, 181, 1244 and the D7 control) on four different target sites loxF8 (SEQ ID NO: 11), HG1 (SEQ ID NO: 12), HG2 (SEQ ID NO: 13), and HG2L (SEQ ID NO: 14). Fig. 3C shows the results of the recombination assay (Fig. 3B) performed in triplicate replicates and quantified. The band intensities of the recombined and non-recombined products of the selected recombinase clusters (138, 181, 1244) and of the control recombinase D7 were determined using the image analysis software Fiji (Schindelin et al., 2012). Band intensity values of the recombined
products were divided by the combined values of the recombined and non-recombined bands to determine the fraction of recombined DNA, which was converted to a percentage value by multiplying with 100. Each dot represents one replicate of the assay of the respective variant.
Fig. 4 schematically shows the adapted substrate linked directed evolution (SLiDE) method used in Examples 4 and 11. The gene of a DNA editing enzyme, which performs base editing is amplified using error-prone PCR or DNA shuffling. This results in multiple copies of the gene with mutations. This gene library is then cloned into a bacterial expression vector, that contains target sites for the base editor, with restriction enzyme (RE) sites located at the position where base change is supposed to happen. The expression vector is transformed into bacteria, which express the DNA modifying enzyme. A sgRNA translated from the same vector guides the base editing enzyme to the target site and causes editing, dependent on the activity of the enzyme. An example for editing can be seen in the bubble in the top right of the figure. Here an A is modified to a G which causes the loss of a RE-site. A digest with the respective RE results in two possible outcomes, the plasmid is digested twice or once. Only single digested plasmids are valid templates for the error-prone PCR (performed with the primers that are indicated as arrows) to start a new cycle of evolution.
Fig. 5A is a schematic illustration of the base editing assay explained in detail in Example 8. Fig. 5B shows agarose gel analysis of the base editing assays performed for Casl2f-ABE (WT, SEQ ID NO: 27) and for the evolved library derived from the WT on target sites 1 to 3 (SEQ ID NOs: 15, 16 and 6). Fig. 5C shows the base editing assay results for the evolved Casl2f-ABE library after four cycles of evolution without mutation on different protein expression levels (induced with L- arabinose).
Fig. 6 shows an overview of the workflow of a further particularly preferred method according to the present invention. Evolved variants of Casl2f-ABE are cloned together with a unique molecular identifier (UMI) into a vector containing three target sites. After transformation into bacteria, transformed bacteria are cultured to express the enzymes. Plasmid DNA is isolated, and the region of interest (comprising regions 1 and 2) is excised and sequenced. Clustering of the UMIs allows consensus sequence polishing and determining the base editing events.
Fig. 7A shows the results of a screen of evolved variants of Casl2f-ABE on three target sites, namely target site 1 (SEQ ID NO: 15), target site 2 (SEQ ID NO: 6), target site 3 (SEQ ID NO: 16). UMI-clusters containing evolved variants of Casl2f-ABE are indicated as grey spots, and WT (SEQ ID NO: 27) control clusters are indicated as black spots. Three selected clusters are highlighted (2 (SEQ ID NO: 28), 3030 (SEQ ID NO: 29), 3301 (SEQ ID NO: 30)) Fig. 7B shows median editing rates of the WT clusters. Fig. 7C shows that 58 non-WT clusters identified with the method of the present invention had editing rates of more than 90% on all three target sites. Fig. 7D shows the percentages of the reads of the Casl2f-ABE screen with correct editing, no editing or other editing outcomes of the WT clusters and the three variants from clusters 2, 3030 and 3301 on the three target sites used in the screen (Figs. 6 and 7A, SEQ ID NOs: 15, 16 and 6).
Fig. 8A shows respective agarose gel analysis of the base editing plasmid assays performed for four different DNA modifying enzymes (Clusters 2, 3030, 3301 and WT control) on three different target sites simultaneously (target site 1 (SEQ ID NO: 15), target site 2 (SEQ ID NO: 6), target site 3 (SEQ ID NO: 16)) in triplicate replicates (rep. 1-3). Fig. 8B shows the quantified results of the base editing plasmid assay (Fig. 8A). The band intensities of the edited and non-edited products of the selected Casl2f-ABE clusters (2, 3030, 3301) and the WT control were determined using the image analysis software Fiji. Band intensity values of the edited products were then divided by the combined values of the edited and non-edited bands to calculate a fraction, which was converted to a percentage value. Each dot represents one replicate of the assay of the respective variant.
Fig. 9A shows editing results of the DNA modifying enzymes from clusters 2, 3030, 3301 and the WT Casl2f-ABE on three E. coli genome target sites (SEQ ID NOs: 8-10) analyzed by Sanger sequencing. Included is also a negative control where no DNA modifying enzymes were expressed in the bacterium. The arrow indicates where base editing is supposed to occur. The sequence of the nonedited genomic DNA is displayed above the curves. Bases were also placed directly above each line graph at the position where editing is supposed to happen. In case of two bases at this position, the upper base corresponds to the higher curve. Fig. 9B shows quantified base editing rates of the Sanger sequencing results from Fig. 9A using the EditR program (Kluesner et al., 2018).
Fig. 10 shows an overview of the workflow of a further particularly preferred method according to the present invention. Sequences encoding different variants of a DNA modifying enzyme are cloned into plasmids comprising respective enzyme target sites. E. coli cells are transformed with the vectors. Transformed bacteria are cultured to express the encoded DNA modifying enzymes. Plasmid DNA is isolated. From all plasmids isolated from the cell cultures, the region of interest (comprising regions 1 and 2) is excised and sequenced by nanopore sequencing.
Fig. 11 shows the results of a screen of naturally occurring tyrosine site-specific recombinases on multiple target sites. Recombinases and target sites have been previously identified, the screen was performed without the use of UMI. Recombination events are displayed by a heatmap of recombination percentage for each possible combination. Target sites are displayed horizontally and ordered based on their similarity. Recombinases tested are aligned based on their homology on the vertical axis.
Sequences
The sequences referred to in the present description are disclosed in the accompanying sequence listing.
Detailed description of the invention
Before the present invention is described in detail below, it is to be understood that this invention is not limited to the particular methodology, protocols and reagents described herein as these may vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and it is not intended to limit the scope of the present invention which will be limited only by the appended claims. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art.
Preferably, the terms used herein are defined as described in "A multilingual glossary of biotechnological terms: (IUPAC Recommendations)", Leuenberger, H.G.W, Nagel, B. and Klbl, H. eds. (1995), Helvetica Chimica Acta, CH-4010 Basel, Switzerland).
Throughout this specification and the claims which follow, unless the context requires otherwise, the word "comprise", and variations such as "comprises" and "comprising", will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps. In the following passages, different aspects of the invention are defined in more detail. Each aspect so defined may be combined with any other aspect or aspects unless clearly indicated to the contrary. Any feature indicated as being optional, preferred or advantageous may be combined with any other feature or features indicated as being optional, preferred or advantageous.
Several documents are cited throughout the text of this specification. Each of the documents cited herein (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions etc.), whether supra or infra, is hereby incorporated by reference in its entirety. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention. Some of the documents cited herein are characterized as being "incorporated by reference". In the event of a conflict between the definitions or teachings of such incorporated references and definitions or teachings recited in the present specification, the text of the present specification takes precedence.
In the following, the elements of the present invention will be described. These elements are listed with specific embodiments; however, it should be understood that they may be combined in any manner and in any number to create additional embodiments. The variously described examples and preferred embodiments should not be construed to limit the present invention to only the explicitly described embodiments. This description should be understood to support and encompass embodiments which combine the explicitly described embodiments with any number of the disclosed and/or preferred elements. Furthermore, any permutations and combinations of all described elements in this application should be considered disclosed by the description of the present application unless the context indicates otherwise.
Definitions
In the following, some definitions of terms frequently used in this specification are provided. These terms will, in each instance of its use, in the remainder of the specification have the respectively defined meaning and preferred meanings.
As used in this specification and the appended claims, the singular forms "a", "an", and "the" include plural referents, unless the content clearly dictates otherwise.
The term "sequence comparison" is used herein to refer to the process wherein one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, if necessary, subsequence coordinates are designated, and sequence algorithm program parameters are designated. Default program parameters are commonly used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters. In case where two sequences are compared and the reference sequence is not specified in comparison to which the sequence identity percentage is to be calculated, the sequence identity is to be calculated with reference to the longer of the two sequences to be compared, if not specifically indicated otherwise. If the reference sequence is indicated, the sequence identity is determined on the basis of the full length of the reference sequence indicated by one of the SEQ ID NOs of the present invention, if not specifically indicated otherwise.
Methods of alignment of sequences for comparison are well known in the art. Optimal alignment of sequences for comparison can be conducted, for example, by the local homology algorithm of Smith and Waterman (Adv. Appl. Math. 2:482, 1970), by the homology alignment algorithm of Needleman and Wunsch (J. Mol. Biol. 48:443, 1970), by the search for similarity method of Pearson and Lipman (Proc. Natl. Acad. Sci. USA 85:2444, 1988), by computerized implementations of these algorithms (e.g., GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by manual alignment and visual inspection (see, e.g., Ausubel et al., Current Protocols in Molecular Biology (1995 supplement)). Algorithms suitable for determining percent sequence identity and sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al. (Nuc. Acids Res. 25:3389-402, 1977), and Altschul et al. (J. Mol. Biol. 215:403-10, 1990), respectively. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http://www.ncbi.nlm.nih.gov/). This algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are extended in both directions along each sequence for as far as the cumulative alignment score can be increased. Cumulative scores are
calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). For amino acid sequences, a scoring matrix is used to calculate the cumulative score. Extension of the word hits in each direction are halted when: the cumulative alignment score falls off by the quantity X from its maximum achieved value; the cumulative score goes to zero or below, due to the accumulation of one or more negative-scoring residue alignments; or the end of either sequence is reached. The BLAST algorithm parameters W, T, and X determine the sensitivity and speed of the alignment. The BLASTN program (for nucleotide sequences) uses as defaults a wordlength (W) of 11, an expectation (E) of 10, M=5, N=-4 and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a wordlength of 3, and expectation (E) of 10, and the BLOSUM62 scoring matrix (see Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915, 1989) alignments (B) of 50, expectation (E) of 10, M=5, N=-4, and a comparison of both strands. The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin and Altschul, Proc. Natl. Acad. Sci. USA 90:5873-87, 1993). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.2, typically less than about 0.01, and more typically less than about 0.001.
The term "nucleic acid" and "nucleic acid molecule" are used synonymously herein and are understood as well -accepted in the art, i.e. as single or double-stranded oligo- or polymers of deoxyribonucleotide or ribonucleotide bases or both. The term "nucleic acids" as used herein includes not only deoxyribonucleic acids (DNA) and ribonucleic acids (RNA), but also all other linear polymers in which the bases adenine (A), cytosine (C), guanine (G) and thymine (T) or uracil (U) are arranged in a corresponding sequence (nucleic acid sequence). The invention also comprises the corresponding RNA sequences (in which thymine is replaced by uracil), complementary sequences and sequences with modified nucleic acid backbone or 3 'or 5 '-terminus. Nucleic acids in the form of DNA are however preferred.
The term "DNA modifying enzyme" as used herein describes any enzyme that is capable of manipulating the structure of a nucleic acid in a genome. More specifically, the term refers to respective enzymes that can catalyse a recombination selected among an excision, integration, inversion and translocation reaction. The term encompasses proteins selected from the group consisting of recombinase, integrase, adenosine base editor (ABE), zine-finger nuclease, transcription activator-like effector nuclease, and Cas nuclease. Such enzymes are often present in the form of dimers or tetramers, meaning that the enzyme comprises two or more subunits that may be the same or different. A specifically preferred example of a DNA modifying enzyme are recombinases which specifically include site-specific recombinases (SSRs) and more specifically tyrosine recombinases
(Y -SSRs). Recombinases can be present in the form of a monomer, a dimer or a tetramer. Dimers and tetramers can comprise two or four monomers of the same protein having recombinase activity (homodimer or homotetramer), respectively, or alternatively two or more different monomers having recombinase activity (heterodimer or heterotetramer).
The term "target site" (also referred to as "recognition site") as used herein refers to a specific nucleotide sequence which a DNA modifying enzyme recognizes, and at which or in the vicinity of which a DNA modification such as breakage and strand exchanges occur. One example of a target site is the recombinase target site. Target sites or recognition sites for recombinases typically range between 30 and 200 base pairs in length and are comprised of two inversely repeated recombinase binding regions flanking a central spacer sequence (Meinke et al., 2016). An example of such a recognition site can be seen in the SSR Cre/loxP binding complex, where the Cre recombinase is bound to the 34 base pair loxP target sequence. The loxP recognition site comprises two 13 base pair inverted repeat Cre binding elements flanking an 8 base pair spacer region. The left half-site is the 13 base pair binding element to the left of the spacer and the right half-site is the 13 base pair binding element to the right of the spacer. Depending on the number and relative orientation of the recognition sites and their spacers, the DNA recombining enzyme either performs an excision, an integration, an inversion or a replacement of genetic content (reviewed in Meinke et al., 2016). Therefore, a target site or recognition site according to the present invention is preferably a nucleotide sequence comprising a first half-site, a second half-site, and a spacer separating the first and the second half-site. A further example of a target site is a target site for Cas like gene editors. The target site for the UnlCasl2fl enzyme for example has a 4 bp PAM site (TTTR) followed by a 20 bp recognition site for the sgRNA. Further target sites of DNA modifying enzymes are e.g. UnlCasl2fl target sites (disclosed e.g. in Xin et al., 2022), Cas9 target sites (disclosed e.g. in Cong et al., 2013), TALEN target sites (disclosed e.g. in Gaj et al., 2013), and homing endonuclease target sites (disclosed e.g. in Stoddard, 2006). All scientific literature cited herein is incorporated by reference.
For a recombination event to occur, the recombinase complex recognizes a first recognition site and a second recognition site on a DNA double strand. The recognition sites are also referred to as upstream and downstream recognition sites, depending on their location on the DNA double strand.
In symmetric target sites, the first half-site (e.g. the left half-site) and the second half-site (e.g. the right half-site) are identical and palindromic (reverse complement). In asymmetric target sites, the first half-site (e.g. the left half-site) and the second half-site (e.g. the right half-site) are not identical and not palindromic, i.e. they differ from each other in at least one nucleotide.
The term "variant(s) of a DNA modifying enzyme" as used herein denotes a DNA modifying enzyme carrying modifications in its amino acid sequence compared to a reference DNA modifying enzyme. Preferably, the reference DNA modifying enzyme is a naturally occurring DNA modifying enzyme or a known DNA modifying enzyme, such as a DNA modifying enzyme described in the art. A variant of a DNA modifying enzyme preferably exhibits at least one amino acid modification
compared to the reference DNA modifying enzyme. The term "at least one amino acid modification" as used in the context of variant(s) of a DNA modifying enzyme is not limited to a specific number of amino acid modifications but includes preferably one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more amino acid modifications compared to the reference DNA modifying enzyme. Amino acid modifications in this context include substitutions, insertions and excisions. Examples of respective expression libraries of DNA modifying enzymes are disclosed in Buchholz and Stewart, 2001, Lansing et al., 2019, Hoersten et al., 2022, and in Lansing et al., 2022. Likewise, the term "variant(s) of one or more target sites" as used herein denotes target sites carrying at least one modification in their nucleic acid sequence compared to a reference target site. The term "at least one nucleic acid modification" or "at least one nucleotide modification" (both terms are used interchangeably herein) as used in the context of variant(s) of a target site is not limited to a specific number of nucleotide modifications but includes at least one or more and preferably one, two, three, four, five, six, seven, eight, nine, ten, eleven, or twelve nucleotide modifications compared to the reference target site. Nucleotide modifications in this context include substitutions, insertions and excisions. Preferably, the reference target site is a naturally occurring target site or a known target site, such as a target site described in the art for a DNA modifying enzyme.
The term "unique molecular identifier" ("UMI") as used herein denotes a type of molecular barcoding. Molecular barcodes are short sequences to uniquely tag a molecule in a sample library. UMIs are known to the skilled person and are described in detail in Zurek et al., 2020, or Karst et al., 2021, both of which are herein incorporated by reference in their entirety. According to a preferred embodiment of the present invention, a UMI is an oligonucleotide comprising at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90 or at least 100 random nucleotides. According to a particularly preferred embodiment, a UMI is an oligonucleotide comprising at least 50 random nucleotides. A UMI may preferably form part of a UMI-tag, which may further comprise one or more sequences downstream and/or upstream of the random nucleotides of the UMI (e.g. flanking the random nucleotides), which sequences may serve as one or more primer binding sites. A UMI-tag may further comprise one or more and preferably at least two restriction sites preferably down-stream and/or upstream of the random nucleotides of the UMI. According to a preferred embodiment, a UMI-tag comprises a first primer binding site, a first restriction site, the random nucleotides of the UMI, a second restriction site, and a second primer binding site.
The term "cell" or "host cell" as used herein relates to an intact cell, i.e. a cell with an intact membrane that has not released its normal intracellular components such as enzymes, organelles, or genetic material. An intact cell preferably is a viable cell, i.e. a living cell capable of carrying out its normal metabolic functions. Preferably said term relates to any cell which can be transfected or transformed with an exogenous nucleic acid.
The term "regulatory nucleic acid sequence" as used herein refers to gene regulatory regions of DNA. In addition to promoter regions, this term encompasses operator regions more distant from the gene as well as nucleic acid sequences that influence the expression of a gene, such as ciselements, enhancers or silencers. The term "promoter region" as used herein refers to a nucleotide sequence on the DNA allowing a regulated expression of a gene. The promoter region allows regulated expression of the nucleic acid encoding for the respective protein. The promoter region is located at the 5'-end of the gene and thus before the coding region. Both, bacterial and eukaryotic promoters are applicable for the present invention.
The term "identical" is used herein in the context of two or more nucleic acids or polypeptide sequences, to refer to two or more sequences or subsequences that are the same, i.e. that comprise the same sequence of nucleotides or amino acids. Sequences are "identical" to each other if they have the same sequence of nucleotides or amino acid residues (= 100% identity). The term "essentially identical" used in the context of two nucleic acids or polypeptides means that the two sequences compared share at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identity over the specified sequence, when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection. The percentage identity for "essentially identical" likewise applies for sequences that are "essentially reverse complementary" to each other.
The term "one or more" as used herein refers to the number of the respective entity denoted with this term. It specifically encompasses one, two, three, four, five, six, seven, eight, nine, ten or more than ten, preferably two, three or four, more preferably two or three. For example, the term "one or more target sites" preferably refers to one target site, two target sites, three target sites, four target sites or more than four target sites such as five, six or seven target sites. According to a preferred embodiment, the term "one or more target sites" refers to two target sites or three target sites. According to a particularly preferred embodiment, the term refers to two target sites.
The terms "library of expression vectors" and "expression library" are used interchangeably herein. An expression library is commonly known as a collection of plasmids (or phages) containing a representative sample of DNA or genomic fragments that are constructed in such a way that they will be transcribed and translated by a host organism into which the expression library is introduced. In the context of the present invention, the expression library is used for expression cloning of variants of the DNA modifying enzyme disclosed herein or for expression cloning of the DNA modifying enzyme and variants the one or more target sites of the DNA modifying enzyme as disclosed herein. Expression cloning is a technique in DNA cloning that uses expression vectors to generate a library of clones, with each clone expressing one protein, in the present case one (variant of the) DNA modifying enzyme. The expression library is screened for the property of interest, and clones of
interest are recovered for further analysis. Expression cloning is a method well-known in the art and described in further detail e.g. in Lodes et al., 2004, incorporated herein by reference.
Description of embodiments
According to the present invention, a method for screening and sequencing of a plurality of DNA modifying enzymes comprises the steps of: providing a library of expression vectors comprising a plurality of different vectors, wherein each vector comprises a first region encoding a variant of a DNA modifying enzyme, and a second region comprising one or more target sites of a DNA modifying enzyme; introducing the library of expression vectors into host cells; culturing the host cells and expressing the DNA modifying enzymes; isolating plasmid DNA from the culture of host cells; sequencing at least the first and the second region of the expression vectors; and determining, based on the sequencing results, whether the DNA sequence of the second region on the expression vector has been altered by the DNA modifying enzyme.
In accordance with said method and as defined above, the DNA modifying enzyme is preferably selected from the group consisting of a recombinase, an integrase, an adenosine base editor (ABE), a zinc-finger nuclease, a transcription activator-like effector nuclease (TALEN) and a Cas nuclease. The DNA modifying enzyme can also be a meganuclease (also termed homing endonuclease). According to a preferred embodiment, the DNA modifying enzyme is a recombinase or an integrase. According to a particularly preferred embodiment, the DNA modifying enzyme is a recombinase.
According to one embodiment of the present invention, the DNA modifying enzyme is a monomer. According to a further preferred embodiment of the present invention, the DNA modifying enzyme is a dimer that comprises at least two protein monomers, also termed subunits. The terms monomer and subunit are used interchangeably herein. A dimer can comprise two monomers of the same type (i.e. two identical protein monomers, homodimer) or two monomers of a different type (i.e. two different protein monomers, heterodimer). According to a further preferred embodiment of the present invention, the DNA modifying enzyme comprises at least four protein monomers, i.e. it is in a tetrameric form (tetramer). Such a tetramer can comprise four monomers of the same type (i.e. four identical protein monomers, homotetramer), or monomers of a different type such as two, three or four different monomers (heterotetramer). According to a specifically preferred embodiment, the DNA modifying enzyme comprises a heterodimer or heterotetramer comprising different protein monomers, i.e. is a heterodimer or a heterotetramer.
According to the present invention, each expression vector may encode one variant of a DNA modifying enzyme or a plurality of variants of a DNA modifying enzyme. In line therewith, each expression vector may encode a first monomer of a DNA modifying enzyme and a second monomer
of a DNA modifying enzyme, which monomers may be the same or different. Likewise, each expression vector may encode a first, a second, a third and a fourth monomer of a DNA modifying enzyme, which monomers may be the same or different. In cases where a vector encodes a plurality of variants of a DNA modifying enzyme, the vector preferably encodes two or four variants (or different monomers) of a DNA modifying enzyme. In such cases, the encoded variants of the DNA modifying enzymes preferably work together to modify the DNA. The two or four variants of the DNA modifying enzyme preferably work together in form of a dimer or a tetramer as defined herein.
The DNA modifying enzyme can be a naturally occurring DNA modifying enzyme or it can be a variant of a naturally occurring DNA modifying enzyme. According to a preferred embodiment of the present invention, the DNA modifying enzyme is an evolved enzyme. Preferably, the DNA modifying enzyme has been evolved applying substrate linked directed evolution (SLiDE) as described in e.g. Buchholz and Stewart 2001, Buchholz and Hauber 2011, Lansing et al. 2019, Hoersten et al. 2021, and Lansing et al. 2022, all of which are incorporated herein by reference in their entirety.
According to a preferred embodiment, the variants of the DNA modifying enzyme are based on a naturally occurring DNA modifying enzyme or on a known DNA modifying enzyme, and comprise one, two, three, four, five, six, seven, eight, nine, ten, eleven, twelve, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more amino acid modifications compared to the naturally occurring DNA modifying enzyme or to the known DNA modifying enzyme, which are also referred to as reference DNA modifying enzyme herein. When evolving DNA modifying enzymes, the variants of the DNA modifying enzyme in a first expression library preferably comprise a small number of amino acid modifications compared to the reference DNA modifying enzyme, such as one, two, three, four, five, six, seven, eight, nine, ten amino acid modifications. During the evolution of the DNA modifying enzyme, the number of amino acid modifications compared to the reference DNA modifying enzyme preferably increases in each further expression library. Examples of respective evolution processes are described in Buchholz and Stewart 2001, Buchholz and Hauber 2011, Lansing et al. 2019, Hoersten et al. 2021, and Lansing et al. 2022, all of which are incorporated herein by reference in their entirety.
According to a specifically preferred embodiment of the present invention, the protein having recombinase activity is a recombinase, more preferably a site-specific recombinase, even more preferably a tyrosine site-specific recombinase.
According to an alternatively preferred embodiment, the DNA modifying enzyme is an integrase.
In accordance with the present invention and as defined above, a target site of a DNA modifying enzyme is a nucleotide sequence. Particularly in cases of a recombinase as DNA modifying enzyme, a target site comprises a first half-site, a second half-site, and a spacer separating the first and the second half-sites as defined herein above. According to a preferred embodiment, the target site can
be the target site known for the respective DNA modifying enzyme. However, the target site is not limited to the specific target site that might be known in the art for the DNA modifying enzyme. The target site can be a modified version of a known target site. As such, the target site can be a naturally occurring target site of a DNA modifying enzyme, or it can be an artificially created target site such as a modified version of a known target site, preferably of a known naturally occurring target site. Such modified versions of a known target site preferably differ in one, two, three, four, five, six, seven, eight, nine, ten, eleven or twelve nucleotides from the known target site. A modified target site may alternatively differ from a known target site in about 2%, about 4%, about 6%, about 8%, about 10%, about 12%, about 14%, about 16%, about 18%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, or in about 50% of the nucleotides.
According to a preferred embodiment, the variants of the target site are based on a naturally occurring target site or on a known target site of a DNA modifying enzyme, and comprise one, two, three, four, five, six, seven, eight, nine, ten, eleven, or twelve nucleotide modifications compared to the naturally occurring target site or to the known target site of a DNA modifying enzyme, which are also referred to as reference target sites herein.
The expression vector in accordance with the present invention is not limited to any specific expression vector. Typically, an expression vector comprises an origin of replication, a promoter, as well as specific gene sequences that allow phenotypic selection of host cells comprising the vector. In accordance with the present invention, the vector comprises a first region encoding a variant of a DNA modifying enzyme, and a second region comprising at least one, at least two, at least three or more target sites of a DNA modifying enzyme and as defined herein above. If more than one target site is present in the vector, the target sites are preferably essentially identical or essentially reverse complementary to each other.
According to a particularly preferred embodiment, the vector comprises two target sites of a DNA modifying enzyme. In such a case, the DNA modifying enzyme is preferably a recombinase or an integrase. According to a particularly preferred embodiment, in cases of two target sites being present in the vector, the DNA recombining enzyme is a recombinase.
A particularly preferred expression vector to be used in accordance with the present invention is the pEVO vector as described in Buchholz and Stewart 2001.
In cases of using evolved DNA modifying enzymes in the method of the present invention, the gene encoding the DNA modifying enzyme is preferably excised from a respective evolution library as shown in Figs. 1 and 6. Respective evolution libraries are described e.g. in Buchholz and Stewart 2001, and in Lansing et al. 2022.
According to a particularly preferred embodiment of the present invention, the expression vector further comprises a unique molecular identifier (UMI) as defined herein above. The position of the UMI in the vector is not particularly limited. However, according to a preferred embodiment, the unique molecular identifier is located in the first region of the expression vector adjacent to the
sequence encoding the DNA modifying enzyme. According to a particularly preferred embodiment, the UMI is located downstream of the sequence encoding the DNA modifying enzyme, and thus upstream of the second region comprising one or more target sites of a DNA modifying enzyme as shown for example in Figs. 1 and 6. According to a particularly preferred embodiment, the UMI forms part of a UMI-tag further comprising in addition to the random nucleotide sequence of the UMI further sequences such as one or more primer binding sites and/or further sequences such as one or more restriction sites. According to one embodiment, the primer binding sites and the restrictions sites are preferably located upstream and downstream of the random nucleotides, such as that a UMI-tag according to a preferred embodiment of the invention has the structure: primer binding site 1 - restriction site 1 - random nucleotides - restriction site 2 - primer binding site 2. An exemplary and preferred UMI-tag in accordance with the present invention has a sequence according to SEQ ID NO: 17 or SEQ ID NO: 18. These exemplary sequences contain 50 random nucleotides at positions 43 to 92 termed the UMI. In said sequences, the random nucleotides are represented by letters "N", "Y", and "R", with "N" representing any nucleotide selected from the group consisting of A, C, G and T; with "Y" representing any nucleotide selected from the group consisting of C and T, and "R" representing any nucleotide selected from the group consisting of A and G. Primer binding sites are located at positions 1 to 20 and 101 to 120. In SEQ ID NO: 17, restriction site 1 (BsiWI) is located at positions 21 to 26 and restriction site 2 (Sbfl) is located at positions 93 to 101. In SEQ ID NO: 18, restriction site 1 (Xbal) is located at positions 21 to 26 and restriction site 2 (Sbfl) is located at positions 93 to 100.
The library of expression vectors may comprise a plurality of different libraries. Each library in the plurality of libraries preferably differs from the other libraries in the number of modifications introduced to the variants of the DNA modifying enzyme encoded by the respective expression vectors or to the variant of the target sites. For example, a first library of the plurality of different libraries may comprise a plurality of different vectors, wherein each vector encodes a variant of a DNA modifying enzyme having one, two or three modifications compared to the reference DNA modifying enzyme, such as a naturally occurring DNA modifying enzyme or a known DNA modifying enzyme. A second library of the plurality of different libraries may comprise a plurality of different vectors, wherein each vector encodes a variant of a DNA modifying enzyme having one, two or three modifications compared to the variants of the DNA modifying enzymes encoded in the first library, or in other words four, five or six modifications compared to the reference DNA modifying enzyme. A third library may then encode variants of a DNA modifying enzyme having one, two or three modifications compared to the variants of DNA modifying enzymes encoded in the second library, or in other words seven, eight or nine modifications compared to the reference DNA modifying enzyme, and so on. It is to be understood that the number of modifications cited in this example is not particularly limited and is for exemplary reasons only. The same applies likewise to variants of target sites according to the method of the present invention, i.e. a first library includes target sites having a limited number of nucleotide
modifications, compared to the reference target site a second library includes target sites having the same or similar number of nucleotide modifications compared to the target sites of the first library and so on.
According to the present invention, the library of expression vectors is introduced into the host cells. A host cell within the meaning of the invention is a naturally occurring cell or a cell line (optionally transformed or genetically modified) that comprises at least one vector as described above. The invention thus includes host cells that include at least one expression vector according to the invention as a plasmid. According to a preferred embodiment, a host cell includes one expression vector according to the invention as a plasmid. This embodiment is particularly useful if bacteria are used as host cells.
According to an embodiment of the present invention, the host cell is selected from the group consisting of bacterial cells, yeast cells, fungal cells, and mammalian cells. Suitable bacterial cells include cells from gram-negative bacterial strains such as strains of Escherichia coli, Proteus, and Pseudomonas, and gram-positive bacterial strains such as strains of Bacillus, Streptomyces, Staphylococcus, and Lactococcus. Suitable fungal cells include cells from species of Trichoderma, Neurospora, and Aspergillus . Suitable yeast cells include cells from species of Saccharomyces (for example Saccharomyces cerevisiae), Schizosaccharomyces (for example Schizo saccharomyces pomhe), Pichia (for example Pichia pastoris and Pichia methanolicd), and Hansenula. Suitable mammalian cells include for example CHO cells, BHK cells, HeLa cells, COS cells, 293 HEK and the like. However, amphibian cells, insect cells, plant cells, and any other cells used in the art for the expression can be used as well. According to a preferred embodiment, the host cell is a bacterial cell. According to a particularly preferred embodiment, the host cell is an E. coli cell.
The introduction of the nucleic acids into the host cells is performed using techniques of genetic manipulation known by a person skilled in the art. Among suitable methods are cell transformation, transfection or viral infection, whereby a nucleic acid sequence encoding the DNA modifying enzyme is introduced into the cell as a component of the vector. The cell culturing is carried out by methods known to a person skilled in the art for the culture of the respective cells. Therefore, cells are preferably transferred into a conventional culture medium, and cultured at temperatures and in a gas atmosphere that is conducive to the survival of the cells and allows expression of the encoded protein. A preferred method for introducing the library of expression vectors into host cells is transformation, preferably by electroporation. As an example of the step of introducing the library of expression vectors into host cells, E. coli cells can be transformed with the vectors using electroporation.
According to one embodiment of the method of the present invention, the host cells comprising the expression vectors are plated on a suitable medium such as an agar plate to allow growing and selection of individual cell colonies. This step may also be used for selecting a specific number of individual colonies instead of culturing all cells simultaneously, thereby decreasing the
overall number of the members of the library and reducing the number of enzyme variants to be screened. Alternatively, the host cells comprising the expression vectors can be briefly cultured in a respective culture medium (e.g. for a limited time, for example 0.5 to 1 hour) and an amount of the culture medium equal to the desired number of variants can be taken from this culture medium for further culturing as described in the following. The number of transformed bacteria present per pl of culture medium can be estimated based on the number of colonies on the plates. If for example 10 pl of the medium were spread on a plate and the plate after incubation shows 100 distinct colonies, one pl of medium contains ten transformed bacteria.
According to the present invention, the host cells comprising the library of expression vectors or the cells selected as explained above are cultured to allow the expression of the plasmid. Culturing conditions depend on the cell line used and are not particularly limited. It is, however, not necessary to culture each selected colony individually. The selected colonies can be cultured together in a single cell culture. Alternatively, the host cells comprising the library of expression vectors or the cells selected as explained above are cultured in more than one separate cell culture. For activation of the expression of the nucleic acid encoding for the DNA modifying enzyme, the nucleic acid encoding for the DNA modifying enzyme further comprises a regulatory nucleic acid sequence, preferably a promoter region. Hence, expression of the nucleic acid encoding for DNA modifying enzyme is initiated or regulated by activating the regulatory nucleic acid sequence. The cells are preferably cultured under conditions allowing cell growth and protein expression. Preferably, the cells are grown for more than two hours up to 12 hours or longer, allowing more than two, preferably more than three, more than four, more than five, more than six, more than seven, more than eight, more than nine or more than ten cell doublings.
Upon expression of the DNA modifying enzymes in the culture of host cells, the DNA modifying enzymes start modifying the DNA at the one or more target sites in the second region of the expression vector, if said DNA modifying enzymes recognizes the respective target site(s). For example, in cases of recombinases, for a recombination event to occur, the recombinase complex recognizes a first target site and a second target site on the plasmid. In cases of excision events of a DNA segment, said segment is flanked by two target sites oriented in the same direction and mediated by the respective recombinase protein encoded on the same vector. If such a recombination event occurs, the DNA segment between the two target sites is excised during the culturing of the cells. Alternative modifications of the DNA in the second region of the vector are also encompassed by the present invention, such as but not limited to an inversion of a segment of the DNA in the second region of the vector, an integration event, or an event on the level of individual nucleotides in the second region of the vector. To determine if any modification in the second region of the vector has occurred, plasmid DNA is isolated from the culture(s) of host cells after culturing of the host cells and expressing the encoded library of DNA modifying enzymes. Isolating of plasmid DNA from a culture of cells can be done by any conventional methods well known in the art. After isolation, the plasmid
DNA is sequenced to determine its nucleotide sequence at least in the first and in the second region. It will be appreciated that more parts or all of the plasmid can be sequenced than the first and the second region containing the sequences encoding the one or more DNA modifying enzyme and the one or more target sites. However, for a high-throughput screening, it is preferred that the first and the second region are sequenced since the information obtained by sequencing the first and the second region is sufficient for the method of the present invention. This sequencing step allows determining the sequence of the specific DNA modifying enzyme, the sequence of the target region and whether not the DNA sequence in the second region on the expression vector has been altered by the DNA modifying enzyme. According to a preferred embodiment, the first and the second region of the vector are sequenced in a single method step, i.e. they are sequenced together and not subsequently.
The sequencing method is not particularly limited. However, for particularly high-throughput methods, long-read sequencing methods are preferred. A particularly preferred sequencing method in accordance with the present invention is nanopore sequencing as described in Wang et al., 2021 (incorporated herein by reference), or as provided by e.g. Oxford Nanopore Technologies, UK. In brief, nanopore sequencing works by monitoring changes to an ionic current as nucleic acids are passed through a protein nanopore. The resulting signal is decoded to provide the specific DNA sequence. Based on the sequencing results, an activity rate for each DNA modifying enzyme on the respective target site(s) can be determined.
In cases in which the sequences of the DNA modifying enzymes are known, sequence comparisons of the screen sequence data to the known sequences of the DNA modifying enzymes are performed to identify which sequence reads belong to which DNA modifying enzyme. DNA modifying enzyme sequences can be known because they have been sequenced previously, or because they were synthesized according to a defined sequence. The sequence reads are separated according to their sequence matches and from here on processed separately for consensus sequence generation and determination of the activity rate. It is preferred that in such cases no UMI and no UMI-tag are used in the method of the invention.
In cases in which a UMI or a UMI-tag is used, the method according to the present invention may further comprise the steps of clustering the unique molecular identifiers, generating and polishing consensus sequences, determining the number of DNA modification events for each DNA modifying enzyme, and determining an activity rate for each DNA modifying enzyme. Clustering and polishing of respective sequencing results are described in Zurek at el., 2020, and Karst et al., 2021, both of which are incorporated herein by reference in their entirety. For example, the UMI sequence of the sequencing reads can be identified by comparing the sequences upstream and/or downstream of the UMI. In case of using the upstream sequence for the identification of the UMI, the DNA sequence with the length of the UMI after the match is collected with the sequence read ID. In case of using the downstream sequence for the identification of the UMI, the DNA sequence with the length of the UMI before the match is collected with the sequence read ID. In case of using the downstream and upstream
DNA for the identification of the UMI, the DNA sequence between the matches is collected with the sequence read ID. The collected UMI sequences are then clustered based on sequence identity, while maintaining their association to their read IDs, for example with vsearch (Rognes et al., 2016). Based on these clusters and their associated read IDs, the sequence reads are then separated into different files, where each file contains reads that corresponds to one UMI-cluster. From each of these clusters of separated reads, the part of the sequence that contains the gene for the DNA modifying enzyme is aligned to a reference gene that contains a high similarity to the screened variants. Typically, such a gene is determined through sequencing of the screened DNA modifying enzyme library or by selecting the DNA modifying enzyme gene that was modified to produce the screened variant library. With the aid of this alignment, a software for consensus sequence generation and sequence polishing can be used to determine the most likely gene sequence of the DNA modifying enzyme associated to this cluster based on all sequence reads. Polishing in this respect refers to additional rounds of processing of the reads, leading to improved sequence accuracy.
The method for determining the activity rate is not particularly limited. For example, particularly for DNA modifying enzymes that cause DNA changes of five or more bases, the sequencing reads, which have been separated through sequence comparison or UMI clustering, are compared to expected sequences from the second region in the changed and non-changed sequences. Based on this sequence comparison, the number of reads matched to the changed DNA can be counted. The counted changes divided by the total number of reads matching the second region gives rise to an activity rate of the DNA modifying enzyme which is associated to the separated reads. A particularly preferred example for DNA changes of five or more bases is the excision of 741 bp in the pEVO expression vector using recombinases (Example 2 and 5).
In cases in which a DNA modifying enzyme causes changes that are less than 5 bp in size, the location of the expected changes in the second region of the sequencing reads can be identified by sequence comparison to the second region. The sequence at the location of the changes in the sequencing reads is collected to count all observed outcomes. The counted outcomes with the desired DNA changes divided by the total number of reads matching the second region gives rise to an activity rate of the DNA modifying enzyme which is associated to the separated reads.
In cases in which a DNA modifying enzyme is screened on multiple target sites, additional references for the second region can be provided to associate the activity rate for the identified DNA modifying enzyme to multiple target sites. The additional references are sequences of the second region with the target sites replaced with the alternative target sites that were used in the screen. These references can also be provided in a changed and non-changed version to determine the activity rate as described above.
According to a preferred embodiment of the present invention, the first and the second region of the expression vector are excised from the expression vector before the sequencing, and the reminder of the vector is discarded. Excision can be performed with any conventional method known
in the art. A preferred method for excising parts of a vector is the use of restriction enzymes. The excision of the first and the second region of the expression vector allows the sequencing reaction to take place only on the relevant parts of the expression vector, i.e. those parts containing the sequence encoding the DNA modifying enzyme, optionally the UMI, the target sequence(s) and the segment prone to modification by the DNA modifying enzyme.
According to a particularly preferred embodiment of the present invention, genes encoding DNA modifying enzymes are cloned together with a unique molecular identifier (UMI) into an expression vector, which further contains one or more of the target sites of interest. The barcoded plasmids are transformed into respective cells, and the cells are spread on one or more culture plates. Based on the number of colonies on the plates, a defined number of clones is cultured in medium that induces the expression of the genes encoding the DNA modifying enzymes. This results in multiple copies of the transformed plasmids, with a fraction of the plasmids being modified by the DNA modifying enzymes. The plasmids are isolated and sequenced, preferably using nanopore sequencing. Using the UMIs, the sequences are then clustered. These clusters are used to construct accurate consensus sequences of the variant genes. A cluster thus refers to an enzyme variant identified with the methods of the present invention. Counting of the recombined plasmids generates an editing rate for the particular variant of the DNA modifying enzyme on the respective target site(s).
The present invention as laid out above can be likewise used for screening and sequencing of a plurality of target sites for DNA modifying enzymes. Accordingly, all embodiments and descriptions provided herein above apply likewise to a further aspect of the present invention, which provides a method for screening and sequencing of a plurality of target sites for DNA modifying enzymes, the method comprising the steps of: providing a library of expression vectors comprising a plurality of different vectors, wherein each vector comprises a first region encoding a DNA modifying enzyme, and a second region comprising variants of one or more target sites of a DNA modifying enzyme; introducing the library of expression vectors into host cells; culturing the host cells and expressing the DNA modifying enzymes; isolating plasmid DNA from the culture of host cells; sequencing the first and the second region of the expression vectors; and determining, based on the sequencing results, whether the DNA sequence of the second region on the expression vector has been altered by the DNA modifying enzyme.
In this second aspect of the present invention, the target sites in the second region of the vector can be different naturally occurring target site of different DNA modifying enzymes. In addition or alternatively, the target sites can be an artificially created target sites such as modified versions of a known target site. Such modified versions of a known target site preferably differ in one, two, three, four, five, six, seven, eight, nine or ten nucleotides from the known target site. A modified target site may alternatively differ from a known target site in about 2%, about 4%, about 6%, about 8%, about
10%, about 12%, about 14%, about 16%, about 18%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, or in about 50% of the nucleotides.
The present invention provides a new sequencing and screening method acquiring the sequence information and the recombination rate of library variants in a high-throughput manner. Specifically, the high-throughput screening method enables characterization of thousands of DNA editing enzyme variants on multiple target sites, and likewise characterization of thousands of target sites for a DNA editing enzyme variant. The method preferably utilizes nanopore technology for sequencing full-length enzyme variants at fast turn-around times. By simultaneously capturing the target site with the enzyme sequence, the DNA editing rate for each enzyme variant can be quantified. Using unique molecular identifiers, the sequencing results can be clustered to a highly accurate consensus sequence. The applicability of the method of the present invention is demonstrated in the following non-limiting examples using two different DNA modifying enzymes, i.e. on variants of the Cre designer-recombinase D7 (Lansing et al., 2022) and on evolved Casl2f-derived mini-ABEs.
All references and patent literature cited herein are incorporated by reference in their entirety.
Examples
Example 1: UMI fragment preparation
To generate a DNA fragment for ligation, a single stranded DNA oligonucleotide was ordered containing primer binding sites, restriction sites, and 50 random bases. There were two variants of this oligonucleotide depending on which pEVO plasmid (Buchholz and Stewart, 2001) it was intended to be used for. For screening of the evolved Casl2f-ABE variants, the "Casl2f UMI-tag" (SEQ ID NO: 18) was used, while for screening the loxF8 recombinase variants the "loxF8 UMI-tag" (SEQ ID NO: 17) was used. To make these oligonucleotides double stranded, a 50 pl PCR was performed with 20 pM of the primers UMIprimer F and UMIprimer R (SEQ ID NOs: 32 and 33), 10 pM of the oligonucleotide, 10 pl of 5x MyTaq buffer and 1 pl MyTaq polymerase (Bioline). The PCR-cycler was set 94°C for 90 seconds, followed by 10 cycles of 15 seconds at 94°C, 15 seconds at 54°C and 15 seconds at 72°C. The resulting PCR product was digested with Sbfl and Xbal for the Casl2f UMI-tag or BsiWI and Sbfl for the loxF8 UMI-tag. The digest was then again cleaned up with the Isolate II PCR and Gel Kit (Bioline) and measured with a Qubit HS dsDNA Kit on a Qubit 2.0 (Thermo Fisher Scientific).
Example 2: Enzyme variant barcoding
Evolved libraries of Casl2f-ABE and loxF8 recombinases were obtained by directed evolution of the fusion protein UnlCasl2fl-Tad8e (short name: Casl2f-ABE, SEQ ID NO: 27) and the tyrosine site-specific recombinase Cre (SEQ ID NO: 31). The resulting enzyme variants contain randomly
acquired mutations in comparison to their origin. The DNA editing enzyme gene variants were acquired from the pEVO plasmids used in the evolution by digesting with Xbal and BsrGI-HF for the evolved Casl2f-ABE library or SacI and BsiWI for the loxF8 library. The Casl2f-ABE library and the Casl2f-ABE controls were ligated in a ratio of 60 ng of Casl2f-ABE fragment, 4.8 ng UMI-tag and 100 ng of Bsrgl-HF and Sbfl digested pEVO-BE plasmid. The loxF8 library and the D7 control (Lansing et al. 2022, SEQ ID NO: 19 and 20) were ligated in a ratio of 40 ng of the recombinase gene fragment, 1.9 ng of the loxF8 UMI-tag, and 120 ng of SacI and Sbfl digested pEVO-loxF8 plasmid.
Ligated plasmids were desalted with MF -Millipore membrane fdters (Merck) on distilled water for 30 minutes and transformed into XL-1 Blue E. coli (Agilent) via electroporation. The transformed bacteria were cultured in SOC medium for 30 minutes at 37°C. 2 pl of this culture was spread on agarose plates with 15 mg/ml chloramphenicol and incubated over night at 37°C. The number of colonies on the plates were counted to calculate the number of transformed bacteria present per pl of SOC culture.
To nominate the number of variants for the screen, an amount of the SOC culture equal to the desired number of variants was cultured overnight in 100 ml LB medium with 25 mg/ml chloramphenicol and a defined amount of L-arabinose. For the Casl2f-ABE screen, around 4,000 transformed bacteria per library and 100 transformed bacteria per control were cultured with 10 pg/ml L-arabinose. For the loxF8 recombinase screen, around 4,000 transformed bacteria from the library and 50 transformed bacteria of the control were cultured with 1 pg/ml L-arabinose. For each sequencing run, the different libraries and controls were cultured together and the plasmid DNA of these cultures was extracted with the GeneJet Plasmid Miniprep Kit (Thermo Fisher Scientific).
To test the same recombinase variants on multiple target sites, the recombinase DNA from the loxF8 screen was digested with SacI and Sbfl and isolated from an agarose gel with the Isolate II PCR and Gel Kit (Bioline). Recombinase dimer fragments were then cloned into pEVO plasmids containing the off-targets of interest (HG1, HG2, HG2L, SEQ ID NOs: 12 to 14; these specific off-targets were identified in the human genome as having high similarity to the on-target sites). The ligated plasmids were desalted with MF -Millipore membrane filters (Merck) and transformed into XL-1 Blue E. coli (Agilent) via electroporation. The transformed bacteria were cultured in SOC medium for 30 minutes at 37°C and then transferred to 100 ml LB medium with 25 mg/ml chloramphenicol and 100 pg/ml L- arabinose. The plasmid DNA of these cultures was subsequently extracted using the GeneJet Plasmid Miniprep Kit (Thermo Fisher Scientific).
Barcoded plasmid extracts from the Casl2f-ABE library were digested with Seal and BsrgI, while barcoded plasmids from the loxF8 screen were digested with Seal and SacI. The resulting fragments containing the evolved gene, the UMI-tag, and the target sites were isolated via agarose gel excision with the Isolate II PCR and Gel Kit (Bioline).
Example 3: Nanopore sequencing and processing of screen libraries
DNA concentrations of the digested and isolated plasmid extracts (Example 2) were measured with a Qubit dsDNA HS Assay Kit on a Qubit 2.0 Fluorometer (Thermo Fisher Scientific). Nanopore sequencing library preparation for the loxF8 library was performed according to the "Amplicons by Ligation (SQK-LSK110)" protocol from Oxford Nanopore Technologies, while the Casl2f-ABE library was prepared using the protocol "Amplicons by Ligation (SQK-LSK112)". The prepared loxF8 library was then loaded on a MinlON FLO-MIN 106 flow cell with r9.4.1 pore (Oxford Nanopore Technologies), while the Casl2f library was loaded onto a MinlON FLO-MIN 110 flow cell with rl0.4 pores (Oxford Nanopore Technologies). Sequencing was performed for 72 hours. Each screen was performed on one flow cell.
Basecalling of the sequence data was performed on guppy version 6.0.1 with the high accuracy model for the loxF8 library and the super accuracy model for the Casl2f-ABE library (Oxford Nanopore Technologies). Processing of the sequence data was performed on a custom developed pipeline (available on https://github.com/ltschmitt/DEQSeq). Reads were first filtered with Filtlong vO.2.1 (https://github.com/rrwick/Filtlong) to be at least 2,900 bp long for the Casl2f-ABE screen, and 3,000 bp long for the loxF8 screen. Further, Filtlong was also used to filter the reads for a minimum mean Phred quality value of 10 for the loxF8 screen and a minimum mean Phred quality value of 18 for the Casl2f-ABE screen. The sequences were then aligned with minimap2 (Li, 2021) to a reference sequence containing UnlCasl2fl-ABE8e (Casl2f-ABE screen, SEQ ID NO: 34) or D7 (F8 dimer screen, SEQ ID NO: 35) and the UMI consisting of 50 random bases. To ensure coverage of the genes and the UMI, the aligned reads were filtered with samtools (Danecek et al., 2021) based on coordinates at the beginning of the enzyme gene and the end of the UMI.
UMIs were then extracted from the filtered alignment using the stackStringsFromBam function from the R package GenomicAlignments (Lawrence et al., 2013). UMIs were subsequently clustered with VSEARCH (Rognes et al., 2016) with a cluster_identity value of 0.7. Sequence reads from clusters with a minimum size of 50 reads were then transferred to separate files and aligned to the gene-UMI reference sequence. These separate read files and alignments were used to construct consensus sequences with racon (Vaser et al. 2017) followed by further polishing with medaka (https://github.com/nanoporetech/medaka), both with standard settings. The polishing process was run in parallel with GNU parallel (Tange 2023). Finally, gene sequences were extracted using the R package GenomicAlignments and translated to amino acids.
For the loxF8 screen, the DNA excision rate of the enzymes was determined by aligning the clustered reads to reference sequences that contain the target site region as a non-recombined and recombined variant. These references start 100 bp upstream of the first target site and end 100 bp downstream of the last target site. For each target site sequence, separate references were provided for identifying the different targets. The additional references are essentially identical to the loxF8 recombined and non-recombined references (SEQ ID NO: 36 and 37), with the sole difference that the
loxF8 target sites were replaced with HG1, HG2 or HG2L (SEQ ID NO: 38 to 43). The recombination rate of the variants was determined based on the read counts of the target site region alignments.
For the Casl2f-ABE screen, the base editing rate of the enzyme variants was determined by aligning the clustered reads to a reference sequence that contains the unedited target site (SEQ ID NO: 44). Using the GenomicAlignments R package (Lawrence et al., 2013), the aligned stacks of four bases from position two to five counting from the TTTG PAM sequence on all three target sites were extracted. The base editing of Casl2f-ABE is expected on position three or four after the PAM sequence. Comparing this region can therefore provide information about potential editing of adjacent positions. Editing rates can then be determined by counting the correctly edited reads, non-edited reads and other editing outcome reads.
Results from each screen were combined and filtered for clusters with 100 reads or more. All further data processing and visualization was performed in R with the tidyverse (Wickham et al. 2019) and stringdist packages (van der Loo, 2014).
Example 4: Extraction of enzyme variants
To validate the screens obtained as described in Example 3, enzyme variants were extracted from the screened libraries using PCR. Reverse primers specific for the UMI of the cluster of interest were designed ("loxF8 UMI-138 R" (SEQ ID NO: 45), "loxF8 UMI-181 R" (SEQ ID NO: 46), "loxF8 UMI-1244 R" (SEQ ID NO: 47), "Casl2f UMI-2 R" (SEQ ID NO: 48), "Casl2f UMI-3030 R" (SEQ ID NO: 49), "Casl2f UMI-3301 R" (SEQ ID NO: 50)) and used together with a universal forward primer (binding on the plasmid before the sequence encoding the enzyme in region 1 ("loxF8 universal F" (SEQ ID NO: 51) for the loxF8 screen, "Evolution F" (SEQ ID NO: 52) for the Casl2f-ABE screen) to amplify enzyme genes. PCRs were performed with a high-fidelity polymerase (Herculase II Fusion DNA Polymerase, Agilent) and PCR products were digested with Xbal and BsrgI (Casl2f- ABE enzymes) or SacI and BsiWI (loxF8 recombinases) for further cloning.
Example 5: Recombination assay
A schematic illustration of the recombination assay used is shown in Fig. 3A pEVO vectors with the different target sites were published previously (Lansing et al., 2022). Recombinases were cloned into pEVO vectors with the respective target sites by utilizing SacI and BsiWI restriction enzymes. Expression of recombinases was controlled by an L-arabinose inducible promoter system (araBAD). Recombination of the respective target sites on the evolution plasmid leads to the excision of a 741 bp fragment from the plasmid. The resulting size difference is mediated by the recombinase activity and is detectable by a restriction digest followed by gel electrophoresis.
Example 6: pEVO vectors for UnlCasl2fl-ABE8e
To generate a plasmid with target sites, oligonucleotides containing the target sites (SEQ ID NOs: 54 and 55) were annealed to form a DNA fragment that was cloned into a Bglll digested pEVO backbone via Cold Fusion (System Biosciences). SgRNA scaffold fragments were synthesized (Twist Bioscience) and cloned into pEVO vector containing the target sites utilizing Nsil and Notl restriction enzymes. The pEVO plasmid containing the sgRNA scaffold and the target sites was then used as a template in a PCR reaction where three different sgRNA spacers were included in the reverse primers. In that way, generated sgRNA fragments were introduced to the pEVO vector in a stepwise manner: by cloning sgRNA 1 using Notl and Nsil, by cloning sgRNA2 using Nsil, and by cloning sgRNA3 using Xhol and Sall. UnlCasl2fl fragments were produced by Twist Bioscience and amplified via PCR, and the TadA gene was prepared via PCR by using the pABE8e-protein plasmid as template (Addgene plasmid # 161788). To create Casl2f-ABE, both PCR products were used as template for an overlap-PCR. The Casl2f-ABE was then cloned into the pEVO-TS-sgl-sg2-sg3 vector using BsrGI and Xbal restriction enzymes. The expression of the Casl2f-ABEs was controlled by the araBAD L- arabinose inducible promoter system.
Example 7: Evolution of Casl2f-ABE
A schematic illustration of the procedure is shown in Fig. 4. In a first step, the Casl2f-ABE library was generated using error-prone-PCR with a low-fidelity DNA polymerase (MyTaq, Bioline and Primers Evolution F and Evolution R (SEQ ID NOs: 52 and 53) and cloned into the vector using BsrGI and Xbal restriction enzymes. After transformation into XL-1 blue E. coli, expression of the enzymes was induced with 200 qg/ml L-arabinose. After expression of the base editors, the plasmids were isolated and digested with restriction enzymes (RE), which target the sgRNAs and their target sites. Because the base editing of the target site takes place on the restriction enzyme site, the digest of these enzymes will linearize the edited plasmid, while the non-edited plasmids will be cut into two fragments (Fig. 4). This process can be visualized and quantified using agarose gel electrophoresis (Fig. 5A). The next round of directed evolution was started with an error-prone-PCR on the enzymatically digested DNA. Only edited plasmids were replicated because primers (small arrows in Fig.4) are placed in a way that only the edited DNA fragment is a valid template for PCR. The PCR product was then cloned into non-edited pEVO vectors, which signified the start a new evolution cycle. For increasing selection pressure, L-arabinose levels and thus induction of enzyme expression were lowered from 200 qg/ml to 10 qg/ml. Finally, to enrich the library for the DNA editing quantification sequencing screen, four cycles of evolution at 10 qg/ml L-arabinose were performed with a high-fidelity polymerase (Herculase II Fusion DNA Polymerase, Agilent) (Fig. 5C).
Example 8: Base editing assay
A schematic illustration of the base editing assay is shown in Fig. 5A. Edited and non-edited pEVO plasmids (circles in top part) for base editing are digested with restriction enzymes (RE) that are specific for a sequence in the target site (quadratic box) and in the sgRNA array (grey). If the target site is edited (quadratic box with line), the RE-site is lost. Due to the different number of cuts, the number and sizes of the fragment changes. Digestion of non-edited plasmids results in two DNA fragments (smallest two fragments), digestion of edited plasmids results in one DNA fragment (biggest fragment). These fragments are visualized using agarose gel electrophoresis (schematic shown at the bottom). A mixture of edited and non-edited plasmids is shown on the right lane of the gel schematic. To the right of the gel, pictograms are used to indicate to which editing status the fragments belong to. A line with a box filled with a horizontal line indicates the edited single DNA fragment, the shorter lines with half a box indicate the two non-edited fragments that have been cut. For the assay, the same pEVO plasmid that was used in the examples before has been used. The Casl2f-ABE variants were cloned utilizing BsrGI and Xbal restriction enzymes. Expression of the variants was controlled by an L-arabinose inducible promoter system (araBAD). Base editing of the target site results in a loss of a restriction enzyme site. Therefore, the restriction digest will linearize the edited plasmid, whereas non-edited plasmids will result in two fragments, which can be detected with gel electrophoresis (Fig. 5A).
Example 9: Modification of the E. Coli genome
Three different sgRNA were cloned into the pEVO plasmid, which target genomic DNA of E. Coli. Casl2f-ABE WT and the three variants were cloned into these vectors utilizing BsrGI and Xbal restriction enzymes. Modified vectors were subsequently transformed into E. Coli cells. After an overnight culture, the cells were spun down and resuspended in 200 pl ddfEO. This suspension was then heated to 95°C for 10 minutes and spun down again. The supernatant, which contained the gDNA, was used for a PCR reaction that produces a DNA fragment containing all three sgRNA target sites (Primers "E_Coli gDNA F" (SEQ ID NO: 56) and "E_Coli gDNA R" (SEQ ID NO: 57)). The PCR products were sequenced via Sanger sequencing and the editing rates were analyzed with EditR (Kluesner et al., 2018).
Example 10: Screening of evolved site-specific recombinases
The screening method was performed on a library of evolved site-specific recombinases that target a sequence in the human FactorVIII gene (loxF8) (Lansing et al., 2022). A scheme of this method is shown in Fig. 1. The D7 variant (SEQ ID NOs: 19 and 20) that had been identified by Lansing et al., 2022 by picking and evaluating 96 random recombinases was used as a control. The aim of the screen was to identify variants that have lower off-target activity compared to D7, while maintaining similar on-target activity. To this end and in addition to the loxF8 target site, the library of
evolved recombinases was also screened simultaneously on three off-target sites that are recognized by D7, namely HG1, HG2 and HG2L (SEQ ID NO: 12, 13 and 14, respectively) (Lansing et al., 2022) (Fig. 2A). In total, the screen yielded 2,515 UMI-clusters with 50 or more reads, from which 53 clusters were identified as D7 control. Analysis of the polished D7 recombinase sequences revealed no sequence errors, indicating an accuracy of close to 100%. Median recombination rates of D7 were 80.2% on its intended target, 5.8% on HG1, 57.7% on HG2, and 72.6% on HG2L (Fig. 2B). Of the 2,476 non-D7 clusters, 70 clusters were identified that had less than 10% off-target activity on the three off-targets and more than 25% activity on the on-target (Fig. 2C).
To validate the results, three variants (clusters 138 (SEQ ID NOs: 21 and 22), 181 (SEQ ID NOs: 23 and 24), 1244 (SEQ ID NOs: 25 and 26)) were extracted from the screened loxF8 library via PCR amplification using primers specific for the UMI of the variants and a universal forward primer (SEQ ID NO: 51) binding on the plasmid before the sequence encoding the enzyme in region 1. The variants were chosen for their different levels of on-target activity, while maintaining low off-target activity. The variants were evaluated using an established plasmid-based recombination assay (Lansing et al., 2022) (Fig. 3A), which confirmed that the recombination rate of the tested variants was comparable in the assay (Figs. 3B, 3C). The method of the present invention thus identifies DNA modifying enzymes having much lower off-target activity than recombinase D7.
Example 11: Screening of evolved Casl2f-ABEs
The method of the present invention was further assessed on a CRISPR-Cas based system, a schematic representation of said method is shown in Fig. 6. Substrate Linked Directed Evolution as described in Buchholz and Stewart, 2001, was adapted to the evolution of ABE systems as described in Sheriff et al., 2022 (CaSLiDE, Fig. 4). Using this method, 46 directed evolution cycles were performed to generate a library of UnlCasl2fl-ABE8e (Casl2f-ABE) with improved editing efficiency compared to the original Casl2f-ABE (Fig. 5B). The evolved library was then further enriched for four cycles, by performing adapted SLiDE with a high-fidelity PCR (Fig. 5C). This library was then screened with the method of the present invention, in which the original UnlCasl2fl- ABE8e (wild type, WT) was used as control. In total, the screen yielded 3,606 UMI-clusters with 50 or more reads, of which 123 clusters were identified as WT control (Fig. 7A). The determined median editing rates of the WT were 0.5% (target site 1), 16% (target site 2), and 6.2% (target site 3) (Fig. 7B). Of the 3,483 non-WT clusters, 58 had editing rates of more than 90% on all 3 target sites (Fig. 7C).
To validate these results, three variants with high activity on all three target sites (clusters 2 (SEQ ID NO: 28), 3030 (SEQ ID NO: 29), 3301 (SEQ ID NO: 30)); Fig. 7A) from the screened Casl2f-ABE library were extracted and evaluated using a base editing assay, schematically shown in Fig. 5 A. In short, the assay is performed by digesting the CaSLiDE evolution plasmid with restriction enzymes that recognize the target sites (quadratic box, Fig 5A). Successfully edited plasmids
(quadratic box with dash, Fig 5A) do not contain the restriction sites. Non-edited target sites are cut by the restriction enzyme (RE), while edited target sites are not cut due to the change in the restriction site. This results in different DNA fragments, which are visualized using agarose gel electrophoresis. The editing dependent difference in the fragments makes it possible to quantify the amount of edited and non-edited fragments. The assay showed that the variants have editing rates on all three target sites simultaneously of 91.1%, 60.8%, and 82.6%, while for the WT 0% was recorded (Figs. 8A and 8B). The variants were also tested on genomic DNA of E. coli. Three different sgRNAs were tested with the selected variants or the WT control. Editing results were analyzed by Sanger sequencing (Figs. 9A and 9B). The quantified base editing rates further validated the superior editing efficiencies of the variants identified using the method of the present invention over WT.
Example 12: Screening of evolved tyrosine site-specific recombinases
High-throughput sequencing of known tyrosine site-specific recombinases (Panto, Dre, Cre, Vika and VCre) and novel tyrosine site-specific recombinases (termed YR1, YR2, etc.). All combinations of the 13 examined recombinases and their respective 13 target sites on 169 (13x13) individual vectors were produced. Specifically, the 13 target site sequences were cloned individually into an expression vector (pEVO vector described in Buchholz and Stewart, 2001). The resulting constructs were pooled and linearized to clone in a pool of the 13 recombinase coding sequences in one ligation reaction. After an overnight culture and induction of recombinase expression, plasmid DNA was retrieved and fragments carrying the recombinase sequence and target site sequence were excised. Using nanopore sequencing, a total of 417,769 reads was obtained containing both, the specified recombinases and target sites. All possible 169 combinations of recombinases and target sites were identified with a minimum coverage of 224 reads. Using this data, the recombination rates for the individual recombinases on all target sites was calculated, providing a specificity profile for each recombinase (Fig. 11). Detailed recombination rates are shown in table 1 below.
Table 1 : Recombination rates for 13 tested recombinases on 13 target sites
Cited non-patent literature
Abi-Ghanem J., Chusainow J., Karimova M., Spiegel C., Hofmann-Sieber H., Hauber J., Buchholz F., Pisabarro M.T. (2012). Engineering of a target site-specific recombinase by a combined evolution- and structure-guided approach. Nucleic Acids Research 41, 2394-2403.
Anzalone A.V., Koblan L.W., Liu D.R. (2020). Genome editing with CRISPR-Cas nucleases, base editors, transposases and prime editors. Nat Biotechnol 38, 824-844.
Buchholz F., Hauber J. (2011). In vitro evolution and analysis of HIV-1 LTR-specific recombinases. Methods 53, 102-109.
Buchholz, F., Stewart, A.F. (2001). Alteration of Cre recombinase site specificity by substrate-linked protein evolution. Nat Biotechnol 19, 1047-1052.
Cong L., Ann Ran F., Cox D., Lin S. et al. (2013). Multiplex Genome Engineering Using CRISPR/Cas Systems. Science 339(6121), 819-823.
Danecek P., Bonfield J.K., Liddle J., Marshall J., Ohan V., Pollard M.O., et al. (2021). Twelve years of SAMtools and BCFtools. GigaScience 10(2): 1-4.
Gaj T., Gersbach C.A., Barbas, C.F. (2013). ZFN, TALEN and CRISPR/Cas-based methods for genome engineering. Trends Biotechnol. 31(7), 397-405.
Hoersten J., Ruiz-Gomez G., Lansing F., Rojo-Romanos T., Schmitt L.T., Sonntag J., Pisabarro M.T., Buchholz F. (2021). Pairing of single mutations yields obligate Cre-type site-specific recombinases. Nucleic Acids Research; available from: https://doi.org/10.1093/nar/gkabl240.
Karpinski J., Hauber I., Chemnitz J., Schafer C., Paszkowski-Rogacz M., Chakraborty D., Beschomer N., Hofmann-Sieber H., Lange U.C., Grundhoff A. (2016). Directed evolution of a recombinase that excises the provirus of most HIV-1 primary isolates with high specificity. Nature Biotechnology 34, 401-409.
Karst S.M., Ziels R.M., Kirkegaard R.H., Sorensen E.A., McDonald D., Zhu Q., Knight R., Albertsen M. (2021). High-accuracy long-read amplicon sequences using unique molecular identifiers with nanopore or PacBio sequencing. Nat Methods 18: 165-169; available from: https://doi.org/10.1038% 2Fs41592-020-01041-y.
Kluesner M.G., Nedveck D.A., Lahr W.S., Garbe J.R., Abrahante J.E., Webber B.R., et al. (2018). EditR: A Method to Quantify Base Editing from Sanger Sequencing. CRISPR J. 1 (3):239— 50.
Lansing F., Mukhametzyanova L., Rojo-Romanos T., Iwasawa K., Kimura M., Paszkowski-Rogacz M., Karpinski J., Grass T., Sonntag J., Schneider P.M. (2022). Correction of a Factor VIII
genomic inversion with designer-recombinases. Nature Communications 13; available from: https://doi.org/10.1038%2Fs41467-022-28080-7.
Lansing F., Paszkowski-Rogacz M., Schmitt L.T., Schneider P.M., Romanos T.R., Sonntag J., Buchholz F. (2019). A heterodimer of evolved designer-recombinases precisely excises a human genomic DNA locus. Nucleic Acids Research 48, 472-485.
Lawrence M., Huber W., Pages H., Aboyoun P., Carlson M., Gentleman R., Morgan M.T., Carey V.J. (2013). Software for computing and annotating genominc ranges. PLoS Comput. Biol. 9, 1-10.
Li H. (2021). New strategies to improve minimap2 alignment accuracy. Bioinformatics 37(23):4572- 4.
Lodes M.J., Dillon D.C., Houghton R.L., Skeiky Y.A. (2004). Expression cloning. Methods in Molecular Medicine. 94: 91-106.
Meinke G., Bohm A., Hauber J., Pisabarro M.T. and Buchholz F. (2016). Cre Recombinase and Other Tyrosine Recombinases. Chem Rev, 116, 12785-12820.
Rognes T., Flouri T., Nichols B., Quince C., Mahe, F. (2016). VSEARCH: a versatile open source tool for metagenomics. PeerJ 2016, 1-22.
Sarkar L, Hauber L, Hauber J., Buchholz F. (2007). HIV-1 Proviral DNA Excision Using an Evolved Recombinase. Science 316, 1912-1915.
Schindelin J., Arganda-Carreras L, Frise, E., Kaynig V. et al. (2012). Fiji: an open-source platform for biological-image analysis. Nature Methods, 9, 676-682.
Sheriff A., Guri L, Zebrowska P. et al. (2022). ABE8e adenine base editor precisely and efficiently corrects a recurrent COL7A1 nonsense mutation. Nature: Scientific Reports 12, 19643.
Stoddard B.L. (2006). Homing endonuclease structure and function. Quaterly Reviews of Biophysics, Cambridge University Press, 38(1).
Tange O. (2018). GNU Parallel 2018; available from: https://zenodo.org/record/1146014.
Van der Loo M.P.J. (2014). The stingdist package for approximate string matching. The R Journal, 6, 111-122.
Vaser R., Sovic L, Nagarajan N., Sikic, M. (2017). Fast and de novo genome assembly from long uncorrected reads. Genome Res. 27, 737-746.
Wang Y ., Zhao Y ., Bellas A., Wang Y, Au K.F. (2021). Nanopore sequencing technology, bioinformatics and applications. Nature Biotechnology 39, 1348-1365.
Wickham H., Averick M., Bryan J., et al. (2019). Welcome to the Tidyverse. J. Open Source Softw. 4, 1686.
Xin C., Yin J., Yuan S., Ou L., Liu M., Zhang W., Hu J. (2022). Comprehensive assessment of miniature CRISPR-Casl2f nucleases for gene disruption. Nature Communications, 13, Article no.: 5623.
Zurek P.J., Knyphausen P., Neufeld K., Pushpanath A., Hollfelder F. (2020). UMI-linked consensus sequencing enables phylogenetic analysis of directed evolution. Nature Communications 11;
available from: https://doi.org/10.1038/s41467-020-19687-9.
Claims
1. Method for screening and sequencing of a plurality of DNA modifying enzymes, the method comprising the steps of: providing a library of expression vectors comprising a plurality of different vectors, wherein each vector comprises a first region encoding a variant of a DNA modifying enzyme, and a second region comprising one or more target sites of a DNA modifying enzyme; introducing the library of expression vectors into host cells; culturing the host cells and expressing the DNA modifying enzymes; isolating plasmid DNA from the culture of host cells; sequencing the first and the second region of the expression vectors; and determining, based on the sequencing results, whether the DNA sequence of the second region on the expression vector has been altered by the DNA modifying enzyme.
2. Method for screening and sequencing of a plurality of target sites for DNA modifying enzymes, the method comprising the steps of: providing a library of expression vectors comprising a plurality of different vectors, wherein each vector comprises a first region encoding a DNA modifying enzyme, and a second region comprising variants of one or more target sites of a DNA modifying enzyme; introducing the library of expression vectors into host cells; culturing the host cells and expressing the DNA modifying enzymes; isolating plasmid DNA from the culture of host cells; sequencing the first and the second region of the expression vectors; and determining, based on the sequencing results, whether the DNA sequence of the second region on the expression vector has been altered by the DNA modifying enzyme.
3. The method according to claim 1 or 2, wherein the first region further comprises a unique molecular identifier.
4. The method according to claim 3, wherein the unique molecular identifier is an oligonucleotide comprising at least 50 random nucleotides.
5. The method according to claim 3 or 4, wherein the unique molecular identifier is located in the first region of the expression vector adjacent to the sequence encoding the DNA modifying enzyme.
6. The method according to any one of claims 3 to 5, further comprising the steps of: clustering the unique molecular identifiers, generating and polishing consensus sequences, determining the number of DNA modification events for each DNA modifying enzyme; and determining an activity rate for each DNA modifying enzyme.
7. The method according to any one of claims 1 to 6, wherein the first and the second region of the vector are sequenced in a single step.
8. The method according to any one of claims 1 to 7, wherein sequencing of the first and the second region of the expression vectors comprises nanopore sequencing.
9. The method according to any one of claims 1 to 8, wherein the DNA modifying enzyme is selected from the group consisting of a recombinase, an integrase, an adenosine base editor (ABE), a zinc -finger nuclease, a transcription activator-like effector nuclease and a Cas nuclease.
10. The method according to any one of claims 1 to 9, wherein the DNA modifying enzyme comprises more than one subunit.
11. The method according to claim 10, wherein the DNA modifying enzyme comprises at least two different subunits.
12. The method according to any one of claims 1 to 11, wherein the first and the second region of the expression vectors are excised from the expression vector before the sequencing.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP23171664 | 2023-05-04 | ||
| PCT/EP2024/060393 WO2024227603A1 (en) | 2023-05-04 | 2024-04-17 | High-throughput screening and sequencing method |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4705449A1 true EP4705449A1 (en) | 2026-03-11 |
Family
ID=86330935
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24720464.7A Pending EP4705449A1 (en) | 2023-05-04 | 2024-04-17 | High-throughput screening and sequencing method |
Country Status (4)
| Country | Link |
|---|---|
| EP (1) | EP4705449A1 (en) |
| CN (1) | CN121152876A (en) |
| TW (1) | TW202505032A (en) |
| WO (1) | WO2024227603A1 (en) |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US7422889B2 (en) | 2004-10-29 | 2008-09-09 | Stowers Institute For Medical Research | Dre recombinase and recombinase systems employing Dre recombinase |
| JP5336592B2 (en) | 2009-06-08 | 2013-11-06 | 公益財団法人かずさDna研究所 | Site-specific recombination method using a novel site-specific recombination enzyme and its recognition sequence |
| EP2576798B1 (en) * | 2010-05-27 | 2016-04-27 | Heinrich-Pette-Institut Leibniz-Institut für experimentelle Virologie-Stiftung bürgerlichen Rechts - | Tailored recombinase for recombining asymmetric target sites in a plurality of retrovirus strains |
| EP2690177B1 (en) | 2012-07-24 | 2014-12-03 | Technische Universität Dresden | Protein with recombinase activity for site-specific DNA-recombination |
-
2024
- 2024-04-17 CN CN202480029004.4A patent/CN121152876A/en active Pending
- 2024-04-17 WO PCT/EP2024/060393 patent/WO2024227603A1/en not_active Ceased
- 2024-04-17 EP EP24720464.7A patent/EP4705449A1/en active Pending
- 2024-04-17 TW TW113114270A patent/TW202505032A/en unknown
Also Published As
| Publication number | Publication date |
|---|---|
| WO2024227603A1 (en) | 2024-11-07 |
| TW202505032A (en) | 2025-02-01 |
| CN121152876A (en) | 2025-12-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| CN107109422B (en) | Genome editing using split Cas9 expressed from two vectors | |
| CN116179513B (en) | A Cpf1 protein and its application in gene editing | |
| US20180258409A1 (en) | Recombinase mutants | |
| US20240287484A1 (en) | Systems, compositions, and methods involving retrotransposons and functional fragments thereof | |
| US20200131504A1 (en) | Plasmid library comprising two random markers and use thereof in high throughput sequencing | |
| CN119842677A (en) | Cytosine deaminase with high editing activity and no sequence preference for structure-oriented mining and application thereof | |
| US20200347441A1 (en) | Transposase compositions, methods of making, and methods of screening | |
| Takenaka et al. | RNA editing in plant mitochondria: assays and biochemical approaches | |
| US20240327871A1 (en) | Systems and methods for transposing cargo nucleotide sequences | |
| EP3436471B1 (en) | Recombinase mutants | |
| CN106589134A (en) | Chimeric protein pAgoE, construction method and applications thereof, chimeric protein pAgoE using guide, and construction method and applications thereof | |
| CN118006584A (en) | Programmable nucleases completely lacking Cas1, Cas2 and Cas4 in CRISPR loci and their applications | |
| EP4458963A1 (en) | Highly active crispr base editors obtained through cas-assisted substrate-linked directed evolution (caslide) | |
| JP2024501892A (en) | Novel nucleic acid-guided nuclease | |
| EP4705449A1 (en) | High-throughput screening and sequencing method | |
| WO2024227911A2 (en) | Highly active crispr base editors obtained through cas-assisted substrate-linked directed evolution (caslide) | |
| US20240360477A1 (en) | Systems and methods for transposing cargo nucleotide sequences | |
| WO2026091549A1 (en) | Cas12a homologue pscas12a and mutant thereof, and use thereof in gene editing | |
| EP4630542A2 (en) | Retrotransposon compositions and methods of use | |
| WO2024124204A2 (en) | Retrotransposon compositions and methods of use | |
| WO2025160202A1 (en) | Engineered large serine recombinases | |
| Ciobanu et al. | Single Cell Genomics and Transcriptomics for Unicellular Eukaryotes | |
| Wang et al. | Genome editing of model oleaginous microalgae | |
| HK1261197B (en) | Recombinase mutants | |
| HK1261197A1 (en) | Recombinase mutants |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251203 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |