EP4615970A1 - Micro-compartmentalized ultra-high throughput screening from single copy gene libraries - Google Patents

Micro-compartmentalized ultra-high throughput screening from single copy gene libraries

Info

Publication number
EP4615970A1
EP4615970A1 EP23804695.7A EP23804695A EP4615970A1 EP 4615970 A1 EP4615970 A1 EP 4615970A1 EP 23804695 A EP23804695 A EP 23804695A EP 4615970 A1 EP4615970 A1 EP 4615970A1
Authority
EP
European Patent Office
Prior art keywords
nucleic acid
microcompartments
protein
dna
activity
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
EP23804695.7A
Other languages
German (de)
French (fr)
Inventor
Yannick RONDELEZ
Thibault DI MEO
Christophe DANELON
Zhanar ABIL
Margarida GOMES
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Centre National de la Recherche Scientifique CNRS
Ecole Superieure de Physique et Chimie Industrielles de Ville de Paris ESPCI
Universite Paris Sciences et Lettres
Technische Universiteit Delft
Original Assignee
Centre National de la Recherche Scientifique CNRS
Ecole Superieure de Physique et Chimie Industrielles de Ville de Paris ESPCI
Universite Paris Sciences et Lettres
Technische Universiteit Delft
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Centre National de la Recherche Scientifique CNRS, Ecole Superieure de Physique et Chimie Industrielles de Ville de Paris ESPCI, Universite Paris Sciences et Lettres, Technische Universiteit Delft filed Critical Centre National de la Recherche Scientifique CNRS
Publication of EP4615970A1 publication Critical patent/EP4615970A1/en
Pending legal-status Critical Current

Links

Classifications

    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12NMICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
    • C12N15/00Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
    • C12N15/09Recombinant DNA-technology
    • C12N15/10Processes for the isolation, preparation or purification of DNA or RNA
    • C12N15/1034Isolating an individual clone by screening libraries
    • C12N15/1075Isolating an individual clone by screening libraries by coupling phenotype to genotype, not provided for in other groups of this subclass
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12PFERMENTATION OR ENZYME-USING PROCESSES TO SYNTHESISE A DESIRED CHEMICAL COMPOUND OR COMPOSITION OR TO SEPARATE OPTICAL ISOMERS FROM A RACEMIC MIXTURE
    • C12P19/00Preparation of compounds containing saccharide radicals
    • C12P19/26Preparation of nitrogen-containing carbohydrates
    • C12P19/28N-glycosides
    • C12P19/30Nucleotides
    • C12P19/34Polynucleotides, e.g. nucleic acids, oligoribonucleotides
    • CCHEMISTRY; METALLURGY
    • C12BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
    • C12PFERMENTATION OR ENZYME-USING PROCESSES TO SYNTHESISE A DESIRED CHEMICAL COMPOUND OR COMPOSITION OR TO SEPARATE OPTICAL ISOMERS FROM A RACEMIC MIXTURE
    • C12P21/00Preparation of peptides or proteins
    • C12P21/02Preparation of peptides or proteins having a known sequence of two or more amino acids, e.g. glutathione

Definitions

  • the present invention relates to the fields of directed evolution of proteins, protein engineering and screening.
  • Directed evolution (DE) of proteins is an engineering approach, inspired by natural evolution, that allows protein and enzyme engineering without requiring precise knowledge concerning the structural, functional or mechanistic aspects of the protein. It is performed by cycles of genetic diversification and enrichment: starting from a known gene with a given sequence, or a small set of these, the first step is to introduce some diversity so as to obtain a collection (library) of genes that are different in their nucleotide sequences but related to the parental one. These genes are called variants. Various methods are then used to enrich this library in variants with the desired properties.
  • This screening or selection stage performed on the encoded protein variant’s properties eventually allows to recover the genes coding for proteins with the most interesting properties (for example improved catalytic rate, or higher stability, non-natural specificities, etc.), which can be used for a subsequent round of diversification and enrichment. It is also possible to couple this approach with DNA sequencing to identify the beneficial or detrimental mutations with respect to the targeted property.
  • directed evolution has been applied to improve or alter the properties of a large variety of enzymes. New methodologies have been introduced to manipulate always larger libraries and to screen various types of activities (such as enzymatic activity, fluorescence, binding to a particular receptor/ligand). In particular, the use of microcompartments has enabled ultrahigh throughput screening methods.
  • Micro-compartmentalized ultrahigh throughput screening (uHTS) methods use random encapsulation of variants in individual microcompartments to manipulate large libraries of variants or candidate genes. This random encapsulation is done at a concentration that is low enough that many microcompartments end up receiving 0 or 1 member of the library, and only a small proportion of the microcompartments receives more than 1 member (e.g. 2 or more). This process is referred as Poissonian partitioning and enables clonality, i.e. the fact that the signal observed in most microcompartments is associated to a unique variant of the library. Because of that, the microcompartments obtained from Poissonian partitioning of variants are sometime called “monoclonal” microcompartments.
  • microcompartments can be for example water-in-oil microdroplets, generated and manipulated using microfluidic devices, or liposomes. The volume of these microcompartments ranges between the femtoliter and the microliter.
  • phage display One approach for genotype- phenotype linkage in the absence of a cell is to create a direct physical connection between the encoding genes and their gene products, so called ‘display systems'.
  • Various display systems have been developed over the years to perform this function: phage display, ribosome display, mRNA display, CIS display, SNAP display, etc.
  • these ‘display systems' provide only one, or just a few copies of the protein of interest attached to each gene variant and in some cases, e.g., phage display, still require an in vivo step which is performed within live transformed cells.
  • Such ‘display system’ can be used for selection based on affinity (panning), but are in many cases insufficient to assess the catalytic properties of an enzyme, or other functional protein properties (e.g., fluorescence). Indeed, it is generally difficult to detect enzymatic activity or other protein properties from a single molecule or just a few molecules.
  • the microcompartments may serve this function, in addition to their role of separating the genetic variants from one-another.
  • a general approach for in vitro uHTS would require to obtain better performance of the IVTT reaction in each compartment, i.e. produce more protein from the unique copy of the encapsulated gene and hence obtain more signal related to the activity of this protein within the sorting apparatus. Indeed, when a sufficient amount of protein is produced in the compartment, it becomes easier to differentiate positive and negative microcompartments (e.g., those displaying an interesting enzymatic activity from those that do not) and therefore to recover only the microcompartments containing the genes encoding active protein variants.
  • current in vitro methods to obtain multiple copies of the same genetic variants in each microcompartment, while maintaining clonality are cumbersome.
  • Several methods have been developed to mitigate the poor expression levels arising from using a single molecule of DNA template, which are based on pre-amplifying the gene variants in a clonal fashion.
  • the first approach uses isothermal amplification that creates concatemers.
  • rolling circle amplification starts from a single circular DNA molecule containing a gene and creates multiple catenated copies of the same gene.
  • a single molecule contains multiple copies of the gene information, thus increasing the gene concentration even for single encapsulation events.
  • preprocessing steps are necessary for concatemer generation and, in the case of a library, it is not guaranteed that single DNA molecules contain copies of a single variant of the library.
  • a second approach uses droplet-PCR amplification from single encapsulated gene followed by picoinjection of the IVTT (in vitro transcription and translation) mix (Mazutis, L. et al. Droplet-Based Microfluidic Systems for High-Throughput Single DNA Molecule Isothermal Amplification and Analysis. Anal Chem 81 , 4813-4821 (2009)).
  • a 2021 study describes a method for cell- free directed evolution at ultrahigh throughput using such a droplet-based microfluidic device.
  • the library of gene variants in a circular form is distributed in microdroplets following Poissonian partitioning (where the concentration of the library was adjusted so that -10% of the droplet would contain a single variant and most droplet would contain no variant) and first amplified by rolling circle amplification.
  • the droplets are then pico-injected with an IVTT mixture for protein expression. Then, these droplets are picoinjected again with the substrate of the enzyme of interest.
  • the resulting droplets can be sorted according to the fluorescence signal associated with the enzymatic transformation of the substrate (Holstein, J. M., Gylstorff, C. & Hollfelder, F. Cell-free Directed Evolution of a Protease in Microdroplets at Ultrahigh Throughput. ACS Synthetic Biology acssynbio.0c00538 (2021 ) doi:10.1021 /acssynbio.0c00538). This tedious protocol with multiple microfluidic steps is technically challenging and limits the throughput of the method.
  • a third approach avoids the cumbersome multiple microfluidic manipulation of droplets, by using so-called bead-display (or microbead display, see A Sepp, DS Tawfik, AD Griffiths.
  • Microbead display by in vitro compartmentalisation selection for binding using flow cytometry. FEBS letters 532 (3), 455-458).
  • the gene amplification for example by PCR
  • expression reactions are performed beforehand, on beads, in a clonal fashion.
  • the resulting microbeads carry both multiple clonal copies of the gene, and multiple copies of the corresponding protein. As such, they physically maintain the phenotype-genotype linkage and can be manipulated in bulk.
  • these beads are subsequently encapsulated in droplets with a fluorogenic substrate to assay the activity of the carried protein, and these droplets can be sorted according to detected activity levels, in order to recover the beads carrying the most active protein variants and ultimately the genes encoding these variants.
  • the protocol for bead display also involves multiple preparatory steps before the screening can be performed.
  • a bead display technique allowing the display of up to millions of copies of the protein of interest together with its gene was developed in 2013, based on the SNAP-tag system (Diamante, L., Gatti-Lafranconi, P., Schaerli, Y. & Hollfelder, F. In vitro affinity screening of protein and peptide binders by megavalent bead surface display. Protein engineering, design & selection: PEDS 26, 713-724 (2013)).
  • This approach requires a complex multistep protocol for bead preparation, encapsulation and sorting, which increases the workload and introduces many possibilities for failure.
  • a fourth option which is particularly desirable in terms of its simplicity, would be a “one- pot” process combining amplification/replication of a single compartmentalized nucleic acid molecule, with protein expression, all in the same mixture without external intervention.
  • Combining replication of DNA with expression systems (IVTT, such as the PURE system) in a single compartment has previously been attempted.
  • IVTT expression systems
  • significant difficulties arose due to the low compatibility of expression systems (including the PURE system) with DNA replication.
  • PCR cannot be used in this context because the thermal cycles would destroy the protein expression machinery.
  • this compatibility can be improved by selecting specific amplification schemes and adjusting the reaction conditions (e.g. decreasing ribonucleotide triphosphate concentration), known approaches still require high DNA concentrations to be efficient.
  • the present invention fulfils this need. Indeed, the present Inventors have developed a new method that is specifically designed to amplify a single copy of a nucleic acid molecule encoding a protein (such as a nucleic acid in a standard linear form (e.g., PCR products)) and simultaneously express the encoded protein and assay its activity within a microcompartment.
  • the method is based on the combined use of an IVTT system and a system capable of performing nucleic acid replication and amplification in vitro (i.e. in a non-cellular context).
  • a replicator linear DNA comprising, in addition to the gene to be encoded and its regulatory sequences, origins of replication (such as phage Phi29 origins of replication or derivatives thereof) in 5’ of each DNA strand is used, and a minimal replication machinery (such as phage Phi29 p2, p3, p5 and p6 proteins) is added, either as purified proteins or as DNA encoding them.
  • origins of replication such as phage Phi29 origins of replication or derivatives thereof
  • a minimal replication machinery such as phage Phi29 p2, p3, p5 and p6 proteins
  • the decoupled replication-expression can be used to express any protein of interest, or any library of DNA fragments (e.g. encoding variants of the protein of interest); ii) that this modified reaction can be performed starting from a very low concentration (preferably using poissonian partitioning of single nucleic acid molecules into microcompartments) of replicating DNA carrying the gene(s) of interest; iii) that it is possible to couple replication with transcription/translation ; iv) that this reaction is compatible with enzymatic assays, in particular fluorogenic assays, v) that when performed in micro-compartment starting from sufficiently low initial concentration of library of DNA fragments (preferably using poissonian partitioning), it results in a population of clonal micro-compartments containing a high copy number of one of DNA fragment from the library and a high copy number of its encoded protein, and producing a signal related to the activity of the encapsulated protein, which can be used in a
  • Poissonian partitioning enables clonality, i.e., ensures that only one copy of nucleic acid molecule (i.e., a unique variant of the library of the library to be tested) is incorporated in the majority of micro-compartments.
  • the present invention thus provides an original, reliable, rapid and efficient method to perform clonal in vitro microcompartmentalized uHTS via a one-step encapsulation of nucleic acid molecules (such as linear DNA constructs) containing the genetic variant (or candidate sequence) to be tested, in a molecular mixture that allows for simultaneous specific gene amplification and protein expression and optionally activity/enzymatic assay.
  • the present invention thus relates to a novel method for screening a library of nucleic acid molecules each encoding a candidate polypeptide for a polypeptide with an activity of interest by direct encapsulation.
  • a first composition comprising an in vitro platform comprising a nucleic acid replication machinery, optionally a nucleic acid transcription machinery, and a nucleic acid translation machinery
  • a second composition comprising a library of nucleic acid molecules each comprising at least one candidate sequence encoding at least one candidate polypeptide, are provided and mixed.
  • the thus obtained mixture is partitioned (preferably using poissonian partitioning) into microcompartments wherein at least one of said microcompartments comprises a single copy of said nucleic acid molecule.
  • Poissonian partitioning enables clonality, i.e., the fact that most microcompartments contains a unique variant of the library (i.e., only one copy of a nucleic acid molecule of the library to be tested is incorporated in most of the microcompartments). Then, in the microcompartments, the candidate sequence is amplified by the nucleic acid replication machinery in the microcompartments, optionally transcribed by the nucleic acid transcription machinery, and translated in at least one candidate polypeptide by the nucleic acid translation machinery.
  • the activity of the thus obtained candidate polypeptide is then detected and the candidate sequence encoding the candidate polypeptide having the activity of interest is finally recovered using a sorting apparatus.
  • the present invention also concerns a kit comprising an in vitro platform as defined above.
  • the kit also comprises at least one ingredient for making microcompartments.
  • directed evolution refers to an engineering approach, inspired by natural evolution, that allows protein and enzyme engineering without requiring precise knowledge concerning the structural, functional or mechanistic aspects of the protein. It is performed by cycles of genetic diversification and enrichment: starting from a known gene with a given sequence, or a small set of these, the first step is to introduce some diversity so as to obtain a collection (library) of genes that are different in their nucleotide sequences but related to the parental one. Various methods are then used to enrich this library in variants with the desired properties. This screening or selection stage performed on the encoded protein variant’s properties eventually allows to recover the genes coding for proteins with the most interesting properties, which can be used for a subsequent round of diversification and enrichment.
  • bioprospection and “functional metagenomics” are used interchangeably and refer to another approach for the identification of proteins with desired activity.
  • bioprospection the screened libraries are libraries of natural microorganism, or cells or can be isolated nucleic acid sequences. These libraries of nucleic acid strands may be obtained by extraction and cloning of nucleic acides coming from microorganisms, plants or animals or a plurality of those.
  • screening refers to a process for testing and selecting compounds/active agents (in particular proteins and/or nucleic acid molecule encoding such proteins) for a specific effect/activity.
  • candidate sequence it is herein meant a genetic sequence (i.e., a nucleotide sequence) encoding a candidate polypeptide whose activity is to be tested in the method of the invention.
  • candidate sequences also referred to as genetic variants or mutants
  • polypeptide variants or mutants i.e the original polypeptide with one or more mutations selected from substitutions, insertions, deletions and permutations
  • polypeptide variants or mutants i.e the original polypeptide with one or more mutations selected from substitutions, insertions, deletions and permutations
  • a “library” of nucleic acid molecules refers to a mixture of several and preferably many (such as one or more hundreds up to billions or more) of distinct nucleic acid molecules.
  • partitioning refers to separating a sample into a plurality of portions, also referred to as “partitions” or “compartments”. Partitions are generally physical, such that a sample in one partition does not, or does not substantially, mix with a sample in an adjacent partition.
  • micro-compartment or “microcompartment” herein means a compartment with an inner volume between the femtoliter and the microliter (i.e., a characteristic size below 1 mm).
  • Microcompartments include, but are not limited to, microdroplets (such as water-in-oil microdroplets), vesicles (such as liposomes), microchambers, etc. It may be advantageous that the compartments used in a method of screening according to the invention are identical in shape and volume, however, collection of microcompartiments of different size (obtained for example by liposome forming protocol, or by bulk emulsification protocols) can also be used. Methods of obtaining microcompartments are well known of the skilled person.
  • microdroplets and emulsions such methods include various emulsification techniques, such as membrane emulsification, mechanical emulsification, microfluidic emulsification, high homogenization, microfluidization and ultrasonication, among others.
  • Emulsions are generally stabilized by appropriate surfactants known to the skilled person, which can provide them with long term stability.
  • microdroplet can be obtained as mono-disperse emulsions using microfluidic approaches or a number of other approaches.
  • micro-com partmentalisation refers to the separation of reactions into multiple microreactors, such as multiple micro-compartments.
  • nucleic acid refers to any polynucleotide or nucleic acid molecule.
  • nucleotide residues also called nucleotide residues
  • Nucleotide monomers are composed of a nucleobase, a sugar or a sugar analogue (such as but not limited to ribose or 2'-deoxyribosefor RNA and DNA, but other sugars or sugar analogues may be used for xeno nucleic acids referred to as XNA, such as 1 ,5-anhydrohexitol for HNA, cyclohexene for CeNA, threose for TNA, glycol for GNA, a ribose modified with an extra bridge connecting the 2' oxygen and 4' carbon for locked nucleic acid or LNA, or N-(2-aminoethyl)-glycine units for peptide nucleic acid or PNA), and one to three phosphate groups.
  • XNA xeno nucleic acids
  • XNA such as 1 ,5-anhydrohexitol for HNA, cyclohexene for CeNA, threos
  • a polynucleotide is formed through phosphodiester bonds linking ribose or deoxyribose groups in the case of natural nucleic acids, but other artificial bonds may be used in the case of XNA (such as peptide bonds in the case of PNA) between the individual nucleotide monomers.
  • Nucleic acid molecules include, but are not limited to, ribonucleic acid (RNA), deoxyribonucleic acid (DNA), and mixtures thereof such as, e.g., RNA-DNA hybrids (mixed polyribo- polydeoxyribonucleotides).
  • RNA-DNA hybrids e.g., RNA-DNA hybrids.
  • a polynucleotide may comprise non- naturally occurring nucleotides and may be interrupted by non-nucleotide components.
  • Exemplary DNA nucleic acids include without limitations, complementary DNA (cDNA), genomic DNA, plasmid DNA, DNA vector, viral DNA (e.g., viral genomes, viral vectors), oligonucleotides, probes, primers, satellite DNA, microsatellite DNA, coding DNA, non-coding DNA, antisense DNA, and any mixture thereof.
  • cDNA complementary DNA
  • genomic DNA genomic DNA
  • plasmid DNA DNA vector
  • viral DNA e.g., viral genomes, viral vectors
  • oligonucleotides e.g., probes, primers
  • satellite DNA e.g., microsatellite DNA
  • coding DNA e.g., non-coding DNA, antisense DNA, and any mixture thereof.
  • RNA nucleic acids include, without limitations, messenger RNA (mRNA), precursor messenger RNA (pre-mRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), RNA vector, viral RNA, guide RNA (gRNA), antisense RNA, coding RNA, non-coding RNA, antisense RNA, satellite RNA, small cytoplasmic RNA, small nuclear RNA, etc.
  • mRNA messenger RNA
  • pre-mRNA precursor messenger RNA
  • siRNA small interfering RNA
  • shRNA short hairpin RNA
  • miRNA microRNA
  • RNA vector viral RNA
  • guide RNA guide RNA
  • antisense RNA antisense RNA
  • coding RNA non-coding RNA
  • antisense RNA satellite RNA
  • small cytoplasmic RNA small nuclear RNA, etc.
  • Polynucleotides described herein may be synthesized by standard methods known in the art, e.g., by use of an automated DNA synthesizer (such as those that are commercially available from Biosearch, Applied Biosystems, etc.) or by DNA assembly and gene synthesis method, or by mutagenesis methods or obtained from a naturally occurring source (e.g., a genome, cDNA, etc.) or an artificial source (such as a commercially available library, a plasmid, etc.) using molecular biology techniques well known in the art (e.g., cloning, PCR, etc.).
  • the nucleic acids can be synthesized chemically, e.g., in accordance with the phosphotriester method (see, for example, Uhlmann, E. & Peyman, A. (1990) Chemical Reviews, 90, 543-584).
  • a “replicator” is a nucleic acid molecule able to replicate under specific conditions in the presence of specific compounds (the “nucleic acid replication machinery”).
  • the replicator acts as a template for its amplification, so that the newly-made nucleic acid molecules are copies or reverse-complemented copies of it.
  • concentration of the replicator will increase. For example, if one starts from a single molecule of the replicator, after some time, 2, then 4, 8 etc molecules may be present, i.e. its concentration will increase exponentially.
  • a replicator can be any type of nucleic acid (DNA, RNA or hybrids or XNA) and may contain specific subsequences or non- canonical modifications.
  • a replicator preferably includes a subsequence selected from an origin of replication (OR), a primer binding site, one or more genes with appropriate regulatory sequences (e.g. a promoter), a ribosome-binding site (RBS), barcodes, etc.
  • OR origin of replication
  • RBS ribosome-binding site
  • protein and “polypeptide” are used interchangeably herein and refer to any peptide-bond-linked polymer of amino acids, regardless of length or post-translational modification. These terms preferably refer to polymers of amino acid residues comprising at least six amino acids covalently linked by peptide bonds.
  • the polymer can be linear, branched or cyclic.
  • the polymer may comprise naturally occurring and/or amino acid analogues and it may be interrupted by non-amino acids. No limitation is placed on the maximum number of amino acids comprised in a polypeptide. As a general indication, the term refers to both short polymers (typically designated in the art as peptide, or protein fragment) and to longer polymers (typically designated in the art as polypeptide or protein).
  • polypeptides encompasses native polypeptides, modified polypeptides (also designated derivatives, analogues, variants, mutants), polypeptide fragments, polypeptide multimers (e.g., dimers), mutated polypeptides, engineered polypeptides, fusion polypeptides among others.
  • a polypeptide is understood to be any translational product of a polynucleotide regardless of size, and whether glycosylated or not, and includes peptides and proteins.
  • Polypeptides/Proteins usable herein can be further modified by chemical or enzymatic modification.
  • chemically modified polypeptide or enzymatically modified polypeptide comprises other chemical groups than the 20 naturally occurring amino acids.
  • Examples of such chemical or enzymatic modifications include post-translational modifications.
  • Chemical or enzymatic modifications of a polypeptide may provide advantageous properties as compared to the parent polypeptide, e.g., one or more of enhanced stability, increased biological half-life, increased water solubility, increased activity, enhanced properties, labelling, etc.
  • amino acid polymer contains more than 50 amino acid residues, it is preferably referred to as a polypeptide or a protein, whereas if the polymer consists of 50 or fewer amino acids, it is preferably referred to as a "peptide".
  • the reading and writing senses of an amino acid sequence of a polypeptide, protein and peptide as used herein are the conventional reading and writing senses.
  • the reading and writing convention for amino acid sequences of a polypeptide, protein and peptide places the amino terminus on the left, with the sequence then being written and read from the amino terminus (N-terminus) to the carboxyl terminus (C-terminus), from left to right.
  • telomere a protein or polypeptide having an activity of interest, by opposition to a non-functional protein or peptide without the activity of interest (e.g. a mutant of a functional polypeptide that has lost the activity of interest).
  • a library of protein variants some of which are functional and some of which are non-functional, is constructed and then screened using the method according to the invention, in order to recover the functional variants.
  • activity refers to any biological activity of the protein that is screened for when using the method according to the invention.
  • Activities of interest notably include enzymatic activity, reporter activity (e.g. fluorescence activity), regulatory activity (e.g. transcription factor activity) and biological activity (e.g. drug or antibiotic activity).
  • Enzymatic activity refers to the specific catalytic activity of an enzyme.
  • Enzyme herein refers to a protein with catalytic properties (enzymatic properties).
  • Most biomolecules capable of catalysing chemical reactions in cells are enzymes; however, some catalytic biomolecules are made of RNA and are therefore distinct from protein enzymes: these are ribozymes.
  • An enzyme works by lowering the activation energy of a chemical reaction, which increases the speed of the reaction. The enzyme is either not modified during the reaction or regenerated unmodified at the end of the catalytic reaction.
  • the initial molecules are the “substrates” of the enzyme, and the molecules formed from these substrates are the products of the reaction. Enzymes are often characterized by their very high specificity.
  • enzymes are generally globular proteins that act alone or in complexes of several enzymes or subunits. Like all proteins, enzymes consist of one or more polypeptide chains folded to form a three-dimensional structure corresponding to their native state. However, some enzyme can be fully or partially disordered proteins. Many enzymes are composed of more than one peptide chain or assemble as multimeric assemblies.
  • catalytic site located in the catalytic domain.
  • the catalytic site may be located in the vicinity of one or more binding sites, at which the substrate(s) is (are) bound and oriented to catalyse the chemical reaction.
  • the catalytic site and the binding sites form the active site of the enzyme.
  • Enzymes perform a large number of functions in living organisms. For example, they can be involved in signal transduction and regulation of cellular processes, in the generation of movement, in active transmembrane transport, in digestion, in metabolism, in the immune system, in nucleic acid digestion or cleavage mechanisms or in nucleic acid production (referred to here as "nucleic acid-acting enzymes"), in prodrug conversion mechanisms (prodrug-to-drug conversion).
  • the enzyme is preferably a prokaryotic, eukaryotic or viral enzyme, preferably an enzyme from an animal, plant, alga, microalgae, insect, microorganism, archaea, bacterium, parasite, yeast, fungus or virus.
  • the enzyme may be a mammalian enzyme, such as a human enzyme.
  • the various categories of enzymes are well known to the person skilled in the art, who can refer in particular to reference works in the field (such as Schomburg D., Schomburg I., Springer Handbook of Enzymes. 2 edn. Heidelberg: Springer; 2001 -2009; Liebecq C., IUPAC-IUBMB Joint Commission on Biochemical Nomenclature (JCBN) and Nomenclature Committee of IUBMB (NC-IUBMB) Biochem. Mol. Biol. Int. 1997;43:1151 -1156; IUBMB (1992), Enzyme Nomenclature 1992, Academic Press, San Diego; and specialized databases as described in Schomburg D, Schomburg I.
  • Enzyme databases in particular, the BRENDA database (available, inter alia, at brenda-enzymes.org), as described, for example, by Chang A, Schomburg I, Placzek S, Jeske L, Ulbrich M, Xiao M, Sensen CW, Schomburg D, Nucleic Acids Res. 2015 Jan;43. Epub 2014 Nov 5. BRENDA in 2015: exciting developments in its 25th year of existence). Enzymes are also classified based on their type of activity in the Enzyme Commission number (EC number) classification, and any enzyme of any subcategory of EC numbers 1 to 7 is of interest in the present invention.
  • EC number Enzyme Commission number
  • an “enzyme fragment” is any portion of an enzyme, preferably provided that such fragment/portion is capable of having an enzymatic activity.
  • the enzyme fragment preferably comprises at least 6 consecutive amino acid residues of the enzyme (and is preferably an enzyme catalytic site) (preferably at least 8 consecutive amino acid residues of the enzyme, preferably at least 10, preferably at least 15, preferably at least 20, preferably at least 30 amino acid residues of the enzyme).
  • Cofactors refers to a non-protein molecule that forms a stable or transient complex with an enzyme and whose presence is essential for the activity of an enzyme.
  • Cofactors can be inorganic ions, particularly metallic ions or other organic molecules, such as FAD or NADH.
  • Coenzymes refers to a non-protein organic molecule that forms a complex with an enzyme and whose presence is essential for the activity of an enzyme.
  • Coenzymes may be divided into two types. The first is called a “prosthetic group”, which consists of a coenzyme that is tightly (or even covalently) and permanently bound to a protein.
  • the second type of coenzymes are called “cosubstrates”, and are transiently bound to the protein. Cosubstrates may be released from a protein at some point, and then rebind later.
  • Reporter activity refers to the specific activity of a reporter protein.
  • a “reporter protein” or “reporter polypeptide” refers to a protein that is directly or indirectly detectable in the screening step of the method of the invention, preferably by optical means. Reporter proteins directly detectable in the screening step notably include coloured or fluorescent proteins.
  • fluorescent proteins are proteins that emit light when exposed to light or other radiation, the emitted light having a longer wavelength than the light to which the protein has been exposed.
  • Non-limiting examples of fluorescent proteins include blue (e.g. EBFP2, mTagBFP2), cyan (e.g.
  • UV-excitable green e.g. mT-Sapphire
  • green e.g. EGFP, mEGFP, Emerald, mEmerald, sfGFP
  • yellow-green e.g. YFP, mPapaya, YPet, Citrine, mCitrine, Venus, mVenus, Topaz, mTopaz, Clover, mClover, mNeonGreen), orange (e.g. mOrange, m0range2, mKO, mK02), orange-red (e.g.
  • Reporter proteins indirectly detectable in the screening step notably include enzymes that produce a product that can be easily detected for example by spectroscopic means (e.g. a colored, fluorescent or luminescent product) from a precursor molecule (e.g. a chromogenic compound, a fluorogenic compound or a luciferin).
  • a chromogenic compound is a compound that produces color when modified by an enzyme
  • a fluorogenic compound is a compound that produces fluorescence when modified by an enzyme
  • a “luciferin” is a compound that produces light when modified by an enzyme.
  • transcription factor refers to a protein that controls the transcription rate of a DNA to an mRNA through a specific binding mechanism. All TFs share an essential feature: they contain at least one DNA-binding domain (DBD). Examples of TFs include, but are not limited to, AP-1 , CREB, C/EBP, c-Myc, NF-1 , TFII factors, etc.
  • DBD DNA-binding domain
  • TFs include, but are not limited to, AP-1 , CREB, C/EBP, c-Myc, NF-1 , TFII factors, etc.
  • the various categories of TFs are well known to the person skilled in the art, who can refer in particular to reference works in the field (such as Yusuf, Dimas, et al.
  • transcription factor encompasses native TFs and derivatives thereof (e.g., mutated and/or engineered TFs), provided that such derivative is capable of controlling the transcription rate of a DNA to an mRNA through a specific binding mechanism.
  • TF fragment is any portion of a TF, preferably provided that such fragment/ portion is capable of controlling the transcription rate of a DNA to an mRNA through a specific binding mechanism (e.g., DBDs, etc.).
  • the TF fragment preferably comprises at least 6 consecutive amino acid residues of the TF (and is preferably a DBD) (preferably at least 8 consecutive amino acid residues of the TF, preferably at least 10, preferably at least 15, preferably at least 20, preferably at least 30 amino acid residues of the TF).
  • peptide or protein fragment or “part of a peptide or protein” or “protein domain” herein mean a portion of a peptide or protein, i.e., a portion of the sequence of consecutive amino acids making up said peptide or protein (referred to as the peptide or protein from which the fragment is derived).
  • the fragment when the fragment is a peptide, protein, the fragment preferably comprises at least 6 consecutive amino acids of the peptide or protein from which it is derived; more preferably at least 8 consecutive amino acids, more preferably at least 10 consecutive amino acids, more preferably at least 12 consecutive amino acids, more preferably at least 15 consecutive amino acids, more preferably at least 20 consecutive amino acids, more preferably at least 30 consecutive amino acids of the peptide or protein from which it is derived.
  • the fragment when the fragment is a peptide or protein fragment, the fragment preferably has a three-dimensional structure, under non-denaturing conditions (e.g., conditions that are usually non-denaturing for proteins, especially in the absence of denaturing and/or chaotropic agents).
  • the fragment is preferably a functional fragment.
  • “Functional fragment” means any peptide or protein fragment, having at least one of the original functions of the peptide or protein from which said fragment is derived.
  • the functional fragment performs said function with an efficiency equal to at least 30% of that of said peptide or protein or molecule, preferably at least 40%, preferably at least 45%, preferably at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, preferably at least 99%, preferably at least 100% of the efficacy of said peptide or protein.
  • protein fragments examples include, e.g., protein domains, protein epitopes, etc.
  • Fragments usable herein can be further modified by chemical or enzymatic modification (e.g., post- translational modification(s)).
  • chemical or enzymatic modification e.g., post- translational modification(s)
  • a chemically/enzymatically modified fragment comprises other chemical groups than the 20 naturally occurring amino acids (e.g., can comprise post-translational modification(s)).
  • protein derivative or “derivative” herein mean a mutated protein and/or an engineered protein and/or a mimetic.
  • the derivative is preferably a functional derivative.
  • “Functional derivative” means any protein derivative having at least one of the original functions of the protein from which it is derived.
  • the functional derivative performs said function with an efficiency equal to at least 30% of that of said protein, preferably at least 40%, preferably at least 45%, preferably at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, preferably at least 99%, preferably at least 100% of the efficacy of said protein.
  • three-dimensional structure or "tertiary structure” it is herein referred to the intrinsic folding of a molecule in space.
  • a molecule with a three-dimensional structure is a molecule with a stable spatial configuration, (commonly called folding), which is its own and which is, in general, intimately linked to its function.
  • folding commonly called folding
  • the molecule is said to be denatured and loses its function.
  • a three-dimensional structure is a structure with limited flexibility.
  • a non-three-dimensional structure can adopt a dynamic set of configurations that constantly change over time.
  • the three-dimensional structure is the folding of the polypeptide chain in space.
  • the three-dimensional structure is not a linear chain of amino acids that can adopt a dynamic set of configurations constantly changing over time (it is not a linear succession of amino acids without any spatial configuration).
  • the three-dimensional structure of proteins, peptides, mixed molecules comprising a protein or a peptide, or fragments of these, is maintained by different interactions which can be: covalent interactions (disulfide bridges between cysteines); electrostatic interactions (ionic bonds, hydrogen bonds); van der Waals interactions; interactions with the solvent and the environment (ions, lipids).
  • post-translational modification refers to a chemical or enzymatic modification occurring naturally or not on a protein or a protein fragment, after or concomitantly to protein translation (e.g., biological or biochemical synthesis, e.g., using cellular machinery), or after or concomitantly to protein synthesis (e.g., artificial and/or chemical synthesis).
  • protein translation e.g., biological or biochemical synthesis, e.g., using cellular machinery
  • protein synthesis e.g., artificial and/or chemical synthesis
  • post-translationally modified protein it is herein referred to a protein having at least one (i.e. , one or more) post-translational modification.
  • post-translationally modified protein fragment it is herein referred to a protein fragment having at least one (i.e., one or more) post-translational modification.
  • platform and the term “system” are herein considered equivalent.
  • a platform according to the invention is preferably a cell-free system.
  • a cell-free system can be selected from cell extracts, mixtures of purified cell components, reconstituted in vitro transcription translation (IVTT) systems, DNA amplification and replication systems, enzymatic and protein activity assay and any combination thereof.
  • the platform is designed to achieve at least two biological processes, namely isothermal genetic replication and protein expression.
  • the platform according to the invention combines an IVTT system and a system capable of performing replication in vitro to obtain an In Vitro Transcription, Translation, and Replication system (hereinafter referred to as “IVTTR” system), in the very particular context of microcompartments comprising only one nucleic acid molecule comprising a candidate sequence encoding a candidate polypeptide.
  • IVTTR In Vitro Transcription, Translation, and Replication system
  • machine refers to a biological equipment made of one or more materials that functionally cooperate to achieve a biological mechanism.
  • nucleic acid replication machinery thus comprises or consists essentially of or consists of all the required materials for completing nucleic acid replication. Particularly, it comprises one or more materials selected from: DNA polymerases, RNA polymerases, DNA terminal proteins, Single-Stranded DNA Binding Proteins (SSBPs), Double- Stranded DNA Binding Proteins (DSBPs), nucleic acid sequences coding for DNA polymerases, nucleic acid sequences coding for RNA polymerases, nucleic acid sequences coding for DNA terminal proteins, nucleic acid sequences coding for SSBPs, nucleic acid sequences coding for DSBPs, the necessary cofactors and substrates (such as e.g.
  • nucleoside triphosphates ammonium sulfate
  • accessory nucleic acids such as primers
  • Other proteins enzymatic activity may be included as well.
  • proteins enzymatic activity may be included as well.
  • those involved in the replication of DNA in living organism or in viruses such as recombinase, topoisomerase, transposases, reverse transcriptase, helicase, nuclease processivity factors, enzyme associated with nucleic acid metabolism, among others, may also be comprised in the in vitro platform of the first composition.
  • nucleic acid transcription machinery thus comprises or consists essentially of or consists of all the required materials for completing nucleic acid transcription starting from a DNA and ending by an mRNA. Particularly, it comprises one or more materials selected from: RNA polymerases, nucleic acid sequences coding for RNA polymerases, and any combination thereof (and substrates and cofactors and accessory enzymes and proteins such as transcription factors (TF), etc).
  • nucleic acid translation machinery thus comprises or consists essentially of or consists of all the required materials for completing nucleic acid translation starting from an mRNA and ending by a polypeptide.
  • ribosomes comprises one or more ribosomes, translation factors (e.g., initiation factors, elongation factors, release factors, NTP-recycling enzymes), aminoacyl tRNA synthetases, tRNAs, and the like.
  • translation factors e.g., initiation factors, elongation factors, release factors, NTP-recycling enzymes
  • aminoacyl tRNA synthetases e.g., initiation factors, elongation factors, release factors, NTP-recycling enzymes
  • tRNAs e.g., initiation factors, elongation factors, release factors, NTP-recycling enzymes
  • aminoacyl tRNA synthetases e.g., tRNAs, and the like.
  • Any one of the above defined machineries can either be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages (such as Phi29, T7, QB, MS2, or T4), bacteria (such as Escherichia coli), yeasts (such as Saccharomyces cerevisiae), viruses (such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus), other eucaryotic cells (e.g. rabbit retyculocytes) or any combination thereof.
  • bacteriophages such as Phi29, T7, QB, MS2, or T4
  • bacteria such as Escherichia coli
  • yeasts such as Saccharomyces cerevisiae
  • viruses such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus
  • a “cellular extract” refers to an extract of cellular components and may be used in crude (i.e. without any prior purification) or more or less purified (at least one purification step enriching the extract in the desired compounds) form.
  • a “reconstituted system” refers to a reconstituted mixture of previously purified or recombinantly expressed components (e.g. the PURE system as a reconstituted transcription and translation system, or the mixture of Phi29 bacteriophage proteins p2, p3, p5 and p6 as a reconstituted replication system, both of which are described in more details below).
  • parameter refers to a feature, in particular a physical, quantifiable feature (such as, concentration, temperature, pH, viscosity, and the like) whose value is used to appropriately represent, define or classify an entity or a situation, in particular at a given time point.
  • substrate herein refers to a starting material that is subject to a chemical or biological reaction, in particular an enzymatic reaction in the presence of an enzyme.
  • the substrate can advantageously be coupled and/or fused to a compound or moiety having detectable optical properties.
  • a reaction e.g., an enzymatic reaction or a chemical reaction
  • a product is the material resulting from a reaction starting with a substrate.
  • optical technique or “optical method” or “optical property” refer to a technique or a method or a property that allows optical detection and/or measure.
  • An optical technique can be selected from absorption -based detection techniques (including single or multi-wavelength measurements, etc.), fluorescence-based detection techniques (including fluorescence intensity, fluorescence lifetime, fluorescence wavelength, FRET, etc.), and any combination thereof.
  • An optical property can be inter alia an electromagnetic and/or color and/or fluorescent property happening in visible or non-visible wavelengths.
  • a “fluorescent compound” refers to a compound that emits light when exposed to light or other radiation, the emitted light having a longer wavelength than the light to which the compound has been exposed. This light emission can be advantageously detected and/or measured by an optical technique as defined above.
  • the term “compound” herein refers to a chemical or biological entity, such as a molecule.
  • Illustrative examples of compounds are, without limitation, reagents, substrates, cosubstrates, cofactors, coenzymes, DNAs, RNAs, hormones, antigens, epitopes, ligands, antibodies, receptors, toxins, and the like.
  • At least one microcompartment comprises a single copy of a nucleic acid molecule herein means that the average number of nucleic acid molecules per microcompartment (corresponding to the ratio between the sum of the individual nucleic acid molecule contents of the microcompartments and the total number of microcompartments) is between 0 and 10, preferably between 0 and 5, yet preferably between 0 and 2, and most preferably between 0.01 and 1.
  • the average number of nucleic acid molecules per microcompartment is 1 following a random distribution of the DNA molecules in microcompartments of identical inner volume, it means that it is actually obtained about 37% of empty microcompartments, about 37% of microcompartments containing one nucleic acid molecule, and about 26% of microcompartments containing 2 or more nucleic acid molecules.
  • A is easily computed knowing the DNA concentration in the encapsulated mixture and the volume of individual compartments.
  • Poissonian partitioning Such partitioning of DNA molecules in microcompartments, that allows many microcompartments to receive only a single molecule of DNA (and where necessarily many compartments do not receive any molecule of DNA) is referred to as “Poissonian partitioning”.
  • This Poissonian partitioning is very important in all microcompartmentalized uHTS methods because it allows to maintain clonality and genotype-phenotype linkage at a statistical level. To do so, the concentration of objects to be encapsulated (DNA molecules here but may be bacteria, or beads, in other approaches) so adjusted so that only a minor fraction of the microcompartments receive more than one object.
  • the result is a collection of microcompartments that, for the most part, contain either 0 or 1 object. This approach thus allows to isolate most objects in individual microcompartments, without resorting to a deterministic partitioning.
  • the expression “at least one nucleic acid molecule comprising at least one candidate sequence encoding at least one polypeptide” means that:
  • the nucleic acid molecule can comprise more than one candidate sequence; and/or Each candidate sequence can encode more than one polypeptide; and/or
  • Each candidate sequence can be present as a single copy in the nucleic acid molecule.
  • nucleic acid molecule such as a nucleic acid in a standard linear form (e.g., PCR products)
  • a single encapsulation step to provide clonal microcompartment libraries containing numerous copies of a given nucleic acid molecule containing one or more gene and a high concentration of the encoded protein(s), which can be directly submitted to the sorting stage.
  • the present invention thus provides an original, reliable and efficient method to perform in vitro uHTS via a one-step encapsulation of nucleic acid molecules (such as linear DNA constructs) containing the genetic variant (or candidate sequence) to be tested in a molecular mixture that allows for simultaneous specific gene amplification and protein expression.
  • nucleic acid molecules such as linear DNA constructs
  • the present Inventors have unexpectedly demonstrated that it is possible to obtain levels of protein activity that are sufficient for uHTS screening from the encapsulation according to a Poisson distribution of a single copy of a nucleic acid molecule, in linear form, in microcompartments containing a platform combining two activities, isothermal genetic replication and protein expression.
  • Poissonian partitioning enables clonality, i.e. the fact that most microcompartments contains a unique variant of the library.
  • the platform thus developed by the Inventors combines an IVTT system with a system capable of performing replication of this unique variant in each microcompartment, to obtain an IVTTR system.
  • This innovative platform relies on the association of (i) a nucleic acid replication machinery, (ii) optionally a nucleic acid transcription machinery (depending on the type of the starting nucleic acid molecule used in the screening method), and (iii) a nucleic acid translation machinery.
  • the data obtained by the Inventors demonstrate that the association of both the replication and the translation machineries in a one-pot reaction results in the concomitant nucleic acid amplification and protein production of any gene, in an amount that is sufficient for automatized detection of the protein activity of interest, even when starting initially from a single copy of the gene-containing nucleic acid molecule within the microcompartment.
  • the Inventors have thus successfully applied this novel method to the mutational screening or bioprospection of any protein of interest.
  • this innovative approach relies on the direct encapsulation of a mixture, and does not involve procedures that are complex or multistep or limited in throughput (such as microfluidic droplet manipulation, bacterial transformation, preamplification or microbead manipulations), the data confirm it can be successfully used to manipulate large libraries in order to enrich them in genes encoding proteins with targeted properties.
  • this novel method is compatible with the enrichment of any protein library, including enzyme libraries and libraires of protein with other activities, such as fluorescent proteins or transcription factors (TF). This is in contrast with uHTS approaches that are only compatible with the selection of binders.
  • a protein having an activity of interest such as an enzyme, a protein of a protein complex, a nucleic acid binding protein, a transcription factor, and any combination thereof
  • a very high rate e.g., millions of variants per hour
  • nucleic acid molecule even when initially present as of a single copy in a microcompartment, undergoes cycles of isothermal exponential amplification driven by the replication machinery, and, simultaneously, the accumulating gene copies serve as a template for protein translation, leading to the synthesis of high and easily detectable concentrations of the gene products;
  • the approach is advantageous for screening applications because it exponentially replicates the nucleic acid molecule in the microcompartment, going from at least one of a single copy at time of encapsulation up to nanomolar concentrations in the microcompartment when entering the sorting stage. It then becomes easier to recover the genetic material at the end of the screening round, when all sorted positive micro-compartments are pooled and their content analyzed or used as starting library for the next directed evolution cycle. This amplification increases the amount of nucleic acid molecule that is recovered after screening, even if just a few compartments are sorted to the positive bin, and simplifies the further processing of the selected library (i.e. , recovery of sorted genes);
  • any type of nucleic acid molecule can be used as starting material, including DNA (such as cDNA, single stranded DNA, double stranded DNA, partially double-stranded DNA, linear DNA, circular DNA, capped DNA, uncapped DNA, etc.) and RNA (such as mRNA, capped RNA, uncapped RNA, DNA/RNA hybrid, vector (such as a plasmid or a viral vector).
  • DNA such as cDNA, single stranded DNA, double stranded DNA, partially double-stranded DNA, linear DNA, circular DNA, capped DNA, uncapped DNA, etc.
  • RNA such as mRNA, capped RNA, uncapped RNA, DNA/RNA hybrid
  • vector such as a plasmid or a viral vector.
  • in vitro replication can be achieved using terminal P3 capped DNA containing specific origins of replication.
  • In vitro replication can alternatively be initiated by 5’phosphate capped DNA, which is easy to obtain (using 5’P primer during the PCR, or
  • the genetic amplification process may be compatible with long DNA or RNA constructs (up to multi kilobases constructs) so that it is possible to include additional functionalities in the replicator, which can be used to ameliorate the protein screening process;
  • the present invention thus relates to a novel method for screening a library of nucleic acid molecules encoding each a candidate polypeptide for one or more nucleic acid molecule(s) encoding a polypeptide with an activity of interest, comprising, or consisting essentially of, or consisting of: a) Providing a first composition comprising an in vitro platform comprising:
  • step b) Providing a second composition comprising a library of nucleic acid molecules each comprising at least one candidate sequence encoding at least one candidate polypeptide; c) Mixing said first composition of step a) with said second composition of step b); d) Partitioning the thus obtained mixture into microcompartments, wherein at least one of said microcompartments comprises a single copy of said nucleic acid molecule; e) Allowing said candidate sequence to be amplified by the nucleic acid replication machinery in the microcompartments; f) Optionally allowing the amplified candidate sequence of step e) to be transcribed by the nucleic acid transcription machinery in the microcompartments; g) Allowing the amplified candidate sequence of step e) or the transcribed candidate sequence of step f) to be translated in at least one candidate polypeptide by the nucleic acid translation machinery in the microcompartments; h) Optionally adding at least one compound in the mixture of step c) or
  • the method of the invention is particularly adapted for high-throughput screenings (HTS), preferably ultra-high-throughput screenings (uHTS).
  • HTS high-throughput screenings
  • uHTS ultra-high-throughput screenings
  • the method of the invention allows high or ultrahigh throughput screening of at least 10 5 nucleic acid molecules comprising at least one candidate sequence, preferably at least 10 6 nucleic acid molecules comprising at least one candidate sequence.
  • the polymer may notably be a plastic polymer, in particular selected from PLA (polylactic acid), PET (polyethylene terephthalate), PE (polyethylene), PP (polypropylene), PCL (polycaprolactone), and PS (polystyrene).
  • PLA polylactic acid
  • PET polyethylene terephthalate
  • PE polyethylene
  • PP polypropylene
  • PCL polycaprolactone
  • PS polystyrene
  • the polypeptide with an activity of interest is selected from the group consisting of an enzyme, a polypeptide of a protein complex, a nucleic acid binding protein, a transcription factor, and any combination thereof (i.e., the method is preferably for screening a library of nucleic acid molecules encoding each a candidate polypeptide selected from the group consisting of a candidate enzyme, a candidate polypeptide of a protein complex, a candidate nucleic acid binding protein, a transcription factor, and any combination thereof).
  • a first composition comprising an in vitro platform comprising:
  • nucleic acid translation machinery (iii) a nucleic acid translation machinery is provided.
  • the nucleic acid replication machinery comprises or consists essentially of or consists of all the required materials for completing nucleic acid replication. Particularly, it may comprise or consist essentially of or consist of one or more materials selected from: DN polymerases, RN polymerases, DN terminal proteins (TP), Single- Stranded DNA Binding Proteins (SSBPs), Double-Stranded DNA Binding Proteins (DSBPs), nucleic acid sequences coding for DNA polymerases, nucleic acid sequences coding for RNA polymerases, nucleic acid sequences coding for DNA terminal proteins, nucleic acid sequences coding for SSBPs, nucleic acid sequences coding for DSBPs, the necessary cofactors and substrates (such as e.g. nucleosides triphosphates), accessory nucleic acids (such as primers) and any combination thereof.
  • DN polymerases RN polymerases, DN terminal proteins (TP), Single- Stranded DNA Binding Proteins (
  • the combination of materials necessary for completing nucleic acid replication varies depending on the specific type of replication machinery that is selected.
  • the nucleic acid replication machinery preferably comprises (or consists essentially of or consists of) a DNA polymerase, a DNA terminal protein (TP), a Single-Stranded DNA Binding Protein (SSBP), and a Double-Stranded DNA Binding Proteins (DSBP). Since the replication reaction happens in an IVTT mixture, one or more of these enzymes can be replaced by its encoding DNA with appropriate regulatory sequence (promotor, Ribosome binding site and terminator). In that case the enzyme(s) required for replication is produced in situ by in vitro transcription and translation before the replication reaction happens. Even more preferably, when the minimal replication machinery of phage Phi29 is used, it comprises (or consists essentially of or consists of):
  • TP DNA terminal protein
  • SSBP Single-Stranded DNA Binding Protein
  • DSBP Double-Stranded DNA Binding Proteins
  • TP DNA terminal protein
  • a general minimal replication machinery will preferably comprise (or consist essentially or consist of) a DNA or RNA polymerase, a Single-Stranded DNA Binding Protein (SSBP), and a Double-Stranded DNA Binding Proteins (DSBP). This minimal replication machinery may then be completed by other materials, depending on the specific type of replication machinery that is selected.
  • SSBP Single-Stranded DNA Binding Protein
  • DSBP Double-Stranded DNA Binding Proteins
  • the proteins included in the replication machinery may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages (such as Phi29, T7, QB, MS2, or T4), bacteria (such as Escherichia coli), yeasts (such as Saccharomyces cerevisiae), viruses (such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus), other eucaryotic cells (e.g. rabbit reticulocytes) or any combination thereof.
  • bacteriophages such as Phi29, T7, QB, MS2, or T4
  • bacteria such as Escherichia coli
  • yeasts such as Saccharomyces cerevisiae
  • viruses such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus
  • other eucaryotic cells e
  • the proteins included in the replication machinery are selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages, and more preferably from bacteriophage Phi29 (it is then referred to as a “Phi29-based replication machinery”).
  • the replication machinery preferably comprises (or consist essentially or consist of) a DNA polymerase, a DNA terminal protein (TP), a Single-Stranded DNA Binding Protein (SSBP), and a Double-Stranded DNA Binding Proteins (DSBP), wherein the DNA polymerase corresponds to bacteriophage Phi29 protein p2, the DNA terminal protein (TP) to bacteriophage Phi29 protein p3, the Single-Stranded DNA Binding Protein (SSBP) to bacteriophage Phi29 protein p5 and the Double-Stranded DNA Binding Proteins (DSBP) to bacteriophage Phi29 protein p6.
  • the replication machinery comprises (or consists essentially of or consists of):
  • Phi29 proteins p2, p3, p5 and p6 preferably purified recombinant Phi29 proteins p2, p3, p5 and p6 of amino acid sequences SEQ ID NO:45 to 48; or
  • nucleic acid sequences coding for Phi29 proteins p2 and p3 preferably pUC57_OriLR_p2p3 of sequence SEQ ID NO:34
  • purified Phi29 proteins p5 and p6 preferably purified recombinant Phi29 proteins p5 and p6 of amino acid sequences SEQ ID NO:47 and 48.
  • replication machineries may be used.
  • a reconstituted plasmidic replication machinery (Ueno H et al. “Amplification of over 100 kbp DNA from Single Template Molecules in Femtoliter Droplets _ACS Synth Biol. 2021 Sep 17; 10(9):2179- 2186)
  • a reconstituted chromosomal replication machinery (Su’etsugu, Masayuki, et al. "Exponential propagation of large circular DNA by reconstitution of a chromosomereplication cycle.” Nucleic acids research 45.20 (2017): 11525-11534) or a reconstituted viral replication machinery (Kulczyk AW et al. “The Replication System of Bacteriophage T7” Enzymes. 2016;39:89-136) may be used.
  • the above-described proteins of the replication machinery may be provided either as a cellular extract or as a reconstituted system.
  • a reconstituted system comprising the protein mixtures defined above is used.
  • proteins with enzymatic activity may be included as well.
  • those involved in the replication of DNA in living organisms or in viruses such as recombinase, topoisomerase, transposases, reverse transcriptase, helicase, nuclease processivity factors, enzyme associated with nucleic acid metabolism, among others, may also be comprised in the nucleic acid replication machinery.
  • Those also may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages (such as Phi29, T7, QB, MS2, or T4), bacteria (such as Escherichia coli), yeasts (such as Saccharomyces cerevisiae), viruses (such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus), other eucaryotic cells (e.g. rabbit reticulocytes) or any combination thereof.
  • bacteriophages such as Phi29, T7, QB, MS2, or T4
  • bacteria such as Escherichia coli
  • yeasts such as Saccharomyces cerevisiae
  • viruses such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus
  • other eucaryotic cells e.g. rabbit
  • the replication machinery should also contain deoxyribonucleotides triphosphate (dNTP).
  • dNTP deoxyribonucleotides triphosphate
  • the in vitro platform in the first composition optionally comprises a nucleic acid transcription machinery, depending on the type of nucleic acids present in the second composition provided in step b).
  • the second composition of step b) comprises mRNA
  • no nucleic acid transcription machinery is needed.
  • the in vitro platform in the first composition also comprises a nucleic acid transcription machinery.
  • the nucleic acid transcription machinery comprises or consists essentially of or consists of all the required materials for completing nucleic acid transcription starting from a DNA and ending by an mRNA. Particularly, it comprises one or more materials selected from: RNA polymerases, nucleic acid sequences coding for RNA polymerases, and any combination thereof (and substrates and cofactors and accessory enzymes and proteins such as transcription factors (TF), etc).
  • the nucleic acid transcription machinery comprises an RNA polymerase, for example T7 RNA polymerase.
  • a nucleic acid translation machinery is also comprised in the in vitro platform of the first composition.
  • This nucleic acid translation machinery comprises or consists essentially of or consists of all the required materials for completing nucleic acid translation starting from an mRNA and ending by a polypeptide. Particularly, it comprises one or more ribosomes, translation factors (e.g., initiation factors, elongation factors, release factors, NTP-recycling enzymes), aminoacyl tRNA synthetases, tRNAs, and the like.
  • the proteins included in or encoded by a nucleic acid molecule included in the nucleic acid translation machinery and included in or encoded by a nucleic acid molecule included in the optional nucleic acid transcription machinery may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages (such as Phi29, T7, QB, MS2, or T4), bacteria (such as Escherichia coli), yeasts (such as Saccharomyces cerevisiae), viruses (such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus), other eucaryotic cells (e.g.
  • rabbit reticulocytes or any combination thereof. They may more particularly may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteria, in particular Escherichia coli.
  • modifications e.g., genetic modifications such as genetic mutations
  • engineered from natural or wild-type machineries from bacteria, in particular Escherichia coli.
  • the in vitro platform comprised in the first composition is preferably a cell-free system, which may be provided as a cellular extract, a reconstituted system and any combination thereof.
  • a reconstituted system is used.
  • the nucleic acid translation machinery and the nucleic acid transcription machinery may preferably be provided in a single reconstituted system, preferably a purified gene expression machinery from E. coli, such as the PURE system (for “Protein synthesis Using Recombinant Elements”), which is a commercially available from various suppliers (for example the “PUREfrex” system from GeneFrontier) reconstituted nucleic acid transcription and translation machinery able to isothermally express any Open Reading Frame under T7 promoter and E.
  • a purified gene expression machinery from E. coli
  • the PURE system for “Protein synthesis Using Recombinant Elements”
  • the PUREfrex for example the “PUREfrex” system from GeneFrontier
  • coli Ribosome Binding Site comprising 32 individually purified components: initiation factors IF1 , IF2, and IF3; elongation factors EF-G, EF-Tu, and EF-Ts; release factors RF1 and RF3; ribosomerecycling factor (RRF), 20 aminoacyl-tRNA synthetases (ARSs), methionyl-tRNA transformylase (MTF), T7 RNA polymerase, and ribosomes.
  • RRF ribosomerecycling factor
  • ARSs aminoacyl-tRNA synthetases
  • MTF methionyl-tRNA transformylase
  • T7 RNA polymerase T7 RNA polymerase
  • the PURE system contains 46 tRNAs, nucleosides triphosphate (NTPs), creatine phosphate, 10-formyl-5, 6,7,8- tetrahydrofolic acid, 20 amino acids, creatine kinase, myokinase, nucleoside-diphosphate kinase, and pyrophosphatase (Shimizu, Y. et al. Cell-free translation reconstituted with purified components. Nat. Biotechnol. 19, 751-755 (2001 )).
  • PUREfrex 2.0 is an upgraded version of PURE 1.0 for protein production
  • PUREfrex 2.1 is a version of PUREfrex 2.0 lacking any reducing agent, allowing to tune oxidoreductive environment of the solution.
  • the first composition provided in step a) may further comprise other components useful in the method of the invention.
  • a substrate, a cofactor, a coenzyme (prosthetic group or cosubstrate) of the enzyme or any combination thereof may be added in the first composition provided in step a).
  • step b) a second composition comprising a library of nucleic acid molecules each comprising at least one candidate sequence encoding at least one polypeptide is provided.
  • the nucleic acid molecules are preferably selected from the group consisting of DNAs (such as gDNA, cDNA), RNAs (such as mRNA), and any combination thereof (such as DNA/RNA hybrids).
  • the nucleic acid molecule may be single-stranded, double-stranded, or partially double-stranded.
  • the nucleic acid molecule may be linear or circular. It can originate from in vitro preparation (e.g., chemical synthesis, PCR, isothermal amplification%) or in vivo preparation (e.g., plasmid, viral vectors). It can contain specific subsequences (in addition to the candidate sequence encoding a polypeptide mentioned) and specific caps, such as terminal proteins.
  • the nucleic acid molecule may additionally contain (non-canonical) modifications, such as non-canonical bases, epigenetic marks, internal modifications or terminal modifications.
  • Each nucleic acid molecule of the library is preferably a replicator, i.e. a nucleic acid molecule able to replicate under specific conditions in the presence of specific compounds (the “nucleic acid replication machinery”).
  • the replicator acts as a template for its amplification, so that the newly-made nucleic acid molecules are copies or reverse- complemented copies of it.
  • the concentration of the replicator will increase. For example, if one starts from a single molecule of the replicator, after some time, 2, then 4, 8 etc molecules may be present, i.e. its concentration will increase exponentially.
  • a replicator can be any type of nucleic acid (DNA, RNA or hybrids) and may contain specific subsequences or non-canonical modifications (see above).
  • a replicator preferably includes one or more subsequence(s) selected from an origin of replication (OR), a primer binding site, one or more genes with appropriate regulatory sequences (e.g. a promoter, terminator and ribosome-binding site (RBS)). More preferably, a replicator comprises an origin of replication (OR), a primer binding site, one or more genes with appropriate regulatory sequences (e.g. a promoter), and a ribosome-binding site (RBS).
  • the replicator is preferably a linear double-stranded DNA comprising at least a Phi29 origin of replication at each extremity, or a derivative thereof (e.g. minimal origin), and one or more sequences encoding a polypeptide with appropriate regulatory sequences (e.g. a promoter, terminator and ribosome-binding site (RBS)), wherein each strand of the replicator is further capped in 5’ by phage Phi29 P3 protein or a phosphate group.
  • Phosphorylated 5’ ends i.e. capping in 5’ with a phosphate group
  • the replicator comprises only one sequence encoding a candidate polypeptide.
  • the nucleic acid molecule acting as replicator may contain multiple sequences encoding a polypeptide.
  • they started from a nucleic acid molecule comprising both a reporter gene (such as a gene encoding a fluorescent or a colored gene), inserted as a fusion with the gene encoding the polypeptide of interest.
  • the colored or fluorescence signal associated with the reporter gene indicates the level of expression of the polypeptide of interest, which can be used to normalize the activity with respect to expression level during the sorting stage.
  • the sorting process selects nucleic acid sequences encoding polypeptides with high specific activity rather than polypeptides with high apparent activity (where apparent activity is the product of specific activity and polypeptide concentration).
  • the replicator comprises several sequences each encoding a distinct polypeptide.
  • at least one of the sequences encodes a candidate polypeptide to be screened, while other sequence(s) may encode other candidate polypeptide(s) to be screened, a reporter protein (for instance a coloured or fluorescent protein) that may be used to normalize the expression level of the polypeptide(s) to be screened, or a substrate protein or a functional variant or fragment thereof (when the candidate polypeptide is an enzyme with a protein substrate).
  • a reporter protein for instance a coloured or fluorescent protein
  • the distinct sequences may be present as one or more fusions, as independent transcriptional units, or as multicistronic (e.g. bicistronic) constructs.
  • a fusion is particularly adapted for normalization as it ensures that the expression level of the fused polypeptides is exactly the same.
  • the candidate polypeptide encoded by the candidate sequence is selected from the group consisting of putative enzymes, putative transcription factors, putative functional derivatives thereof, and putative functional fragments thereof.
  • the candidate sequence is preselected based on at least one of: sequence thereof, a portion of sequence thereof, structural properties, functional properties, and any combination thereof.
  • nucleic sequences encoding a candidate (or other additional) polypeptide are preferably under the control of a T7 promoter.
  • a replicator may further comprise additional subsequences, notably selected from a ribosome-binding site (RBS), barcodes, restriction enzyme binding sites, aptamers.
  • RBS ribosome-binding site
  • barcodes barcodes
  • restriction enzyme binding sites aptamers
  • the nucleic sequences encoding a candidate (or other additional) polypeptide preferably comprise an E. coli Ribosome Binding Site.
  • a particularly preferred replicator has 5’ -phosphate capped ends and comprises on one strand the following nucleic acid sequences from 5’ to 3’: oriL (preferably Phi29 oriL), T7 promoter, RBS (ribosome binding site, preferably E. coli RBS), gene encoding the candidate polypeptide, terminator, and oriR (preferably Phi29 oriR).
  • the Inventors have applied this novel method in combination with DNA-barcoding techniques by including randomized DNA subsequences called barcodes or Unique Molecular Identifiers (UMI) in the replicator.
  • barcodes or UMI can be used in combination with sequencing, for example to compute the changes in frequency of a given variant upon cycles of screening, to correct for amplification biases, to compute high quality consensuses from multiple reads, to assemble short reads into a longer complete sequence, or to check for the absence of recombination during the amplification steps or library preparation, or to “call out” one particular variant (for instance by performing a PCR with the barcode sequence as primer(s)).
  • the method is a high-throughput screening (HTS) method, more preferably an ultra-high-throughput screening (uHTS) method.
  • HTS high-throughput screening
  • uHTS ultra-high-throughput screening
  • the method allows high or ultrahigh throughput screening of at least 10 5 nucleic acid molecules comprising at least one candidate sequence (preferably replicators as defined herein), preferably at least 10 6 nucleic acid molecules comprising at least one candidate sequence (preferably replicators as defined herein).
  • the second composition provided in step b) may further comprise other components useful in the method of the invention.
  • a substrate, a cofactor, a coenzyme (prosthetic group or cosubstrate) of the enzyme or any combination thereof may be added in the second composition provided in step b).
  • step c) the first composition provided in step a) and the second composition provided in step b) are mixed. To avoid premature initiation of the replication or expression reactions, this step is preferably performed at cold temperature and/or just before step d).
  • Step d) Partitioning the thus obtained mixture into microcompartments
  • step d) the mixture obtained in step c) is partitioned into microcompartments, wherein at least one of said microcompartments comprises a single copy of said nucleic acid molecule.
  • a substantial fraction of said microcompartments comprises a single copy of said nucleic acid molecule. More preferably, essentially all non-empty microcompartments comprises a single copy of said nucleic acid molecule.
  • microcompartments are selected from the group consisting of microdroplets (such as water-in-oil microdroplets), microvesicles (such as liposomes), microchambers, and any combination thereof.
  • microdroplets such as water-in-oil microdroplets
  • microvesicles such as liposomes
  • microchambers and any combination thereof.
  • the desired type(s) of microcompartments may be generated using standard techniques well-known to those skilled in the art.
  • microdroplets may be obtained by microfluidic techniques or by bulk emulsification techniques.
  • monodisperse water-in-oil microdroplets may be obtained using microfluidic system (Christopher, G. F. & Anna, S. L. Microfluidic methods for generating continuous droplet streams. Journal Of Physics D-Applied Physics 40, R319-R336 (2007)).
  • the average number of nucleic acid molecules per microcompartment (corresponding to the ratio between the sum of the individual nucleic acid molecule contents of the microcompartments and the total number of microcompartments) after partitioning is between 0 and 10, preferably between 0 and 5, yet preferably between 0 and 2, and most preferably between 0.1 and 1.
  • many microcompartments receive no nucleic acid molecule (they are referred to as empty microcompartments), and most non-empty microcompartments comprise a single nucleic acid molecule.
  • the mean number of nucleic acid molecules (preferably replicators) per droplet would be 0.2.
  • a Poissonian partitioning is expected, with 82% of droplets being empty, 16% containing one nucleic acid molecule (preferably a replicator), and approximatively 2% containing two or more nucleic acid molecules (preferably replicators).
  • step d) the mixture obtained in step c) is partitioned into microcompartments, using Poissonian partitioning (i.e., poissonian distribution), preferably so that a substantial fraction of said microcompartments comprises a single copy of said nucleic acid molecule, more preferably wherein most of said microcompartments comprises a single copy of said nucleic acid molecule, even more preferably, essentially all non-empty microcompartments comprises a single copy of said nucleic acid molecule.
  • Poissonian partitioning enables clonality, i.e., ensures that only one copy of nucleic acid molecule (i.e., a unique variant of the library) is incorporated in the majority of micro-compartments.
  • the partitioning of nucleic acid molecules may deviate from a Poissonian distribution (e.g., power-law distribution).
  • experiments can be designed to obtain an empirical validation confirming that few microcompartments contain more than one nucleic acid variant molecule at the time of generation.
  • step e) optional step f) and step g), the microcompartments are incubated in conditions suitable for the replication machinery, the optional transcription machinery and the translation machinery to perform their function, leading to amplification of both the candidate nucleic sequence encoding the candidate polypeptide and of the candidate polypeptide itself.
  • Suitable conditions depend on the specific selected replication machinery, optional transcription machinery and translation machinery, and are known in the art.
  • conditions used in step e) are preferably between 20 °C and 50 °C, more preferably temperatures ranging from 25 to 37 °C during 1 hour to 16 hours in the presence of 20 mM ammonium sulfate.
  • reaction temperature can be below 40 °C, more preferably temperatures ranging from 25 to 37 °C during 1 hour to 16 hours.
  • Step h) Optionally adding at least one compound in the mixture of step c) or in the microcompartments of step g), and/or modifying at least one parameter within the microcompartments of step d) or g)
  • step h) at least one compound may be added in the mixture of step c) or in the microcompartments of step g), and/or at least one parameter may be modified within the microcompartments of step d) or g).
  • the candidate polypeptide is a candidate enzyme
  • at least one compound that may be added in the mixture of step c) or in the microcompartments of step g) may be selected from reagents, substrates, cofactors, coenzymes (including prosthetic groups and cosubstrates) of the candidate polypeptide encoded by the candidate sequence and any combination thereof.
  • the activity of interest is polymer enzymatic degradation
  • the at least one compound added the mixture of step c) or in the microcompartments of step g) is a fluorogenic (e.g. Fluorescein Diacetate substrate (FDA), Resorufin Di-B-D- Galactopyranoside (FDG), etc), chromogenic (e.g. 3, 3’-diaminobenzidine tetrahydrochlorure), solvatochromic (Nile Red) or environment-sensitive (e.g.
  • Bromocresol Green compound embedded in one or more polymer particles, wherein degradation of the polymer particles catalyzed by the candidate polypeptide leads to release and conversion of the compound into a fluorescent or colored product, or a change in spectral properties in the microcompartment. Sorting may then be done based on the amount of fluorescent or colored product detected in step i), depending whether it is higher or lower a user-defined threshold value.
  • the polymer may notably be a plastic polymer, in particular selected from PLA (polylactic acid), PET (polyethylene terephthalate), PE (polyethylene), PP (polypropylene), PCL (polycaprolactone), and PS (polystyrene).
  • PLA polylactic acid
  • PET polyethylene terephthalate
  • PE polyethylene
  • PP polypropylene
  • PCL polycaprolactone
  • PS polystyrene
  • At least one compound that may be added in the mixture of step c) or in the microcompartments of step g) may be a nucleic acid molecule comprising a nucleic acid sequence encoding a reporter polypeptide (preferably a fluorescent or colored polypeptide) under the control of said transcription factor.
  • At least one parameter may be modified within the microcompartments of step d), in order to put the microcompartments in conditions optimal for replication in step e), optional transcription in step f) and translation in step g).
  • at least one parameter may be modified within the microcompartments of step g), in order to put the microcompartments in conditions optimal for detecting the candidate polypeptide activity in next step i).
  • At least one parameter that may be modified within the microcompartments of step d) or g) may be selected from concentration, temperature, pH, viscosity, and any combination thereof.
  • Step i) - Detecting the activity of the candidate polypeptide of step g) in the conditions that are optionally modified in step h) within the microcompartments
  • step i) the activity of the candidate polypeptide of step g) is detected in the microcompartments.
  • the activity of the candidate polypeptide encoded by the candidate sequence is determined in step i) by measuring the intensity of one or more signals.
  • This measure can be made using one or more appropriate techniques selected from optical techniques, electrical techniques (such as capacitance, resistance, impedance measurements, and the like), and magnetic techniques.
  • the detection may be a direct detection assay (preferably a fluorogenic assay or a colorimetric assay), i.e. the measured activity directly generates a detectable optical and/or electrical and/or magnetic changes in the microcompartments.
  • a direct detection assay preferably a fluorogenic assay or a colorimetric assay
  • the measured activity directly generates a detectable optical and/or electrical and/or magnetic changes in the microcompartments.
  • a direct detection assay preferably a fluorogenic assay or a colorimetric assay
  • the measured activity directly generates a detectable optical and/or electrical and/or magnetic changes in the microcompartments.
  • a substrate comprising a part with optical may be used, and the reaction of the candidate enzyme encoded by the candidate sequence on the substrate will generate detectable optical changes in the microcompartments if the candidate enzyme has enzymatic activity.
  • a substrate coupled/fused to a compound with detectable optical properties may be used.
  • the activity of the candidate enzyme encoded by the candidate sequence may be determined by detecting a product obtainable from the substrate such as a product that either is converted by the candidate enzyme from the substrate, or is released from the reaction of the candidate enzyme with the substrate.
  • a direct detection assay may also be used when the candidate polypeptide is a candidate reporter protein, such as a fluorescent or colored protein.
  • an indirect activity assay may be used, preferably a fluorogenic assay or a colorimetric assay, wherein the activity of the polypeptide encoded by the candidate sequence is determined by detecting a product of the indirect activity assay.
  • a nucleic acid molecule comprising a nucleic acid sequence encoding a reporter polypeptide with optical and/or electrical and/or magnetic properties (preferably a fluorescent or colored polypeptide) under the control of said transcription factor may be present in the microcompartments, either due to their presence in the first or second composition provided in step a) or b), or by addition in step h).
  • detection of the reporter polypeptide expressed only when the candidate transcription factor is active indirectly permits to detect the candidate transcription factor activity.
  • a product released by the activity of the candidate enzyme may be used as a substrate by a second enzyme, which has been included in the reaction mixture.
  • reaction of this second enzyme with the product of the candidate enzyme optionally in presence of a specific cosubstrate, generate a secondary product which may be easily detected for example via spectroscopic method.
  • Many indirect enzyme detection assays for example with fluorescent readouts, have been described and are well-known.
  • a product released by the activity of the candidate enzyme may bind to one (or more) macromolecules and increase its (their) fluorescence by irreversible photoactivation/photoconversion or reversible photoswitching, such as, but not restricted to, fluorescent biosensors proteins or aptamers (Sam Duwe, Peter Dedecker, Optimizing the fluorescent protein toolbox and its use, Current Opinion in Biotechnology, Volume 58, 2019, Pages 183-191 , ISSN 0958-1669; Koveal, D., Rosen, P.C., Meyer, D.J. et al. A high-throughput multiparameter screen for accelerated development and optimization of soluble genetically encoded fluorescent biosensors. Nat Commun 13, 2919 (2022)).
  • an absorption-based detection method including single or multi-wavelength measurements, turbidometry, etc.
  • a fluorescence- based detection method including fluorescence intensity, fluorescence lifetime, fluorescence wavelength, fluorescence quenching/unquenching, BRET, FRET, relaxation etc.
  • the optical method of step i) is preferably selected from the group consisting of absorption-based detection techniques (including single or multi-wavelength measurements, etc.), fluorescence-based detection techniques (including fluorescence intensity, fluorescence lifetime, fluorescence wavelength, FRET, etc.), and any combination thereof.
  • absorption-based detection techniques including single or multi-wavelength measurements, etc.
  • fluorescence-based detection techniques including fluorescence intensity, fluorescence lifetime, fluorescence wavelength, FRET, etc.
  • candidate polypeptide is a candidate reporter protein, such as a fluorescent or colored protein
  • its activity may be directly detected using fluorescence- based or absorption-based detection techniques (see Example 1 below).
  • candidate polypeptide is a candidate enzyme
  • fluorescence- based or absorptionbased detection techniques may still be used.
  • fluorescence- based detection techniques may be used when a fluorogenic substrate is used (see Example 2 below) or when the candidate enzyme is fused to a fluorescent polypeptide (i.e. the candidate nucleic acid molecule comprises at least one candidate sequence encoding a fusion of a candidate polypeptide and a fluorescent polypeptide, see Examples 3-5 below).
  • absorption-based detection techniques may still be used when a chromogenic substrate is used or when the candidate enzyme is fused to a colored polypeptide (i.e. the candidate nucleic acid molecule comprises at least one candidate sequence encoding a fusion of a candidate polypeptide and a colored polypeptide).
  • fluorescence- based detection techniques may be also used to detect the product of an enzyme, for example using a fluorogenic probe (or a probe coupled to a fluorogenic compound) is used, wherein the probe is specific to the enzyme’s product (for instance see Example 8 below).
  • absorption-based detection techniques may still be used when a chromogenic probe (probe specific to the product of the enzyme) is used or when the probe (specific to the product of the enzyme) is fused to a colored compound.
  • a candidate polypeptide is a candidate regulatory protein such as a transcription factor
  • fluorescence- based or absorption-based detection techniques may still be used.
  • a nucleic acid molecule comprising a nucleic acid sequence encoding a reporter polypeptide (preferably a fluorescent or colored polypeptide) under the control of said transcription factor may be present in the microcompartments, either due to their presence in the first or second composition provided in step a) or b), or by addition in step h).
  • the method of the invention can be a flow-based method (e.g., wherein detection in the microcompartments is sequential) or a non-flow-based method (e.g., wherein detection in the microcompartments is parallel).
  • a flow-based method is preferably performed using a microfluidic device (such as a fluorescence activated droplet sorting (FADS) device (Baret, J.-C. et al. Fluorescence-activated droplet sorting (FADS): efficient microfluidic cell sorting based on enzymatic activity. Lab on a chip 9, 1850-1858 (2009)) or a fluorescence activated cell sorting (FACS) device.
  • FADS fluorescence activated droplet sorting
  • FACS fluorescence activated cell sorting
  • liposomes are immobilized on a flat surface, each on a small electrode, whereby the activity of the protein to be detected changes the electrical property of the droplets, which triggers release of the vesicle in solution.
  • Step j) - Using a sorting apparatus, recovering the candidate sequence encoding a polypeptide having the activity of interest from microcompartments displaying specific activity signal
  • step j the candidate sequences encoding a polypeptide having the activity of interest (as detected in step i)) are recovered from microcompartments displaying specific activity signal, using a sorting apparatus.
  • the detecting device used in step i) for detecting the activity of the candidate polypeptides and the sorting apparatus used in step j) will generally be included in a single apparatus comprising several modules (it will then be referred to a detecting/sorting apparatus).
  • the detecting/sorting apparatus combines a sensor module that detects the activity of interest in each microcompartment, and a sorting module that separates the contents of microcompartments depending on the level of activity measured by the sensor module, based on a user-defined threshold value.
  • the sorting module comprises a software able to compare the level of activity measured by the sensor module with the user-defined threshold value and means for separating the contents of microcompartments depending whether the level of activity measured by the sensor module is above or below the user- defined threshold value.
  • the threshold value is preferably at least equal to and preferably higher than the level of activity measured for a control polypeptide.
  • the threshold value may for instance be at least 1.2, at least 1.3, at least 1.4, at least 1.5, at least 1.6, at least 1.7, at least 1.8, at least 1.9, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 10 3 , at least 10 4 , at least 10 5 or at least 10 6 times higher than the level of activity measured for a control polypeptide.
  • the threshold value is preferably at least equal to and preferably higher than (any value mentioned above) the level of activity measured
  • the threshold value is preferably at least equal to and preferably higher than (any value mentioned above) the level of activity measured for a known polypeptide, more preferably it is higher than the level of activity measured for the polypeptide with the best known level of activity.
  • a fluorescence activated droplet sorting (FADS) device may preferably be used.
  • FACS fluorescence activated cell sorting
  • Such sorting apparatus are commercially available and should be used as recommended by the manufacturer.
  • the sorting apparatus will pool the contents of all microcompartments for which the level of activity measured by the sensor module is higher than the user-defined threshold value.
  • the nucleic acid molecules of the pool may be analyzed or, in the context of directed evolution, used as starting library for the next directed evolution cycle.
  • the initial library (comprised in the second composition that after mixing with the first composition will be partitioned through all microcompartments) is typically obtained from one or a few candidate sequences (for example from natural or engineered sequences encoding proteins with known properties), which are further diversified with mutagenesis (obtained notably by shuffling, random mutations, targeted mutations, combinatorial mutations, or any combination thereof).
  • the invention thus also relates to the use of the method of screening according to the invention in a directed evolution method, said directed evolution method preferably comprising multiple cycles of mutagenesis (obtained notably by shuffling, random mutations, targeted mutations, combinatorial mutations, or any combination thereof) followed by screening using the method of screening according to the invention.
  • the invention also relates to a method of directed evolution, comprising: a) generating a library of nucleic acid molecules encoding each a candidate polypeptide by mutagenesis (obtained notably by shuffling, random mutations, targeted mutations, combinatorial mutations, or any combination thereof) of one or more nucleic acid molecule(s) encoding each a known polypeptide with an activity of interest, b) screening the library generated in step a) for one or more nucleic acid molecule(s) encoding a polypeptide with the activity of interest using the method of screening according to the invention; c) generating a library of nucleic acid molecules encoding each a candidate polypeptide by mutagenesis (obtained notably by shuffling, random mutations, targeted mutations, combinatorial mutations, or any combination thereof) of the one or more nucleic acid molecule(s) selected in step b); d) screening the library generated in step c) for one or more nucleic acid molecule(
  • the directed evolution method may start from one or more nucleic acid molecule(s) encoding each a known enzyme with the activity of interest.
  • known enzymes with the ability to degrade PLA include cutinases (such as Humicola insolens cutinase (HiC)).
  • cutinases such as Humicola insolens cutinase (HiC)
  • Known enzymes with the ability to degrade PET include some proteases and cutinases or IsPETase from Ideonella sakaiensis.
  • Known enzymes with the ability to degrade PE (polyethylene) and PP (polypropylene) include cytochrome P450.
  • PCL polycaprolactone
  • HiC Humicola insolens cutinase
  • proteases Known enzymes with the ability to degrade PS (polystyrene) include hydroquinone peroxidase HM121 from Azotobacter beijerinckii.
  • the initial library is typically a set of natural sequences obtained from an environment where it is likely that genes encoding polypeptides with the desired properties exist. For example, if biomass-degrading enzyme are sought, it may be interesting to use a metagenomic library coming from ruminant gut.
  • the invention thus also relates to the use of the method of screening according to the invention in a functional metagenomics method.
  • the invention also relates to a functional metagenomics method comprising: a) generating a library of natural nucleic acid molecules encoding each a candidate polypeptide obtained from an environment where it is likely that genes encoding polypeptides with the desired activity are present; and b) screening the library generated in step a) for one or more nucleic acid molecule(s) encoding a polypeptide with the activity of interest using the method of screening according to the invention.
  • the method of screening according to the invention may also be used in a method starting by a first step of functional metagenomics and using the selected nucleic acid molecules in a second step of directed evolution.
  • the method of screening according to the invention may be used for the screening of the functional metagenomics step, for the screening of the directed evolution step, or for the screening of both steps.
  • the present invention also relates to a kit, preferably a kit for implementing the method described above, comprising, or consisting essentially of, or consisting of: a) an in vitro system /platform comprising:
  • nucleic acid translation machinery (iii) a nucleic acid translation machinery; b) at least one material for making microcompartments,; and c) optionally instructions of use.
  • the in vitro platform may be selected from any general or preferred embodiment disclosed above with respect to the first composition provided in step a) of the method according to the invention.
  • a preferred in vitro platform comprised in the kit according to the invention comprises both a bacteriophage Phi29-based replication machinery (as described herein) and a PURE transcription and translation system (as described herein).
  • the kit according to the invention further comprises at least one material for making microcompartments.
  • the material for making microcompartments may be selected from reagents for making microcompartments and devices for making microcompartments.
  • the kit according to the invention may comprise at least one partitioning agent, which may notably be selected from suitable continuous phases and surfactants for generating emulsions of microdroplets.
  • the kit according to the invention may comprise a microfluidic microarray.
  • Both a reagent for making microcompartments and a device for making microcompartments may be included in the kit.
  • FIG. 1 Demonstration of IVTTR reaction using PUREfrex 1.0 and the reconstituted replication machinery of phage Phi29 (P2 P3 P5 P6). End-point GFP fluorescence of 10 pL of IVTTR (grey bar) and qPCR quantification of final ⁇ gfp> concentration (black stroke) after 5 h incubation at 30 ° C.
  • A, B and C, D Replication is necessary for high protein production starting from low amounts of template (3 pM or 12 pM of ⁇ gfp> gene, respectively).
  • E, F, G high replication and protein production is possible with purified P2 and P3 proteins, with a balance between protein production and replication.
  • FIG. 1 End-point GFP fluorescence of 10 pL of IVTTR 1.0 (grey bar) and qPCR quantification of final ⁇ gfp> concentration (black stroke) after 5 h IVTTR at 30 °C.
  • A-D increasing amount of pUC57_OriLR_p2p3 plasmid have non monotonous effect on GFP yield, with a maximum at 200 pM plasmid (B). the effect on final ⁇ gfp> concentration is positive and saturates above 300 pM plasmid concentration (C-G). The optimal pUC57_OriLR_p2p3 concentration is around 200 pM.
  • FIG. 3 10X magnified brightfield (A&C) and green fluorescence (B&D) images of IVTTR 1.0 experiment initiated with ⁇ gfp>replicator at a concentration of less than one copy per compartment.
  • Condition supplemented with dNTP shows a discrete number of fluorescent droplets, indicating replication and expression in droplet for droplet having received one or more copy of ⁇ gfp>, whereas condition without dNTP (C&D) does not show any fluorescence.
  • Figure 4 End-point mCherry fluorescence for 5 pL of IVTTR 2.0 containing ⁇ rfp> replicator and different amounts of PLA-FDL microparticles after 8 h incubation at 33 °C. Adjacent bars represent the result of two replicates. Little difference can be seen in mCherry yield, proving a low toxicity of PLA-FDL microparticles on IVTTR 2.0.
  • FIG. 6 4X magnified brightfield (A&D&G) and green fluorescence (B&E&H) and red fluorescence (C&F&I) images of IVTTR 2.0 (A-l) or PUREfrex 2.0 only (G-l) experiment initiated with ⁇ rfp-hic> replicator at 1 pM in 40 pm droplets.
  • Condition supplemented with dNTP shows a discrete number of fluorescent droplets, indicating replication and expression in droplet for droplet having received one or more copy of ⁇ rfp-hic>. Colocalization of green and red fluorescence indicates that cutinase activity is related to protein production.
  • Condition without dNTP (D-F) does not show any fluorescence.
  • Condition with PUREfrex shows a higher background of green fluorescence (H) than IVTTR 2.0 without dNTP (E) but not comparable with IVTTR 2.0 with fluorescence.
  • FIG. 7 4X magnified green (A), red (B), far-red (C) fluorescence, corresponding respectively to Hie esterase activity, mCherry fluorescence, and a fluorescent dye homogeneously distributed in the initial master mix for droplet tracking, and brightfield (D) images of IVTTR 2.0 mock library experiment after overnight incubation and before sorting.
  • a majority of non -fluorescent droplets in green and red channel indicates a Poisson distribution with a parameter A ⁇ 1.
  • the independence of green and red fluorescence confirms the Poisson distribution and implies clonality of DNA after encapsulation.
  • Figure 8 Scatter plot of screened droplets according to far-red fluorescence (x-axis) and green fluorescence (y-axis). Horizontal line represents green fluorescence threshold above which droplets are sorted (positive hints).
  • Figure 9. qPCR quantification of hie percentage in gene libraries during mock library screening experiment. Standard deviations are calculated from technical triplicate qPCR quantification.
  • FIG. 10 A. Microscopic image of polyesterase activity (top: red fluorescence; bottom: green fluorescence) in microdroplets containing the wild-type HiC replicator in fusion with mCherry, via IVTTR reaction performed at less than one copy per compartment.
  • D Histogram of the ratio of fluorescence signals for all non-empty droplets, showing that a sharp peak for specific activity was obtained by normalizing enzymatic activity with the red expression signal.
  • FIG. 11 A) Flow cytometry analysis of YFP-expressing liposomes in a mock enrichment experiment.
  • the initial library contained a 1 :10 ratio of the yfp and minD genes, except for the 1 nM ⁇ yfp> control sample. Starting from 1 pM ⁇ yfp>, IVTTR leads to higher levels of expressed YFP compared to IVTT samples.
  • IVTTR for a non-catalytic and non-fluorescent protein.
  • A Schematic of inliposomes ori-p3 DNA amplification and expression via IVTTR.
  • B Absolute quantitation of ori-p3 DNA by qPCR in lysed liposomes. Total 10 pM input template DNA concentration was used, which was reduced due to externally supplied DNase I.
  • C Populational variation of dsGreen fluorescence in IVTTR-liposomes measured by flow cytometry.
  • D Quantitation of the fraction of liposomes showing above-background dsGreen fluorescence estimated by the diagonal gate in C.
  • Figure 13 Application of IVTTR for improving lipid synthesis by a membrane-associated enzyme.
  • A Schematic of CDP-DAG conversion to PS by PssA.
  • B Schematic of IVTTR- liposomes expressing PssA enzyme and detection of PS-positive liposomes by LactC2-mCherry binding.
  • C Absolute DNA quantitation by qPCR of lyzed CADGE liposome samples. ** P ⁇ 0.01 ; *** P ⁇ 0.001.
  • D Box plots and (E) quantitation of PS-positive CADGE liposomes expressing PssA, as assayed by flow cytometry.
  • FIG. 14 GFP fluorescence (grey bar) and final replicator concentration (black stroke) in IVTTR 2.0 with ⁇ gfp> replicator and different dNTP compositions in presence (A, C, E) or absence (B, D, F) of DNK gene.
  • a and B conditions have equal concentration dATP, dCTP, dGTP and dTTP.
  • C and D conditions have dTTP replaced by an equal concentration of dTMP
  • E and F conditions have no dNTP.
  • FIG. 15 Brightfield (left) and fluorescent (right) images showing the detection of DNK activity from single copy of encapsulated DNA fragment containing the DNK gene and a fluorescent protein.
  • EXAMPLE 1 In vitro coupled replication-expression reactions provide high expression levels even when starting from picomolar gene concentrations
  • the Inventors first showed that a coupled replication-expression in vitro reaction can produce high level of protein expression even when starting from very low concentration of the encoding gene, whereas a simple in vitro expression reaction cannot, when starting from the same low concentration.
  • Two different ways to introduce the replication machinery were tested: either with purified P2 and P3 proteins, or by in situ production of P2 and P3 from their encoding genes on the plasmid pUC57_OriLR_p2p3. In situ production takes advantage of the protein expression machinery in the mixture to express some of the proteins of the replication machinery from their gene.
  • a GFP encoding replicator was assembled by combining in order the following nucleic acid sequences: Phi29 oriL, T7 promoter, an E. coli ribosome binding site, the gfp gene, a terminator, and Phi29 oriR.
  • the DNA was capped with 5’ -phosphate introduced by PCR using 5’-phosphate primers. This molecule is indicated as ⁇ gfp> (in the following, brackets ⁇ ...> indicate DNA molecules that are competent for replication via the for reconstituted Phi29 machinery).
  • a typical coupled replication-expression reaction was performed with PURE/rex 1.0 (GeneFrontier) and a starting replicator DNA concentration of either 3 pM or 12 pM in 10 pL reaction volume.
  • One picomolar concentration of replicator corresponds to a concentration close to one molecule per micro-compartment, in the case where the microcompartment has an inner volume around 1.5 picoliter (which for example is the volume of droplets of 15- 20 pm in diameter, a size typically used in droplet screening applications).
  • Concentration of pUC57_OriLR_p2p3 was fixed at 200 pM in in situ protein production condition. In purified protein condition, P2 protein (NEB) was fixed at 1 U/pL while P3 (homemade) varied between 0.8 mg/mL and 3.2 mg/mL.
  • Green fluorescence was monitored during 5 h at 30 °C in a standard 96-well plate fluorescence spectrometer (here we use Biorad cfx thermocycler running at constant temperature). Values of final green fluorescence (mature GFP production, grey bars) and final ⁇ gfp> concentration (black strokes) are reported in Figure 1. In the absence of dNTP in the mix, ⁇ gfp> replication is impossible, and GFP fluorescence is not detectable (A&C). In presence of dNTP, GFP fluorescence is easily detected, and we measure that ⁇ gfp> concentration increase by orders of magnitude. This is true using either in situ produced P2 and P3, or purified P2 and P3 (B&D and E&F&G respectively). In the case of purified P3 (F), the existence of an optimal concentration, with a maximal protein production at 1.6 mg/mL P3, suggests a competition effect between replication and protein production.
  • Figure 2 clearly demonstrates an increase followed by a decrease in GFP production with increased pUC57_OriLR_p2p3 concentration, with a maximum fluorescence at 200 pM pUC57_OriLR_p2p3. Replication yield is also increased with increasing pUC57_OriLR_p2p3 concentration but stays high even at the highest concentration (700 pM).
  • the mastermix was split in two: one half was supplemented with 0.3 mM dNTP and the other half with MilliQ water only.
  • Flow focusing device tailored to produce 10 pm diameter droplet was used with fluorinated oil (HFE7500) supplemented with 64 mg/mL (4%) Fluosurf (Emulseo).
  • Emulsions were incubated in PCR machine for 4h at 30° C and then kept at 4°C. Emulsions were spread between hydrophobized slide and coverslip, sealed with epoxy glue to prevent evaporation. Images were taken on epifluorescence microscope (Nikon) with GFP filter at 10X magnification, using the same parameters for both samples. Fluorescence images were inversed for easier interpretation.
  • Humicola insolens cutinase is an enzyme from the esterase family, more precisely a cutinase, with hydrolysis activity toward various esters or polyesters. It is known to degrade cutin and polyesters.
  • PKA solid polyester poly-lactic acid
  • FDL fluorogenic compound fluorescein di-laurate
  • HiC cutinase leads to the release of FDL in solution, which is followed by the hydrolysis of FDL into fluorescein, also catalyzed by the esterase activity of HiC. Therefore, the apparition of green fluorescence indicates polyesterase activity and indirectly the presence of HiC.
  • Fluorescein having absorption and emission spectra close to GFP, we used mCherry protein as reporter and negative control.
  • mCherry replicator is noted ⁇ rfp> for “red fluorescent protein”, while HiC replicator is noted ⁇ hic>.
  • a typical IVTTR reaction was performed with PUREfrex 2.0 (GeneFrontier) and a starting replicator DNA concentration of 10 pM. Reaction volume was set at 5 L.
  • Concentration of pUC57_OriLR_p2p3 was fixed at 200 pM. Green and red fluorescence were monitored during 8 h at 33 °C in CFX. IVTTR experiment was carried with either ⁇ hic> or ⁇ rfp> replicator. Plastic microparticles were supplemented to a final concentration of 4,3 mM, 2.1 mM, 1.1 mM or 0 mM. Each plastic concentration pipetting was performed in duplicate. Red fluorescence produced by ⁇ rfp> replicator was used as a control for plastic toxicity to the replication and expression reaction. Baseline corrected red fluorescence of ⁇ rfp> samples are shown in Figure 4. Figure 4 shows that final mCherry fluorescence does not diminishes with increased amount of PLA-FDL microparticles and thus, there is no significant toxicity of the substrate toward IVTTR.
  • Green fluorescence increases strongly in tubes performing ⁇ hic> replication and expression, showing that it is possible to detect the activity of the enzyme encoded in the replicator, that is, to combine replication, expression and enzymatic assay in one-pot condition.
  • the fluorescence still increases after 8 h of incubation, and samples with higher substrate concentration saturate the CFX sensors, so we assess plastic degrading activity by Vmax measurement.
  • the steepest slope in green fluorescence time trace is shown on Figure 5.
  • Figure 5 shows that ⁇ rfp> replicator shows low fluorescein hydrolysis whereas ⁇ hic> replicator shows increasing fluorescein hydrolysis with substrate concentration. This proves that hydrolysis is performed by HiC cutinase synthesized in situ from the replicator.
  • Figure 5 also shows that that Vmax is approximately doubling when PLA-FDL concentration is doubling, consistent with the HiC enzyme being unsaturated.
  • This example shows that a single ⁇ hic> gene molecule in an IVTTR microcompartment containing a fluorogenic substrate is able to produce an observable amount of fluorescence, but that it is not the case for IVTT conditions (i.e. in absence of genetic replication).
  • Water- in-oil monodisperse microdroplets were used, with an inner volume around the picoliter, for the micro-compartmentalization step.
  • this example shows that genetic replication is required to be able to perform uHTS screening in microcompartments directly from a library of linear genes.
  • DNA constructs were designed to encode two proteins, mCherry and HiC cutinase, expressed as a single fusion peptide.
  • the link between the two proteins includes a thrombine cleaving site LVPRGS flanked by flexible peptide sequences GGGGS, but many other linkers would be suitable.
  • the colocalization between protein production (red fluorescence) and cutinase activity (green fluorescence) will prove that cutinase activity is dependent on ⁇ rfp-hic> presence.
  • a typical IVTTR reaction was performed with PURE/rex 2.0 (GeneFrontier) and a starting replicator DNA concentration of 1 pM in 40 micrometer diameter droplets. Concentration of pUC57_OriLR_p2p3 was fixed at 200 pM. Droplet. The mastermix was split in two: one half was supplemented with 0.3 mM dNTP and the other half with MilliQ water only. A third sample was produced containing only the three solutions from PUREfrex 2.0 kit and the ⁇ rfp- hic> replicator at 1 pM concentration.
  • Figure 6 shows that in condition with dNTP (A-C) some droplets, but not all, exhibit green and red fluorescence, denoting a sub-Poissonian distribution (i.e. an initial distribution of the replicator DNA molecules in the droplets such that some droplet did not receive any replicator molecule, some got only one and a few more than one, according to the Poisson distribution), and the fact that the reaction initiates from single DNA molecules.
  • a sub-Poissonian distribution i.e. an initial distribution of the replicator DNA molecules in the droplets such that some droplet did not receive any replicator molecule, some got only one and a few more than one, according to the Poisson distribution
  • green and red fluorescence are colocalized, meaning that cutinase activity is correlated with mCherry protein, proving that fusion protein is present in the droplet.
  • Condition without dNTP with PURE 2.0 does not show any fluorescence, proving that replication is necessary in an IVTTR process to observe high protein production and enzymatic activity.
  • Condition without dNTP with PUREfrex shows a slightly higher background of green fluorescence (H) than IVTTR with PURE 2.0 without dNTP (E).
  • this background appears in all droplets, and the level is much lower than positive droplets in IVTTR with PURE 2.0 condition, thus it can be interpreted as a higher background hydrolysis of the polyester particles in PUREfrex conditions.
  • EXAMPLE 4 Clonal expression in microdroplets, followed by sorting and enrichment of genes encoding active enzymes using IVTTR of single gene molecules in microdroplets containing a fluorogenic substrate
  • a mock library consisting of a mixture of two DNA sequences, one encoding an active polyesterase (cutinase HiC) and one encoding a catalytically inactive protein (the fluorescent protein mCherry). Both sequences were inserted between the replication origins and the expression elements (promoter, terminator and RBS) as explained above, and thus were able to replicate and produce high level of protein expression.
  • the starting library contained 10% ⁇ hic> and 90% ⁇ rfp> replicators.
  • the reaction mixture was prepared using PUREfrex 2.0 kit and fluorogenic particles to detect esterase activity.
  • the reaction was encapsulated in 47-pm droplets (54 pL volume) in surfactant-containing fluorinated oil. Far- red fluorescent dye was added to the solution to allow detection of all droplets.
  • FIG. 9 summarizes qPCR data of the ratio of the hiC gene in each sample.
  • Starting HiC proportion did not change between mock library DNA mixture and droplets before screening, at around 10%.
  • HiC proportion in the sorted population increased up to 54%, while it decreased in waste population down to 4%.
  • the examples above demonstrate that the IVTTR system can amplify and express longer DNA constructs, for example containing more than one gene, and that this can be used to significantly improve the determination of enzyme activity in microdroplets.
  • a rfp-hic fusion gene (where rfp is the fluorescent protein mCherry and hie is the enzyme) was used to measure specific enzymatic activity, instead of apparent enzymatic activity, as in the previous examples.
  • IVTTR was implemented in droplets from a single DNA template, two fluorescence signals could be observed, one corresponding to the enzymatic activity and one (the red signal) indicating the expression level of the construct. This allowed us to measure in droplets the specific activity.
  • the overall enzymatic activity can be calculated as the product of the specific activity and the enzyme concentration. Since the concentration of expressed protein varies between micro-compartments due to stochastic effects in the encapsulation and dynamics of the expression machinery, we observed a distribution of activity values, even when all the droplets contained the same wild-type gene ( Figure 10). In addition, the expression level can vary from variant to variant, blurring possible differences in specific activity.
  • the concentration of enzyme present in the sample (C) was evaluated and the ratio A/C was computed.
  • a fluorescent tag was attached as an N terminal fusion. Droplet-to-droplet heterogeneity can thereby be reduced by normalizing activity by protein concentration.
  • a construct was created by the genetic fusion of mCherry fluorescent protein and HiC cutinase, connected by a flexible and cleavable linker. Both proteins being covalently coupled, each cutinase enzyme carries an mCherry tag, which allowed us to measure the cutinase concentration via the red fluorescence level.
  • IVTTR mixture was assembled using PURE 2.0 kit and fluorogenic microparticles.
  • IVTTR solution was encapsulated in 42-pm (39 pL) droplets in surfactant-containing fluorinated oil.
  • the emulsion was incubated at 33 °C overnight. The day after, the emulsion was imaged on a microscope slide with appropriate filters for mCherry and fluorescein detection.
  • a control experiment was performed in the same conditions using a mock library of ⁇ hic> and ⁇ rfp> replicator DNA.
  • This example shows that partitioned IVTTR from single genes can be applied to non-catalytic and non-fluorescent proteins, by making use of an indirect fluorescent readout method.
  • TP Terminal Protein
  • phage phi29 which is encoded by p3 gene.
  • the p3 gene was inserted between the replication origins downstream a T7 promotor.
  • IVTTR and detection of the encoded enzyme activity is demonstrated in liposomes containing on average less than a copy of the replicator per compartment.
  • PssA from the E. coli Kennedy phospholipid biosynthesis pathway. This enzyme conjugates cytidine diphosphate-diacylglycerol (CDP-DAG) with L-serine to produce cytidine monophosphate and phosphatydilserine (PS), a precursor of phosphatidylethanolamine (Figure 13A).
  • CDP-DAG cytidine diphosphate-diacylglycerol
  • PS phosphatydilserine
  • Figure 13A The activity of in-vesiculo synthesized PssA enzyme in PURE system from the pssA gene was assayed by preparing liposomes containing 5 mol% CDP-DAG ( Figure 13B).
  • PS lipid was detected by externally staining the liposomes with a PS-specific probe which consists in the C2-domain of lactadherin protein (LactC2) fused to a fluorescent protein like mCherry or eGFP ( Figure 13B) (Blanken et al., 2019).
  • a PS-specific probe which consists in the C2-domain of lactadherin protein (LactC2) fused to a fluorescent protein like mCherry or eGFP
  • Figure 13B Boken et al., 2019.
  • This example demonstrates the possibility to apply the invention to the screening of yet another enzymatic activity and it shows that this screen can be done in microcompartments much larger than liposomes (picoliter scale).
  • This picoliter scale is typical of water-in-oil droplets that are often used for directed evolution protocols based on microfluidic dropletsorting chips.
  • deoxyribonucleoside monophosphate kinase from T5 phage (Uniprot: Q6QGP4). Its activity is the phosphorylation of deoxythymidine monophosphate (dTMP).
  • dTTP nucleoside triphosphate
  • dTMP monophosphate counterpart
  • dTDP deoxynucleoside diphosphate
  • NDK Nucleoside Diphosphate Kinase
  • ADK Adenylate kinase
  • Both enzymes are present in PURE system, so it is not necessary to add them.
  • Newly produced dTTP can finally be used by Phi29 polymerase to amplify DNA flanked by origins of replication.
  • the replicating DNA includes the gene of a fluorescent reporter protein, which gets expressed in high yield and provide a fluorescent readout.
  • the bulk experiment reported in Figure 14 shows that in situ expression of DNK gene can trigger DNA amplification in IVTTR when dTTP was replaced by dTMP, and that no DNA amplification happens in absence of DNK gene.
  • the T5 phage wildtype DNK protein coding sequence was inserted downstream a T7 promoter and synthesized as a double stranded DNA fragment (GeneStrand, Eurofins), diluted in ultrapure water (Merck Milli-Q) and used without further purification. This fragment is not flanked by origins of replication so it is not recognized by Phi29 replication machinery.
  • a IVTTR reaction was prepared using PURE/rex 2.0 (GeneFrontier) and a starting concentration of 12 pM of ⁇ gfp> as the only replicator.
  • P2 and p3 are introduced via an expression plasmid (concentration of pUC57_oriLR_p2p3 plasmid was fixed at 100 pM) while p5 and p6 purified protein were kept at final concentration of 0.375 mg/mL and 0.105 mg/mL respectively and supplemented with 20 mM ammonium sulfate.
  • Six conditions were produced, crossing between presence or absence of 25 pM of DNK gene fragment and composition of dNTP mixture.
  • dNTP mixture was either 0.3 mM of each four dNTP, or 0.3 mM of dATP, dGTP, dCTP and dTMP, or absence of dNTP. Fluorescence of 5 pL reaction mixture was monitored in CFX thermocycler for twenty hours at 33 °C. As green fluorescence of positive conditions saturated detector for FAM channel (>65000 RFU), Cal Gold 540 channel was used. Final fluorescence of reaction mixtures containing all dNTP irrespectively of presence of DNK gene (A and B in Figure 14), and in condition where dTTP was replaced to dTMP in presence of DNK gene (C in Figure 14), increased their fluorescence to around 8.10 3 RFU.
  • the gene encoding DNK enzyme in C-terminal fusion with mCherry red fluorescent protein, downstream T7 promoter and followed by T7 terminator and flanked by origins of replication was obtained as a clonal plasmid was isolated and the sequence was verified via Sanger sequencing, the replicator was then amplified from the plasmid using PCR with phosphorylated primers, giving linear fragment competent for replication.
  • a typical IVTTR reaction was prepared using PURE/rex 2.0 (GeneFrontier).
  • the concentration of pUC57_oriLR_p2p3 plasmid was fixed at 200 pM while p5 and p6 purified protein were kept at final concentration of 0.375 mg/mL and 0.105 mg/mL respectively, and the solution was supplemented with 20 mM ammonium sulfate.
  • 2 pM Dextran, Alexa FluorTM 647; 10 000 MW (Invitrogen) and 2 pM Dextran Alexa FluorTM 680; 10 000 MW (Invitrogen) were added for pipetting control purposes and do not interfere with the reaction.
  • dNTP mixture consisted of dATP, dGTP, dCTP and dTMP, each at a final concentration of 0.3 mM.
  • the replicator containing a single copy of DNK gene and a single copy of mCherry gene, was added at a final concentration of 1 pM, giving a theoretical Poisson parameter . - 1 in 15 pm diameter droplets. This mean that, on average, each droplet contains one copy of the replicator.
  • Flow focusing device tailored to produce 15 pm diameter droplets was used with fluorinated oil (HFE7500) supplemented with 64 mg/mL (4%) Fluosurf (Emulseo). Emulsions were incubated in PCR machine overnight at 33 °C. Emulsions were spread between hydrophobized slide and coverslip, sealed with epoxy glue to prevent evaporation.
  • Enzymes and bacterial strains were purchased from New England Biolabs (NEB) unless specified otherwise.
  • Strain for plasmid isolation is E. coli NEB® 5-alpha.
  • oligonucleotides were purchased from Eurofins with high purity salt free grade purification, except for T22 and T23 which were purified with HPLC grade purification. Sequences of said oligonucleotides are shown in paragraph 7.14 below (Table 1 ).
  • Plasmids pUC57_OriLR_p2p3 was isolated from E. coui by midipreparation using Plasmid Midi kit (Qiagen) and pR045-plVEX-mCherry, pUC57_OriLR_gfp, pUC57_OriLR_rfp-hic_AC and pEX_A128_pANobar were was isolated from E. coli using Plasmid Mini Kit (Macherey Nagel), eluted in MilliQ water. All plasmids were Sanger sequenced for the region of interest. They are reported in Table 4 below. 10.6. Preparation of DNA constructs
  • All replicators are linear and phosphorylated to be able to trigger the P2P3 replication system.
  • All DNA constructs were amplified by PCR using Q5® HotStart polymerase and T22/T23 phosphorylated primers listed in Table 1 below. Typical PCR conditions are as follow: In a PCR microtube were mixed a given amount of DNA template (typically 1 fmol), 2 units of Q5® Hot Start DNA polymerase (NEB), 0.2 mM dNTPs and 500 nM forward and reverse primers in a total volume of 50 pL of 1X Q5 High Fidelity buffer.
  • Replicating DNA encoding GFP protein was obtained by typical PCR described above, using 0.2 fmol pUC57_OriLR_gfp and an elongation time of 37s, yielding a single 1229 bp product visible on agarose gel. Replicator was purified using PCR cleanup kit (Macherey Nagel) following manufacturer’s protocol, eluting DNA with 30 pL MilliQ water. Nucleic acid sequences of gfp gene used is shown in Table 5 below.
  • HiC cutinase from soft-rot fungus Humicola insolens (HiC), known to possess esterase and polyesterase activity (A.M. Ronkvist, 2009)), rfp, and mCherry genes used are shown in Table 5 below.
  • GGA fragments were obtained by PCR reaction with using primers listed in Table 1 below. Regular PCR reactions were performed with 1 fmol of DNA template, with 2 units of Q5® Hot Start DNA polymerase (NEB), 0.2 mM dNTPs and 500 nM forward and reverse primers in a total volume of 50 pL.
  • NEB Q5® Hot Start DNA polymerase
  • PCR reaction consisted of a first stage of 2 cycles at low annealing temperature, then annealing temperature was raised at 72 °C for 20 cycles. Cycles consisted of a denaturing step at 98 °C for 10 s, followed by an annealing step for 10 s, followed by an elongation step at 72 °C for 30 s per kb of amplicon product. Low annealing temperatures and elongation times are reported in Table 2 below. After cycling, a completion step at 72 °C for 2 min was performed.
  • Reaction products were purified with PCR cleanup kit (Macherey Nagel) following manufacturer’s protocol, eluting DNA with 30 pL MilliQ water.
  • Golden Gate Assemblies were performed with 50 fmol of each purified PCR products in a final volume of 20 pL using Bsal-HF®v2 Golden Gate Assembly Kit (NEB). GGA were incubated at 37 °C for 15 min, followed by denaturation at 60 °C for 5 min, yielding circular assembly products called pLR-mCherry or pLR-HiC.
  • GGA product was directly used for PCR amplification of DNA constructs using phosphorylated primers T22 and T23 as described above, using 45 s elongation time, yielding a single product of 1408 bp and 1507 bp for ⁇ hic> and ⁇ rfp> respectively.
  • PCR products were purified with PCR cleanup kit (Macherey Nagel) following manufacturer’s protocol, eluting DNA with 30 pL MilliQ water.
  • PCR cycles consisting in a denaturing step at 98 °C for 10 s, followed by an annealing step at 68 °C for 10 s and an elongation step at 72 ° C for 60 s. After cycling, a completion step at 72 °C for 2 min was performed, yielding a single product of 2072 bp visible on agarose gel.
  • PCR was purified using Zymo -5 kit (Zymo research) following manufacturer’s protocol, eluting DNA with 30 pL of water.
  • PCR product was then used for a second PCR amplification, using primers T22/M75 or M74/T23 for OriL and OriR respectively.
  • Primers M74 and M75 bear a barcode of degenerated sequence NNNNNWNNNNNWNNNNN, where N means an equal amount of ACGT and W means an equal amount of AT during chemical synthesis (sequences of M74 and M75 in Table 1 below).
  • N means an equal amount of ACGT
  • W means an equal amount of AT during chemical synthesis
  • PCR cycles consisting of a denaturing step at 98 °C for 10 s, followed by an annealing step at 55 °C for 10 s and an elongation step at 72 ° C for 20 s.
  • a completion step at 72 °C for 2 min was performed, yielding a single product of 346 bp and 343 bp for OriL and OriR respectively.
  • PCR products were purified using Zymo -5 kit (Zymo research) following manufacturer’s protocol, eluting DNA with 30 pL of water. At this step, half of DNA molecules bear a mismatched barcode.
  • PCR products were purified using Zymo -5 kit (Zymo research) following manufacturer’s protocol, eluting DNA with 30 pL of water. These barcoded products were called XOriLbar and XoriRbar.
  • Barcoded fusion protein replicators ⁇ rfp-hic> were obtained by Golden Gate Assembly from four different DNA fragments called XOriLbar, XOriRbar, prom_rfp and hic_ter. Construction of XoriLbar and XoriRbar is described above. Promrfp and hicter were obtained with PCR on plasmid pUC57_OriLR_rfp-hic_AC using primers M105/M106 or M104/M84 for prom_rfp and hic_ter respectively.
  • PCR reaction consisted of a first stage of 4 cycles 59 °C annealing temperature, then annealing temperature was raised at 72 °C for 18 cycles.
  • Cycles consisted of a denaturing step at 98 °C for 10 s, followed by an annealing step for 10 s, followed by an elongation step at 72 °C for 30 s. After cycling, a completion step at 72 °C for 2 min was performed, yielding a single amplicon product of 813 bp ad 756 bp for prom_rfp and hic_ter respectively.
  • Golden Gate Assembly was performed as follows: in a PCR microtube, 2.5 fmol of each of the four parts were mixed with 0.5 pL of BsmBI-v2 Golden Gate Enzyme Mix (NEB) in 10 pL of 1X ligase buffer.
  • GGA mix were incubated in thermocycler for 30 cycles of 1 min at 42 ° C and 1 min at 16 °C, followed by denaturation step at 65 °C for 5 min.
  • One microliter of GGA mix was directly used for typical PCR amplification using phosphorylated T22/T23 primers as described above, with 66 s elongation time, yielding a single product of 2199 bp.
  • PCR product was purified using Zymo -5 kit (Zymo research) following manufacturer’s protocol, eluting DNA with 30 pL of water.
  • the DNA binding proteins P5 and P6 were prepared according to (Soengas, Gutierrez and Salas, 1995) and (Mencia et al. , “Terminal protein-primed amplification of heterologous DNA with a minimal replication system based on phage ⁇ D29”.
  • PNAS 108, 18655-18660 (2011 ) respectively, and have the following stock concentrations and storage buffers: P5 (10 mg/mL in 50 mM Tris pH 7.5, 60 mM ammonium sulfate, 1 mM EDTA, 7 mM BME, 50% glycerol), P6 (10 mg/ml in 50 mM Tris pH 7.5, 0.1 M ammonium sulfate, 1 mM EDTA, 7 mM BME, 50% glycerol). The proteins were aliquoted and stored at -80 °C.
  • the in vitro transcription and translation systems (IVTT) PURE/rex 1.0 and PURE/rex 2.0 were purchased from Euromedex (France). Three main constituents of the kit are: energy solution, enzymes solution and ribosome solution. They were added to the final mixture as specified by the manufacturer: For PURE/rex 1.0 the proportions are one 1 /2 of energy solution, 1 /20 of enzymes solution and 1 /20 of ribosome solution. For PUREfrex 2.0, the proportions are 1 /2 of energy solution, 1 /20 of enzymes solution and 1 /10 of ribosome solution. For high protein production mimicking a successful replication, the linear DNA was added at a final concentration in the nanomolar range, typically 2 nM.
  • IVTTR In Vitro Transcription Translation and Replication
  • IVTTR The basis of IVTTR is the same as for IVTT.
  • the same kits were used in the same proportions as previously depicted.
  • the solution was typically supplemented with 10 mM ammonium sulfate, 100 pM of pUC57_OriLR_p2p3, 38 pg /mL of purified P5 protein, 112 pg/mL of purified P6 protein and 300 pM of dNTP solution (New England Biolabs).
  • the remaining of the reaction mixture was composed of replicator DNA and substrate particles at different concentrations. In some cases, purified P2 and P3 were used instead of pUC57_OriLR_p2p3.
  • any piece of DNA competent for replication in the IVTTR system is herein called a “replicator” and is noted between two outward pointing chevrons in small letters (e.g. ⁇ rfp> for red fluorescent protein reporter).
  • the DNA construct must be double stranded, linear, and flanked by active origins of replication. The origins are named OriL and OriR, conventionally put on the left side and the right side of the construct, respectively.
  • the coding sequence must be placed downstream of a T7 promoter.
  • a T7 terminator is generally positioned downstream of the gene. We used the promoter and terminator previously described in IVTTR experiment (Van Nies et al.
  • PVA Poly-lactic acid
  • FDA Fluorescein Diacetate substrate
  • solid or PLA-FDL (4%) particles were washed with 3 cycles of: 10000 g, 20 min centrifugation to pellet the particles, removal and addition of 10 mL of clean MilliQ water to remove the residual PVA.
  • a wash with ethanol can be added to remove non-encapsulated FDA.
  • the particles suspension was then filtered with a 5-pm filter to remove potential aggregates.
  • Polyester concentrations were quantified through the fluorescein signal obtained after total degradation by proteinase K.
  • a Low-Profile 0.2 ml 8-Tube Strips BioRad
  • activity buffer 100 mM Tris-HCl pH 7.8
  • Proteinase K New England Biolabs
  • Fluorescence was monitored in CFX machine at 33 °C (50° C lid temperature) until plateau was reached.
  • Final fluorescein concentration was determined using a calibration curve of Dextran FITC 2 MDa (Invitrogen D7137) of known concentration in the same buffer.
  • polyester microparticle concentration was calculated according to Lactic acid monomer equivalent: C3H4O2 (MW 72.06 g/mol).
  • ⁇ yfp> (encoding YFP, a yellow-fluorescent protein) and ⁇ minD> (encoding MinD, a non- fluorescent protein) constructs were mixed at 1 :10 molar ratio and IVTTR reactions were assembled, at a final and total DNA concentration of 10 pM in either expression buffer (PURE/rex 2.0: 50% V/V solution I, 5 % V/V solution II, and 10% V/V solution III and 0.6 units/ul of Superase- In RNase inhibitor) or in replication buffer (PURE/rex 2.0 with an addition of 20 mM ammonium sulfate, 300 pM dNTPs, 375 pg/ml purified P5 protein, 52.5 pg/ml purified P6 protein, 3 ng/pl purified P2 protein, 3 ng/pl purified P3 protein, and 0.6 units/pl of Superase- In RNase inhibitor).
  • expression buffer PURE/rex 2.0: 50% V/V solution I, 5
  • the well-mixed solution was encapsulated in liposomes by adding 10 mg lipid-coated beads and rotating on an automatic tube rotator (VWR) at 4 °C for 30 minutes. The mixtures were then subjected to four freeze/thaw cycles. Then, 10 pl of bead-free liposome suspension was transferred to a PCR tube, where it was mixed with 0.5 units of Proteinase K (Thermo Scientific), and incubated at 30 °C for 16 h. Three microliter of liposome suspension was mixed with 497 pl buffer and filtered through the a 35 pm nylon mesh of the cell-strainer cap from the 5 ml round-bottom polystyrene test tubes (Falcon).
  • Fluorescence-activated cell sorting was conducted on FACSMelody (BD Biosciences). Lasers PE-CF594(YG) and FITC-BB515, 100-micron nozzle, 23.14 PSI pressure and 34.2 kHz drop frequency were used. Photon multiplier tube voltages applied were 320 V for forward scatter, 455 for side scatter, 337 V for Texas Red, and 673 V for GFP, and a threshold of 359 V at the side scatter was applied. Liposomes with 1% highest YFP signal were sorted out from liposomes prepared in expression buffer (gate 1 ), and the same gate was applied to the liposomes prepared in replication buffer (sort 1 ) or an adjusted gate including only 0.2% highest YFP signal (sort 2).
  • Liposome suspension was next used as a template for PCR amplification using phosphorylated primers (ChD 491 /ChD 492). Reactions were set up in 100 pl volume, 300 nM each primer, 400 pM dNTP, 10 pl diluted liposome suspension, and 2 units of KOD Xtreme Hotstart DNA polymerase in Xtreme buffer, and thermal cycling was performed as follows: 2 min at 94 °C for polymerase activation, and thermal cycling at (98 °C for 10 s, 65 C for 20 s, 68 °C for 1 .5 min)x30.
  • the amplified PCR fragments were purified using QIAquick PCR purification buffers (Qiagen) and RNeasy MinElute Cleanup columns (Qiagen) using the manufacturer’s guidelines for QIAquick PCR purification, except for longer pre-elution column drying step (4 min at 10000 g with open columns), and elution with 14 pl ultrapure water (Merck Milli-Q) in the final step.
  • the purified DNA was quantified by the Nanodrop 2000c spectrophotometer (Isogen Life Science).
  • Table 2 List of primers parameters used for PCR amplification of Golden Gate Assembly parts for ⁇ hic> and ⁇ rfp> construction (sequences are in Table 1 above)
  • Table 3 Sequences of GeneStrands (Eurofins) used in the examples
  • Table 4 Sequences of isolated clonal plasmids used in the examples. Uppercase letters are regions verified by Sanger sequencing
  • Table 5 Sequences of gene-encoding replicator DNA used in the examples.

Landscapes

  • Life Sciences & Earth Sciences (AREA)
  • Genetics & Genomics (AREA)
  • Health & Medical Sciences (AREA)
  • Chemical & Material Sciences (AREA)
  • Engineering & Computer Science (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Organic Chemistry (AREA)
  • Zoology (AREA)
  • Biomedical Technology (AREA)
  • Wood Science & Technology (AREA)
  • Biotechnology (AREA)
  • General Engineering & Computer Science (AREA)
  • Crystallography & Structural Chemistry (AREA)
  • Plant Pathology (AREA)
  • Molecular Biology (AREA)
  • Microbiology (AREA)
  • Biophysics (AREA)
  • Physics & Mathematics (AREA)
  • Biochemistry (AREA)
  • General Health & Medical Sciences (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)

Abstract

The present invention provides a new method that is specifically designed to start statistically from a single copy of a nucleic acid molecule and requires only a single encapsulation step to provide clonal microcompartment libraries displaying a high concentration of encoded polypeptides. The method of the invention enables one to perform in vitro uHTS via a one-step encapsulation of linear DNA constructs containing a candidate sequence to be tested in a molecular mixture that allows for simultaneous specific gene amplification and protein expression.

Description

MICRO-COMPARTMENTALIZED ULTRA-HIGH THROUGHPUT SCREENING FROM SINGLE COPY GENE LIBRARIES
TECHNICAL FIELD OF THE INVENTION
The present invention relates to the fields of directed evolution of proteins, protein engineering and screening.
BACKGROUND ART
Directed evolution of proteins
Directed evolution (DE) of proteins (such as enzymes) is an engineering approach, inspired by natural evolution, that allows protein and enzyme engineering without requiring precise knowledge concerning the structural, functional or mechanistic aspects of the protein. It is performed by cycles of genetic diversification and enrichment: starting from a known gene with a given sequence, or a small set of these, the first step is to introduce some diversity so as to obtain a collection (library) of genes that are different in their nucleotide sequences but related to the parental one. These genes are called variants. Various methods are then used to enrich this library in variants with the desired properties. This screening or selection stage performed on the encoded protein variant’s properties eventually allows to recover the genes coding for proteins with the most interesting properties (for example improved catalytic rate, or higher stability, non-natural specificities, etc.), which can be used for a subsequent round of diversification and enrichment. It is also possible to couple this approach with DNA sequencing to identify the beneficial or detrimental mutations with respect to the targeted property. In the past, directed evolution has been applied to improve or alter the properties of a large variety of enzymes. New methodologies have been introduced to manipulate always larger libraries and to screen various types of activities (such as enzymatic activity, fluorescence, binding to a particular receptor/ligand...). In particular, the use of microcompartments has enabled ultrahigh throughput screening methods. Micro-compartmentalized ultrahigh throughput screening (uHTS) methods use random encapsulation of variants in individual microcompartments to manipulate large libraries of variants or candidate genes. This random encapsulation is done at a concentration that is low enough that many microcompartments end up receiving 0 or 1 member of the library, and only a small proportion of the microcompartments receives more than 1 member (e.g. 2 or more). This process is referred as Poissonian partitioning and enables clonality, i.e. the fact that the signal observed in most microcompartments is associated to a unique variant of the library. Because of that, the microcompartments obtained from Poissonian partitioning of variants are sometime called “monoclonal” microcompartments. (Holstein, J. M., Gylstorff, C. & Hollfelder, F. Cell-free Directed Evolution of a Protease in Microdroplets at Ultrahigh Throughput. ACS Synthetic Biology acssynbio.0c00538 (2021 ) doi:10.1021 /acssynbio.0c00538). By recovering the microcompartments displaying the most interesting signals, one can therefore enrich the library in interesting variants. These microcompartments can be for example water-in-oil microdroplets, generated and manipulated using microfluidic devices, or liposomes. The volume of these microcompartments ranges between the femtoliter and the microliter.
In vitro methods for directed evolution
Traditionally, living organisms are used to express the protein of interest from its encoding gene. The expressed recombinant protein is generally trapped in the cytoplasm or the periplasm of the transformed cell, or displayed on the cell surface. In all these cases, the cell itself maintains a strong phenotype-genotype linkage, because it encapsulates both the encoding gene (for example as a multi-copy number plasmid) and the encoded protein. However, it may be beneficial to avoid the use of cells.
In vitro evolution of proteins has been developed as a subset of directed evolution techniques, where the DE process is carried out in a cell-free environment. The advantage of performing DE in vitro lies in the ability to engineer cytotoxic proteins, proteins with functions requiring unusual, unnatural, or cell-impermeable substrates, or use of lethal selection conditions. In addition, the introduction of the gene variants inside the living hosts (a step called transformation) is often a throughput-limiting step in workflow using in vivo protein expression, meaning that it limits the size of the library that can be manipulated. Removing this constraint would open the route to DE protocols with a higher throughput than currently available with in vivo approaches.
Genotype-phenotype linkage
In the absence of a living organism, the genotype-phenotype linkage needs to be maintained artificially. In addition, a requirement for efficient screening of a library of proteins is that most microcompartments do not contain more than one of the gene library members, as explained above.
One approach for genotype- phenotype linkage in the absence of a cell is to create a direct physical connection between the encoding genes and their gene products, so called ‘display systems'. Various display systems have been developed over the years to perform this function: phage display, ribosome display, mRNA display, CIS display, SNAP display, etc. However, these ‘display systems' provide only one, or just a few copies of the protein of interest attached to each gene variant and in some cases, e.g., phage display, still require an in vivo step which is performed within live transformed cells. Such ‘display system’ can be used for selection based on affinity (panning), but are in many cases insufficient to assess the catalytic properties of an enzyme, or other functional protein properties (e.g., fluorescence). Indeed, it is generally difficult to detect enzymatic activity or other protein properties from a single molecule or just a few molecules.
In-droplet amplification and expression
Another option to maintain the genotype- phenotype linkage than establishing a physical interaction between the gene and its protein product, is to confine both of them within the same microcompartment. In the case of uHTS techniques via emulsions microdroplets, liposomes, or microfabricated chambers, the microcompartments may serve this function, in addition to their role of separating the genetic variants from one-another.
It is possible to express a protein from its gene without resorting to a living cell but using an in vitro expression platform such as cell extracts or the PURE system (Shimizu, Y. et al. Cell- free translation reconstituted with purified components. Nat. Biotechnol. 19, 751-755 (2001 )). This reaction is called IVTT (for “in vitro transcription and translation”). In this case, no living cell is involved. To use IVTT reaction in microcompartments for screening applications, it is necessary than no more than one variant gene is present in most compartments, which, in in vitro settings, is obtained by dilution of the gene library before encapsulation, so that compartments typically receive no more than one gene molecule (Poissonian partitioning). However, the use of high dilution, necessary to ensure Poissonian partitioning with many compartments receiving one or less variant, also results in low yields of protein expression by the IVTT reaction and therefore in low screening efficiency.
A general approach for in vitro uHTS would require to obtain better performance of the IVTT reaction in each compartment, i.e. produce more protein from the unique copy of the encapsulated gene and hence obtain more signal related to the activity of this protein within the sorting apparatus. Indeed, when a sufficient amount of protein is produced in the compartment, it becomes easier to differentiate positive and negative microcompartments (e.g., those displaying an interesting enzymatic activity from those that do not) and therefore to recover only the microcompartments containing the genes encoding active protein variants. However, current in vitro methods to obtain multiple copies of the same genetic variants in each microcompartment, while maintaining clonality, are cumbersome. Several methods have been developed to mitigate the poor expression levels arising from using a single molecule of DNA template, which are based on pre-amplifying the gene variants in a clonal fashion.
The first approach uses isothermal amplification that creates concatemers. For example rolling circle amplification (RCA), starts from a single circular DNA molecule containing a gene and creates multiple catenated copies of the same gene. Thus a single molecule contains multiple copies of the gene information, thus increasing the gene concentration even for single encapsulation events. In this case however, preprocessing steps are necessary for concatemer generation and, in the case of a library, it is not guaranteed that single DNA molecules contain copies of a single variant of the library.
A second approach uses droplet-PCR amplification from single encapsulated gene followed by picoinjection of the IVTT (in vitro transcription and translation) mix (Mazutis, L. et al. Droplet-Based Microfluidic Systems for High-Throughput Single DNA Molecule Isothermal Amplification and Analysis. Anal Chem 81 , 4813-4821 (2009)).
However, this method utilize multiple steps. This increases complexity and requires costly droplet-based microfluidic handling. For example, a 2021 study describes a method for cell- free directed evolution at ultrahigh throughput using such a droplet-based microfluidic device. In this method, the library of gene variants in a circular form is distributed in microdroplets following Poissonian partitioning (where the concentration of the library was adjusted so that -10% of the droplet would contain a single variant and most droplet would contain no variant) and first amplified by rolling circle amplification. The droplets are then pico-injected with an IVTT mixture for protein expression. Then, these droplets are picoinjected again with the substrate of the enzyme of interest. Finally, after incubation, the resulting droplets can be sorted according to the fluorescence signal associated with the enzymatic transformation of the substrate (Holstein, J. M., Gylstorff, C. & Hollfelder, F. Cell-free Directed Evolution of a Protease in Microdroplets at Ultrahigh Throughput. ACS Synthetic Biology acssynbio.0c00538 (2021 ) doi:10.1021 /acssynbio.0c00538). This tedious protocol with multiple microfluidic steps is technically challenging and limits the throughput of the method.
A third approach avoids the cumbersome multiple microfluidic manipulation of droplets, by using so-called bead-display (or microbead display, see A Sepp, DS Tawfik, AD Griffiths. Microbead display by in vitro compartmentalisation: selection for binding using flow cytometry. FEBS letters 532 (3), 455-458). In this case, the gene amplification (for example by PCR) and expression reactions are performed beforehand, on beads, in a clonal fashion. The resulting microbeads carry both multiple clonal copies of the gene, and multiple copies of the corresponding protein. As such, they physically maintain the phenotype-genotype linkage and can be manipulated in bulk. These beads are subsequently encapsulated in droplets with a fluorogenic substrate to assay the activity of the carried protein, and these droplets can be sorted according to detected activity levels, in order to recover the beads carrying the most active protein variants and ultimately the genes encoding these variants. However, the protocol for bead display also involves multiple preparatory steps before the screening can be performed.
For example, a bead display technique allowing the display of up to millions of copies of the protein of interest together with its gene was developed in 2013, based on the SNAP-tag system (Diamante, L., Gatti-Lafranconi, P., Schaerli, Y. & Hollfelder, F. In vitro affinity screening of protein and peptide binders by megavalent bead surface display. Protein engineering, design & selection: PEDS 26, 713-724 (2013)). This approach requires a complex multistep protocol for bead preparation, encapsulation and sorting, which increases the workload and introduces many possibilities for failure.
A fourth option, which is particularly desirable in terms of its simplicity, would be a “one- pot” process combining amplification/replication of a single compartmentalized nucleic acid molecule, with protein expression, all in the same mixture without external intervention. Combining replication of DNA with expression systems (IVTT, such as the PURE system) in a single compartment has previously been attempted. However, significant difficulties arose due to the low compatibility of expression systems (including the PURE system) with DNA replication. For example, PCR cannot be used in this context because the thermal cycles would destroy the protein expression machinery. Although this compatibility can be improved by selecting specific amplification schemes and adjusting the reaction conditions (e.g. decreasing ribonucleotide triphosphate concentration), known approaches still require high DNA concentrations to be efficient. In the case of the microcompartments required for ultrahigh throughput screening applications, these required high concentrations would necessitate many initial DNA copies in each compartment, incompatible with random Poissonian partitioning and clonal expression (Han et al, Applied microbiology and biotechnology, 106(24), 2022; Libicher et al, Nature Communication, 11 (1 ), 2020).
It would thus be advantageous to have a simple method for in vitro uHTS screening that could work by direct encapsulation of single gene copies (notably in linear form), such as PCR products, and could be directly followed by screening after incubation, without additional steps, while avoiding complex microfluidic fusion or pico-injection protocols.
Thus, known methods to perform micro-compartmentalized in vitro uHTS for functional protein (e.g., enzymes, fluorescent proteins, binders...) require cumbersome protocols in order to create microcompartments that contain both a single genetic variant (in one or multiple copies) and the corresponding protein /enzyme in sufficient amount so that its activity is detectable by the sorting apparatus within the microcompartment, allowing high frequency sorting. In particular, current methods to obtain multiple copies of the same genetic variants in each microcompartment, while maintaining clonality, are burdensome and sources of errors, as they require to perform multiple complex steps before or after compartmentalization. As detailed above, current strategies rely on a preamplification stage on a solid support (bead display) or a complex microfluidic protocol using multiple fusion/picoinjection stages.
Therefore, there is a pressing need to provide novel micro-compartmentalized uHTS methods for screening proteins having activities of interest, free from the above-mentioned cumbersome constraints and which work by direct Poissonian partitioning of single DNA molecules into microcompartments.
The present invention fulfils this need. Indeed, the present Inventors have developed a new method that is specifically designed to amplify a single copy of a nucleic acid molecule encoding a protein (such as a nucleic acid in a standard linear form (e.g., PCR products)) and simultaneously express the encoded protein and assay its activity within a microcompartment. The method is based on the combined use of an IVTT system and a system capable of performing nucleic acid replication and amplification in vitro (i.e. in a non-cellular context). In the examples provided, but this is not limitative, a replicator linear DNA comprising, in addition to the gene to be encoded and its regulatory sequences, origins of replication (such as phage Phi29 origins of replication or derivatives thereof) in 5’ of each DNA strand is used, and a minimal replication machinery (such as phage Phi29 p2, p3, p5 and p6 proteins) is added, either as purified proteins or as DNA encoding them.
Van Nies et al. (“Self-replication of DNA by its encoded proteins in liposome-based synthetic cells”. NATURE COMMUNICATIONS (2018)9:1583) recently presented an in vitro reaction called IVTTR (i.e. IVTT coupled to an in vitro system of Replication) where the reconstituted phage Phi29 replication machinery enables replication of linear DNA fragments simultaneously to the expression of the encoded proteins in an IVTT mixture. This coupled replication-expression reaction was however only demonstrated for the expression of specific proteins from phage Phi29 (replication proteins p2, p3, p5 and p6) and not for other non-Phi29 proteins. Moreover, it required a large initial concentration of the replicating DNA, incompatible with Poissonian partitioning in micro-compartments, such as microdroplets or liposomes.
In the context of the invention, the inventors surprisingly discovered that i) the decoupled replication-expression can be used to express any protein of interest, or any library of DNA fragments (e.g. encoding variants of the protein of interest); ii) that this modified reaction can be performed starting from a very low concentration (preferably using poissonian partitioning of single nucleic acid molecules into microcompartments) of replicating DNA carrying the gene(s) of interest; iii) that it is possible to couple replication with transcription/translation ; iv) that this reaction is compatible with enzymatic assays, in particular fluorogenic assays, v) that when performed in micro-compartment starting from sufficiently low initial concentration of library of DNA fragments (preferably using poissonian partitioning), it results in a population of clonal micro-compartments containing a high copy number of one of DNA fragment from the library and a high copy number of its encoded protein, and producing a signal related to the activity of the encapsulated protein, which can be used in a ultra-High Throughput Sorting apparatuses, and vi) that altogether, this approach enables the in vitro ultra-High Throughput Sorting of libraries of proteins or enzymes directly by encapsulation of their gene libraries, without intermediate steps. Poissonian partitioning enables clonality, i.e., ensures that only one copy of nucleic acid molecule (i.e., a unique variant of the library of the library to be tested) is incorporated in the majority of micro-compartments. The present invention thus provides an original, reliable, rapid and efficient method to perform clonal in vitro microcompartmentalized uHTS via a one-step encapsulation of nucleic acid molecules (such as linear DNA constructs) containing the genetic variant (or candidate sequence) to be tested, in a molecular mixture that allows for simultaneous specific gene amplification and protein expression and optionally activity/enzymatic assay.
SUMMARY OF THE INVENTION
The present invention thus relates to a novel method for screening a library of nucleic acid molecules each encoding a candidate polypeptide for a polypeptide with an activity of interest by direct encapsulation. In this method, (i) a first composition comprising an in vitro platform comprising a nucleic acid replication machinery, optionally a nucleic acid transcription machinery, and a nucleic acid translation machinery, and (ii) a second composition comprising a library of nucleic acid molecules each comprising at least one candidate sequence encoding at least one candidate polypeptide, are provided and mixed. The thus obtained mixture is partitioned (preferably using poissonian partitioning) into microcompartments wherein at least one of said microcompartments comprises a single copy of said nucleic acid molecule. Poissonian partitioning enables clonality, i.e., the fact that most microcompartments contains a unique variant of the library (i.e., only one copy of a nucleic acid molecule of the library to be tested is incorporated in most of the microcompartments). Then, in the microcompartments, the candidate sequence is amplified by the nucleic acid replication machinery in the microcompartments, optionally transcribed by the nucleic acid transcription machinery, and translated in at least one candidate polypeptide by the nucleic acid translation machinery.
The activity of the thus obtained candidate polypeptide is then detected and the candidate sequence encoding the candidate polypeptide having the activity of interest is finally recovered using a sorting apparatus.
The present invention also concerns a kit comprising an in vitro platform as defined above. The kit also comprises at least one ingredient for making microcompartments.
DETAILED DESCRIPTION OF THE INVENTION
Definitions
Unless specifically defined herein, all technical and scientific terms used herein have the same meaning as commonly understood by a skilled scientist in chemistry, biochemistry, cellular biology, molecular biology, protein engineering and medical sciences.
As used herein, when used to define products, compositions, cell lines, uses and methods, the term "comprising" (and any form of comprising, such as "comprise" and "comprises"), "having" (and any form of having, such as "have" and "has"), "including" (and any form of including, such as "includes" and "include") or "containing" (and any form of containing, such as "contains" and "contain") are open-ended and do not exclude additional, unrecited elements or method steps. "Consisting of" means excluding any other components or steps "Consisting essentially of" means excluding other components or steps of any essential significance (however, other minor/insignificant components or steps are not excluded). In the present disclosure, the terms “comprising”, “consisting of” and “consisting essentially of” may be replaced with each other, if required.
As used herein, “directed evolution” refers to an engineering approach, inspired by natural evolution, that allows protein and enzyme engineering without requiring precise knowledge concerning the structural, functional or mechanistic aspects of the protein. It is performed by cycles of genetic diversification and enrichment: starting from a known gene with a given sequence, or a small set of these, the first step is to introduce some diversity so as to obtain a collection (library) of genes that are different in their nucleotide sequences but related to the parental one. Various methods are then used to enrich this library in variants with the desired properties. This screening or selection stage performed on the encoded protein variant’s properties eventually allows to recover the genes coding for proteins with the most interesting properties, which can be used for a subsequent round of diversification and enrichment. It is also possible to couple this approach with DNA sequencing to identify the beneficial or detrimental mutations with respect to the targeted property. As used herein, “bioprospection” and “functional metagenomics” are used interchangeably and refer to another approach for the identification of proteins with desired activity. In bioprospection, the screened libraries are libraries of natural microorganism, or cells or can be isolated nucleic acid sequences. These libraries of nucleic acid strands may be obtained by extraction and cloning of nucleic acides coming from microorganisms, plants or animals or a plurality of those.
As used herein, the term “screening” or “sorting” or “selecting” (all terms are herein considered synonyms) refers to a process for testing and selecting compounds/active agents (in particular proteins and/or nucleic acid molecule encoding such proteins) for a specific effect/activity.
By “candidate sequence”, it is herein meant a genetic sequence (i.e., a nucleotide sequence) encoding a candidate polypeptide whose activity is to be tested in the method of the invention. In the context of directed evolution, candidate sequences (also referred to as genetic variants or mutants) generally encode polypeptide variants or mutants (i.e the original polypeptide with one or more mutations selected from substitutions, insertions, deletions and permutations) of a known polypeptide with the desired activity, with the purpose to identify variants with increased desired activity or improved properties compared to the known polypeptide.
As used herein, a “library” of nucleic acid molecules refers to a mixture of several and preferably many (such as one or more hundreds up to billions or more) of distinct nucleic acid molecules.
As used herein, the term “partitioning” refers to separating a sample into a plurality of portions, also referred to as “partitions” or “compartments”. Partitions are generally physical, such that a sample in one partition does not, or does not substantially, mix with a sample in an adjacent partition.
The term “micro-compartment” or “microcompartment” herein means a compartment with an inner volume between the femtoliter and the microliter (i.e., a characteristic size below 1 mm). Microcompartments include, but are not limited to, microdroplets (such as water-in-oil microdroplets), vesicles (such as liposomes), microchambers, etc. It may be advantageous that the compartments used in a method of screening according to the invention are identical in shape and volume, however, collection of microcompartiments of different size (obtained for example by liposome forming protocol, or by bulk emulsification protocols) can also be used. Methods of obtaining microcompartments are well known of the skilled person. In the case of microdroplets and emulsions, such methods include various emulsification techniques, such as membrane emulsification, mechanical emulsification, microfluidic emulsification, high homogenization, microfluidization and ultrasonication, among others. Emulsions are generally stabilized by appropriate surfactants known to the skilled person, which can provide them with long term stability. For example, microdroplet can be obtained as mono-disperse emulsions using microfluidic approaches or a number of other approaches.
As used herein, the term “micro-com partmentalisation” refers to the separation of reactions into multiple microreactors, such as multiple micro-compartments.
The terms "polynucleotide", “nucleic acid molecule”, and "nucleic acid" are used interchangeably herein and are understood as a polymeric or oligomeric macromolecule made from nucleotide monomers (also called nucleotide residues). Nucleotide monomers are composed of a nucleobase, a sugar or a sugar analogue (such as but not limited to ribose or 2'-deoxyribosefor RNA and DNA, but other sugars or sugar analogues may be used for xeno nucleic acids referred to as XNA, such as 1 ,5-anhydrohexitol for HNA, cyclohexene for CeNA, threose for TNA, glycol for GNA, a ribose modified with an extra bridge connecting the 2' oxygen and 4' carbon for locked nucleic acid or LNA, or N-(2-aminoethyl)-glycine units for peptide nucleic acid or PNA), and one to three phosphate groups. Typically, a polynucleotide is formed through phosphodiester bonds linking ribose or deoxyribose groups in the case of natural nucleic acids, but other artificial bonds may be used in the case of XNA (such as peptide bonds in the case of PNA) between the individual nucleotide monomers. Nucleic acid molecules include, but are not limited to, ribonucleic acid (RNA), deoxyribonucleic acid (DNA), and mixtures thereof such as, e.g., RNA-DNA hybrids (mixed polyribo- polydeoxyribonucleotides). These terms encompass single or double-stranded, linear or circular, natural or synthetic, unmodified or modified versions thereof (e.g., genetically modified polynucleotides; optimized polynucleotides), sense or antisense polynucleotides, chimeric mixture (e.g., RNA-DNA hybrids). Moreover, a polynucleotide may comprise non- naturally occurring nucleotides and may be interrupted by non-nucleotide components. Exemplary DNA nucleic acids include without limitations, complementary DNA (cDNA), genomic DNA, plasmid DNA, DNA vector, viral DNA (e.g., viral genomes, viral vectors), oligonucleotides, probes, primers, satellite DNA, microsatellite DNA, coding DNA, non-coding DNA, antisense DNA, and any mixture thereof. Exemplary RNA nucleic acids include, without limitations, messenger RNA (mRNA), precursor messenger RNA (pre-mRNA), small interfering RNA (siRNA), short hairpin RNA (shRNA), microRNA (miRNA), RNA vector, viral RNA, guide RNA (gRNA), antisense RNA, coding RNA, non-coding RNA, antisense RNA, satellite RNA, small cytoplasmic RNA, small nuclear RNA, etc. Polynucleotides described herein may be synthesized by standard methods known in the art, e.g., by use of an automated DNA synthesizer (such as those that are commercially available from Biosearch, Applied Biosystems, etc.) or by DNA assembly and gene synthesis method, or by mutagenesis methods or obtained from a naturally occurring source (e.g., a genome, cDNA, etc.) or an artificial source (such as a commercially available library, a plasmid, etc.) using molecular biology techniques well known in the art (e.g., cloning, PCR, etc.). The nucleic acids can be synthesized chemically, e.g., in accordance with the phosphotriester method (see, for example, Uhlmann, E. & Peyman, A. (1990) Chemical Reviews, 90, 543-584).
In the context of this invention, a “replicator” is a nucleic acid molecule able to replicate under specific conditions in the presence of specific compounds (the “nucleic acid replication machinery”). The replicator acts as a template for its amplification, so that the newly-made nucleic acid molecules are copies or reverse-complemented copies of it. By this replication, the concentration of the replicator will increase. For example, if one starts from a single molecule of the replicator, after some time, 2, then 4, 8 etc molecules may be present, i.e. its concentration will increase exponentially. A replicator can be any type of nucleic acid (DNA, RNA or hybrids or XNA) and may contain specific subsequences or non- canonical modifications. For example, a replicator preferably includes a subsequence selected from an origin of replication (OR), a primer binding site, one or more genes with appropriate regulatory sequences (e.g. a promoter), a ribosome-binding site (RBS), barcodes, etc.
The terms "protein" and "polypeptide" are used interchangeably herein and refer to any peptide-bond-linked polymer of amino acids, regardless of length or post-translational modification. These terms preferably refer to polymers of amino acid residues comprising at least six amino acids covalently linked by peptide bonds. The polymer can be linear, branched or cyclic. The polymer may comprise naturally occurring and/or amino acid analogues and it may be interrupted by non-amino acids. No limitation is placed on the maximum number of amino acids comprised in a polypeptide. As a general indication, the term refers to both short polymers (typically designated in the art as peptide, or protein fragment) and to longer polymers (typically designated in the art as polypeptide or protein). This term encompasses native polypeptides, modified polypeptides (also designated derivatives, analogues, variants, mutants), polypeptide fragments, polypeptide multimers (e.g., dimers), mutated polypeptides, engineered polypeptides, fusion polypeptides among others. A polypeptide is understood to be any translational product of a polynucleotide regardless of size, and whether glycosylated or not, and includes peptides and proteins. Polypeptides/Proteins usable herein (including protein derivatives, protein variants, protein fragments, protein domains, protein epitopes and protein domains) can be further modified by chemical or enzymatic modification. This means that such a chemically modified polypeptide or enzymatically modified polypeptide comprises other chemical groups than the 20 naturally occurring amino acids. Examples of such chemical or enzymatic modifications include post-translational modifications. Chemical or enzymatic modifications of a polypeptide may provide advantageous properties as compared to the parent polypeptide, e.g., one or more of enhanced stability, increased biological half-life, increased water solubility, increased activity, enhanced properties, labelling, etc.
As a general indication and without being bound therein, if the amino acid polymer contains more than 50 amino acid residues, it is preferably referred to as a polypeptide or a protein, whereas if the polymer consists of 50 or fewer amino acids, it is preferably referred to as a "peptide". The reading and writing senses of an amino acid sequence of a polypeptide, protein and peptide as used herein are the conventional reading and writing senses. The reading and writing convention for amino acid sequences of a polypeptide, protein and peptide places the amino terminus on the left, with the sequence then being written and read from the amino terminus (N-terminus) to the carboxyl terminus (C-terminus), from left to right.
The terms “functional protein” or “functional polypeptide” refer to a protein or polypeptide having an activity of interest, by opposition to a non-functional protein or peptide without the activity of interest (e.g. a mutant of a functional polypeptide that has lost the activity of interest). In the case of directed evolution of proteins, a library of protein variants, some of which are functional and some of which are non-functional, is constructed and then screened using the method according to the invention, in order to recover the functional variants.
When referring to a protein, the term “activity”, “activity of interest” or “function” refers to any biological activity of the protein that is screened for when using the method according to the invention. Activities of interest notably include enzymatic activity, reporter activity (e.g. fluorescence activity), regulatory activity (e.g. transcription factor activity) and biological activity (e.g. drug or antibiotic activity).
“Enzymatic activity” refers to the specific catalytic activity of an enzyme. “Enzyme" herein refers to a protein with catalytic properties (enzymatic properties). Most biomolecules capable of catalysing chemical reactions in cells are enzymes; however, some catalytic biomolecules are made of RNA and are therefore distinct from protein enzymes: these are ribozymes. An enzyme works by lowering the activation energy of a chemical reaction, which increases the speed of the reaction. The enzyme is either not modified during the reaction or regenerated unmodified at the end of the catalytic reaction. The initial molecules are the “substrates” of the enzyme, and the molecules formed from these substrates are the products of the reaction. Enzymes are often characterized by their very high specificity. Moreover, most enzyme have the characteristic of being reusable, in the sense that the enzyme can catalyse many conversions of substrate into product. Enzymes are generally globular proteins that act alone or in complexes of several enzymes or subunits. Like all proteins, enzymes consist of one or more polypeptide chains folded to form a three-dimensional structure corresponding to their native state. However, some enzyme can be fully or partially disordered proteins. Many enzymes are composed of more than one peptide chain or assemble as multimeric assemblies.
Their size can vary from about 50 residues to more than 2000 residues. Only a very small part of the enzyme - between two and four residues most often, sometimes more - is directly involved in catalysis, the so-called catalytic site located in the catalytic domain. The catalytic site may be located in the vicinity of one or more binding sites, at which the substrate(s) is (are) bound and oriented to catalyse the chemical reaction. The catalytic site and the binding sites form the active site of the enzyme.
Enzymes perform a large number of functions in living organisms. For example, they can be involved in signal transduction and regulation of cellular processes, in the generation of movement, in active transmembrane transport, in digestion, in metabolism, in the immune system, in nucleic acid digestion or cleavage mechanisms or in nucleic acid production (referred to here as "nucleic acid-acting enzymes"), in prodrug conversion mechanisms (prodrug-to-drug conversion). The enzyme is preferably a prokaryotic, eukaryotic or viral enzyme, preferably an enzyme from an animal, plant, alga, microalgae, insect, microorganism, archaea, bacterium, parasite, yeast, fungus or virus. Among animals, the enzyme may be a mammalian enzyme, such as a human enzyme. The various categories of enzymes are well known to the person skilled in the art, who can refer in particular to reference works in the field (such as Schomburg D., Schomburg I., Springer Handbook of Enzymes. 2 edn. Heidelberg: Springer; 2001 -2009; Liebecq C., IUPAC-IUBMB Joint Commission on Biochemical Nomenclature (JCBN) and Nomenclature Committee of IUBMB (NC-IUBMB) Biochem. Mol. Biol. Int. 1997;43:1151 -1156; IUBMB (1992), Enzyme Nomenclature 1992, Academic Press, San Diego; and specialized databases as described in Schomburg D, Schomburg I. Methods Mol Biol. 2010;609:113-28). Enzyme databases; in particular, the BRENDA database (available, inter alia, at brenda-enzymes.org), as described, for example, by Chang A, Schomburg I, Placzek S, Jeske L, Ulbrich M, Xiao M, Sensen CW, Schomburg D, Nucleic Acids Res. 2015 Jan;43. Epub 2014 Nov 5. BRENDA in 2015: exciting developments in its 25th year of existence). Enzymes are also classified based on their type of activity in the Enzyme Commission number (EC number) classification, and any enzyme of any subcategory of EC numbers 1 to 7 is of interest in the present invention. It is contemplated that the term enzyme encompasses native enzymes and derivatives thereof (e.g. mutated and/or engineered enzymes), provided that such derivative is capable of having an enzymatic activity. As used herein, an “enzyme fragment” is any portion of an enzyme, preferably provided that such fragment/portion is capable of having an enzymatic activity. In the case of a protein enzyme, the enzyme fragment preferably comprises at least 6 consecutive amino acid residues of the enzyme (and is preferably an enzyme catalytic site) (preferably at least 8 consecutive amino acid residues of the enzyme, preferably at least 10, preferably at least 15, preferably at least 20, preferably at least 30 amino acid residues of the enzyme).
Some enzymes use “cofactors”, which term refers to a non-protein molecule that forms a stable or transient complex with an enzyme and whose presence is essential for the activity of an enzyme. Cofactors can be inorganic ions, particularly metallic ions or other organic molecules, such as FAD or NADH.
Some enzymes use “coenzymes”, which term refers to a non-protein organic molecule that forms a complex with an enzyme and whose presence is essential for the activity of an enzyme. Coenzymes may be divided into two types. The first is called a "prosthetic group", which consists of a coenzyme that is tightly (or even covalently) and permanently bound to a protein. The second type of coenzymes are called "cosubstrates", and are transiently bound to the protein. Cosubstrates may be released from a protein at some point, and then rebind later.
“Reporter activity” refers to the specific activity of a reporter protein. A “reporter protein” or “reporter polypeptide” refers to a protein that is directly or indirectly detectable in the screening step of the method of the invention, preferably by optical means. Reporter proteins directly detectable in the screening step notably include coloured or fluorescent proteins. “Fluorescent proteins” are proteins that emit light when exposed to light or other radiation, the emitted light having a longer wavelength than the light to which the protein has been exposed. Non-limiting examples of fluorescent proteins include blue (e.g. EBFP2, mTagBFP2), cyan (e.g. mTurquoise, mTurquoise2, mCerulean, mCerulean3, mCerulean3), UV-excitable green (e.g. mT-Sapphire), green (e.g. EGFP, mEGFP, Emerald, mEmerald, sfGFP), yellow-green (e.g. YFP, mPapaya, YPet, Citrine, mCitrine, Venus, mVenus, Topaz, mTopaz, Clover, mClover, mNeonGreen), orange (e.g. mOrange, m0range2, mKO, mK02), orange-red (e.g. tdTomato, TagRFP, TagRFP-T, DsRed2), and red (e.g. mRuby, mRuby2, mApple, mRFP1 , mCherry, FusionRed), and far-red (e.g. mKate2, mNeptune, mCardinal, mPlum) fluorescent proteins (Cranfill PJ, Sell BR, Baird AAA, Allen JR, Lavagnino Z, de Gruiter HM, Kremers GJ, Davidson MW, Ustione A, Piston DW. Quantitative assessment of fluorescent proteins. Nat Methods. 2016 Jul;13(7): 557-62). Reporter proteins indirectly detectable in the screening step notably include enzymes that produce a product that can be easily detected for example by spectroscopic means (e.g. a colored, fluorescent or luminescent product) from a precursor molecule (e.g. a chromogenic compound, a fluorogenic compound or a luciferin). A “chromogenic compound” is a compound that produces color when modified by an enzyme, a “fluorogenic compound” is a compound that produces fluorescence when modified by an enzyme, and a “luciferin” is a compound that produces light when modified by an enzyme.
“Regulatory activity” and “transcription factor activity” are used interchangeably and refer to the specific transcription-inducing activity of a transcription factor. The term “transcription factor” (or “TF”) as used therein refers to a protein that controls the transcription rate of a DNA to an mRNA through a specific binding mechanism. All TFs share an essential feature: they contain at least one DNA-binding domain (DBD). Examples of TFs include, but are not limited to, AP-1 , CREB, C/EBP, c-Myc, NF-1 , TFII factors, etc. The various categories of TFs are well known to the person skilled in the art, who can refer in particular to reference works in the field (such as Yusuf, Dimas, et al. "The transcription factor encyclopedia." Genome biology 13.3 (2012): 1 -25). It is contemplated that the term transcription factor encompasses native TFs and derivatives thereof (e.g., mutated and/or engineered TFs), provided that such derivative is capable of controlling the transcription rate of a DNA to an mRNA through a specific binding mechanism. As used herein, an “TF fragment” is any portion of a TF, preferably provided that such fragment/ portion is capable of controlling the transcription rate of a DNA to an mRNA through a specific binding mechanism (e.g., DBDs, etc.). In the case of a protein TF, the TF fragment preferably comprises at least 6 consecutive amino acid residues of the TF (and is preferably a DBD) (preferably at least 8 consecutive amino acid residues of the TF, preferably at least 10, preferably at least 15, preferably at least 20, preferably at least 30 amino acid residues of the TF).
The terms “peptide” or “protein fragment" or "part of a peptide or protein" or “protein domain” herein mean a portion of a peptide or protein, i.e., a portion of the sequence of consecutive amino acids making up said peptide or protein (referred to as the peptide or protein from which the fragment is derived). When the fragment is a peptide, protein, the fragment preferably comprises at least 6 consecutive amino acids of the peptide or protein from which it is derived; more preferably at least 8 consecutive amino acids, more preferably at least 10 consecutive amino acids, more preferably at least 12 consecutive amino acids, more preferably at least 15 consecutive amino acids, more preferably at least 20 consecutive amino acids, more preferably at least 30 consecutive amino acids of the peptide or protein from which it is derived. When the fragment is a peptide or protein fragment, the fragment preferably has a three-dimensional structure, under non-denaturing conditions (e.g., conditions that are usually non-denaturing for proteins, especially in the absence of denaturing and/or chaotropic agents). When the fragment is a peptide or protein fragment, the fragment is preferably a functional fragment. “Functional fragment" means any peptide or protein fragment, having at least one of the original functions of the peptide or protein from which said fragment is derived. Preferably, the functional fragment performs said function with an efficiency equal to at least 30% of that of said peptide or protein or molecule, preferably at least 40%, preferably at least 45%, preferably at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, preferably at least 99%, preferably at least 100% of the efficacy of said peptide or protein. Examples of protein fragments (in particular of functional fragments) include, e.g., protein domains, protein epitopes, etc. Fragments usable herein (including protein fragments, peptide fragment, protein domain, protein epitopes and protein domains) can be further modified by chemical or enzymatic modification (e.g., post- translational modification(s)). When the fragment is a peptide or protein fragment, a chemically/enzymatically modified fragment comprises other chemical groups than the 20 naturally occurring amino acids (e.g., can comprise post-translational modification(s)).
The terms “protein derivative" or "derivative" herein mean a mutated protein and/or an engineered protein and/or a mimetic. The derivative is preferably a functional derivative. “Functional derivative" means any protein derivative having at least one of the original functions of the protein from which it is derived. Preferably, the functional derivative performs said function with an efficiency equal to at least 30% of that of said protein, preferably at least 40%, preferably at least 45%, preferably at least 50%, preferably at least 55%, preferably at least 60%, preferably at least 65%, preferably at least 70%, preferably at least 75%, preferably at least 80%, preferably at least 85%, preferably at least 90%, preferably at least 91%, preferably at least 92%, preferably at least 93%, preferably at least 94%, preferably at least 95%, preferably at least 96%, preferably at least 97%, preferably at least 98%, preferably at least 99%, preferably at least 100% of the efficacy of said protein.
By "three-dimensional structure" or "tertiary structure" it is herein referred to the intrinsic folding of a molecule in space. A molecule with a three-dimensional structure is a molecule with a stable spatial configuration, (commonly called folding), which is its own and which is, in general, intimately linked to its function. When this structure is dissociated by the use of denaturing or chaotropic agents, the molecule is said to be denatured and loses its function. In particular, a three-dimensional structure is a structure with limited flexibility. In contrast, a non-three-dimensional structure can adopt a dynamic set of configurations that constantly change over time. In the case of a protein, a peptide, or a fragment of these, the three-dimensional structure is the folding of the polypeptide chain in space. In this case, the three-dimensional structure is not a linear chain of amino acids that can adopt a dynamic set of configurations constantly changing over time (it is not a linear succession of amino acids without any spatial configuration). The three-dimensional structure of proteins, peptides, mixed molecules comprising a protein or a peptide, or fragments of these, is maintained by different interactions which can be: covalent interactions (disulfide bridges between cysteines); electrostatic interactions (ionic bonds, hydrogen bonds); van der Waals interactions; interactions with the solvent and the environment (ions, lipids...).
As used herein, “post-translational modification” refers to a chemical or enzymatic modification occurring naturally or not on a protein or a protein fragment, after or concomitantly to protein translation (e.g., biological or biochemical synthesis, e.g., using cellular machinery), or after or concomitantly to protein synthesis (e.g., artificial and/or chemical synthesis). This means that at least one of the naturally occurring amino acid of the protein or protein fragment is modified by the addition of at least one chemical group and/or the modification (including, but not limited to, the removal) of at least one chemical group of the naturally occurring amino acid. Examples of such chemical or enzymatic modifications include without limitation glycosylation, phosphorylation, acylation, carboxylation, acetylation, biotinylation, hydroxylation, lipoylation, amidation, ubiquitination, sumoylation, deamination, etc. By “post-translationally modified protein” it is herein referred to a protein having at least one (i.e. , one or more) post-translational modification. By “post-translationally modified protein fragment” it is herein referred to a protein fragment having at least one (i.e., one or more) post-translational modification. The term “platform” and the term “system” are herein considered equivalent.
It refers to an in vitro material equipment made of one or more elements that are operatively linked together to achieve a particular purpose. A platform according to the invention is preferably a cell-free system. Such a cell-free system can be selected from cell extracts, mixtures of purified cell components, reconstituted in vitro transcription translation (IVTT) systems, DNA amplification and replication systems, enzymatic and protein activity assay and any combination thereof. In the context of the present invention, the platform is designed to achieve at least two biological processes, namely isothermal genetic replication and protein expression.
In particular, the platform according to the invention combines an IVTT system and a system capable of performing replication in vitro to obtain an In Vitro Transcription, Translation, and Replication system (hereinafter referred to as “IVTTR” system), in the very particular context of microcompartments comprising only one nucleic acid molecule comprising a candidate sequence encoding a candidate polypeptide.
As used herein, the term “machinery” refers to a biological equipment made of one or more materials that functionally cooperate to achieve a biological mechanism.
It is provided in the method of the present invention a “nucleic acid replication machinery”. This nucleic acid replication machinery thus comprises or consists essentially of or consists of all the required materials for completing nucleic acid replication. Particularly, it comprises one or more materials selected from: DNA polymerases, RNA polymerases, DNA terminal proteins, Single-Stranded DNA Binding Proteins (SSBPs), Double- Stranded DNA Binding Proteins (DSBPs), nucleic acid sequences coding for DNA polymerases, nucleic acid sequences coding for RNA polymerases, nucleic acid sequences coding for DNA terminal proteins, nucleic acid sequences coding for SSBPs, nucleic acid sequences coding for DSBPs, the necessary cofactors and substrates (such as e.g. nucleoside triphosphates, ammonium sulfate), accessory nucleic acids (such as primers) and any combination thereof. Other proteins enzymatic activity may be included as well. For example, those involved in the replication of DNA in living organism or in viruses, such as recombinase, topoisomerase, transposases, reverse transcriptase, helicase, nuclease processivity factors, enzyme associated with nucleic acid metabolism, among others, may also be comprised in the in vitro platform of the first composition.
It is optionally provided in the method of the present invention a “nucleic acid transcription machinery”. This nucleic acid transcription machinery thus comprises or consists essentially of or consists of all the required materials for completing nucleic acid transcription starting from a DNA and ending by an mRNA. Particularly, it comprises one or more materials selected from: RNA polymerases, nucleic acid sequences coding for RNA polymerases, and any combination thereof (and substrates and cofactors and accessory enzymes and proteins such as transcription factors (TF), etc...)
It is provided in the method of the present invention a “nucleic acid translation machinery”. This nucleic acid translation machinery thus comprises or consists essentially of or consists of all the required materials for completing nucleic acid translation starting from an mRNA and ending by a polypeptide.
Particularly, it comprises one or more ribosomes, translation factors (e.g., initiation factors, elongation factors, release factors, NTP-recycling enzymes), aminoacyl tRNA synthetases, tRNAs, and the like.
Any one of the above defined machineries can either be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages (such as Phi29, T7, QB, MS2, or T4), bacteria (such as Escherichia coli), yeasts (such as Saccharomyces cerevisiae), viruses (such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus), other eucaryotic cells (e.g. rabbit retyculocytes) or any combination thereof.
Any one of the above defined machineries may be provided either as a cellular extract or as a reconstituted system. A “cellular extract” refers to an extract of cellular components and may be used in crude (i.e. without any prior purification) or more or less purified (at least one purification step enriching the extract in the desired compounds) form. A “reconstituted system” refers to a reconstituted mixture of previously purified or recombinantly expressed components (e.g. the PURE system as a reconstituted transcription and translation system, or the mixture of Phi29 bacteriophage proteins p2, p3, p5 and p6 as a reconstituted replication system, both of which are described in more details below).
The term “parameter” refers to a feature, in particular a physical, quantifiable feature (such as, concentration, temperature, pH, viscosity, and the like) whose value is used to appropriately represent, define or classify an entity or a situation, in particular at a given time point.
The term “substrate” herein refers to a starting material that is subject to a chemical or biological reaction, in particular an enzymatic reaction in the presence of an enzyme. The substrate can advantageously be coupled and/or fused to a compound or moiety having detectable optical properties. Usually, a reaction (e.g., an enzymatic reaction or a chemical reaction) starts from one or more substrates and ends with one or more products. Thus, in this context, a “product” is the material resulting from a reaction starting with a substrate. As used herein, the terms “optical technique” or “optical method” or “optical property” refer to a technique or a method or a property that allows optical detection and/or measure. An optical technique can be selected from absorption -based detection techniques (including single or multi-wavelength measurements, etc.), fluorescence-based detection techniques (including fluorescence intensity, fluorescence lifetime, fluorescence wavelength, FRET, etc.), and any combination thereof. An optical property can be inter alia an electromagnetic and/or color and/or fluorescent property happening in visible or non-visible wavelengths. A “fluorescent compound” refers to a compound that emits light when exposed to light or other radiation, the emitted light having a longer wavelength than the light to which the compound has been exposed. This light emission can be advantageously detected and/or measured by an optical technique as defined above. The term “compound” herein refers to a chemical or biological entity, such as a molecule. Illustrative examples of compounds are, without limitation, reagents, substrates, cosubstrates, cofactors, coenzymes, DNAs, RNAs, hormones, antigens, epitopes, ligands, antibodies, receptors, toxins, and the like.
The expression “at least one microcompartment comprises a single copy of a nucleic acid molecule” herein means that the average number of nucleic acid molecules per microcompartment (corresponding to the ratio between the sum of the individual nucleic acid molecule contents of the microcompartments and the total number of microcompartments) is between 0 and 10, preferably between 0 and 5, yet preferably between 0 and 2, and most preferably between 0.01 and 1.
For example, if the average number of nucleic acid molecules per microcompartment is 1 following a random distribution of the DNA molecules in microcompartments of identical inner volume, it means that it is actually obtained about 37% of empty microcompartments, about 37% of microcompartments containing one nucleic acid molecule, and about 26% of microcompartments containing 2 or more nucleic acid molecules. These values are computed from the well-known “Poisson distribution” formula P(k)= Ak.e-k/k! where P(k) is the probability to observe a compartment containing k DNA molecules, and A is the average number of DNA per droplet. A is easily computed knowing the DNA concentration in the encapsulated mixture and the volume of individual compartments. Such partitioning of DNA molecules in microcompartments, that allows many microcompartments to receive only a single molecule of DNA (and where necessarily many compartments do not receive any molecule of DNA) is referred to as “Poissonian partitioning”. This Poissonian partitioning is very important in all microcompartmentalized uHTS methods because it allows to maintain clonality and genotype-phenotype linkage at a statistical level. To do so, the concentration of objects to be encapsulated (DNA molecules here but may be bacteria, or beads, in other approaches) so adjusted so that only a minor fraction of the microcompartments receive more than one object. The result is a collection of microcompartments that, for the most part, contain either 0 or 1 object. This approach thus allows to isolate most objects in individual microcompartments, without resorting to a deterministic partitioning.
As used herein, the expression “at least one nucleic acid molecule comprising at least one candidate sequence encoding at least one polypeptide” means that:
The nucleic acid molecule can comprise more than one candidate sequence; and/or Each candidate sequence can encode more than one polypeptide; and/or
Each candidate sequence can be present as a single copy in the nucleic acid molecule. Method of screening
While known in vitro uHTS methods for screening functional protein (e.g. enzymes, fluorescent proteins, binders...) require cumbersome protocols, the present Inventors have developed a new system that is specifically designed to start statistically from a single copy of a nucleic acid molecule (such as a nucleic acid in a standard linear form (e.g., PCR products)) and requires only a single encapsulation step to provide clonal microcompartment libraries containing numerous copies of a given nucleic acid molecule containing one or more gene and a high concentration of the encoded protein(s), which can be directly submitted to the sorting stage. The present invention thus provides an original, reliable and efficient method to perform in vitro uHTS via a one-step encapsulation of nucleic acid molecules (such as linear DNA constructs) containing the genetic variant (or candidate sequence) to be tested in a molecular mixture that allows for simultaneous specific gene amplification and protein expression.
Indeed, the present Inventors have unexpectedly demonstrated that it is possible to obtain levels of protein activity that are sufficient for uHTS screening from the encapsulation according to a Poisson distribution of a single copy of a nucleic acid molecule, in linear form, in microcompartments containing a platform combining two activities, isothermal genetic replication and protein expression. Poissonian partitioning enables clonality, i.e. the fact that most microcompartments contains a unique variant of the library. The platform thus developed by the Inventors combines an IVTT system with a system capable of performing replication of this unique variant in each microcompartment, to obtain an IVTTR system. This innovative platform relies on the association of (i) a nucleic acid replication machinery, (ii) optionally a nucleic acid transcription machinery (depending on the type of the starting nucleic acid molecule used in the screening method), and (iii) a nucleic acid translation machinery. The data obtained by the Inventors demonstrate that the association of both the replication and the translation machineries in a one-pot reaction results in the concomitant nucleic acid amplification and protein production of any gene, in an amount that is sufficient for automatized detection of the protein activity of interest, even when starting initially from a single copy of the gene-containing nucleic acid molecule within the microcompartment. The Inventors have thus successfully applied this novel method to the mutational screening or bioprospection of any protein of interest.
Because this innovative approach relies on the direct encapsulation of a mixture, and does not involve procedures that are complex or multistep or limited in throughput (such as microfluidic droplet manipulation, bacterial transformation, preamplification or microbead manipulations), the data confirm it can be successfully used to manipulate large libraries in order to enrich them in genes encoding proteins with targeted properties. Importantly, since it is a compartmentalized approach, this novel method is compatible with the enrichment of any protein library, including enzyme libraries and libraires of protein with other activities, such as fluorescent proteins or transcription factors (TF). This is in contrast with uHTS approaches that are only compatible with the selection of binders.
More specifically, the Inventors have shown that a protein having an activity of interest (such as an enzyme, a protein of a protein complex, a nucleic acid binding protein, a transcription factor, and any combination thereof) can be efficiently and reliably screened at a very high rate (e.g., millions of variants per hour), using this system.
This novel method provides the following benefits, as supported by the experimental data:
(i) the nucleic acid molecule, even when initially present as of a single copy in a microcompartment, undergoes cycles of isothermal exponential amplification driven by the replication machinery, and, simultaneously, the accumulating gene copies serve as a template for protein translation, leading to the synthesis of high and easily detectable concentrations of the gene products;
(ii) the approach is advantageous for screening applications because it exponentially replicates the nucleic acid molecule in the microcompartment, going from at least one of a single copy at time of encapsulation up to nanomolar concentrations in the microcompartment when entering the sorting stage. It then becomes easier to recover the genetic material at the end of the screening round, when all sorted positive micro-compartments are pooled and their content analyzed or used as starting library for the next directed evolution cycle. This amplification increases the amount of nucleic acid molecule that is recovered after screening, even if just a few compartments are sorted to the positive bin, and simplifies the further processing of the selected library (i.e. , recovery of sorted genes);
(iii)this approach is compatible with various micro-encapsulation strategies, such as liposomes or water-in-oil microdroplets;
(iv)any type of nucleic acid molecule can be used as starting material, including DNA (such as cDNA, single stranded DNA, double stranded DNA, partially double-stranded DNA, linear DNA, circular DNA, capped DNA, uncapped DNA, etc.) and RNA (such as mRNA, capped RNA, uncapped RNA, DNA/RNA hybrid, vector (such as a plasmid or a viral vector). For example, in vitro replication can be achieved using terminal P3 capped DNA containing specific origins of replication. In vitro replication can alternatively be initiated by 5’phosphate capped DNA, which is easy to obtain (using 5’P primer during the PCR, or restriction of adapters, or enzymatic phosphorylation of DNA). Therefore, a library composed of standard linear 5’P DNA molecules, which are easily obtained from a PCR performed with 5’P primers, or by enzymatic phosphorylation of 5 ’OH DNA, can be directly used;
(v) the genetic amplification process may be compatible with long DNA or RNA constructs (up to multi kilobases constructs) so that it is possible to include additional functionalities in the replicator, which can be used to ameliorate the protein screening process;
(vi)it is possible to change the experimental conditions at any stage (such as between the IVTTR phase and the protein property assaying phase), without resorting to complex microfluidic manipulation (such as picoinjection droplet merging, etc.). For example, the temperature can be easily adjusted, during the replication and IVTT reactions, and then during the assessment of the targeted protein activity. It is also possible to change the concentrations in the microcompartments at any stage, using evaporation, osmotic exchange or diffusion of a compound from the continuous phase without breaking the droplets or liposomes. When liposomes are used, they may also contain pores that allow small molecule to diffuse in from the continuous solution. Finally, some method allowing to deliver compounds with microemulsion compartments using nano-emulsions have been described (Bernath, Kalia, Shlomo Magdassi, and Dan S. Tawfik. "Directed evolution of protein inhibitors of DNA- nucleases by in vitro compartmentalization (IVC) and nano-droplet delivery." Journal of molecular biology 345.5 (2005): 1015-1026) In other words, the activity assay conditions are not limited to the one that is required for IVTTR reactions.
The present invention thus relates to a novel method for screening a library of nucleic acid molecules encoding each a candidate polypeptide for one or more nucleic acid molecule(s) encoding a polypeptide with an activity of interest, comprising, or consisting essentially of, or consisting of: a) Providing a first composition comprising an in vitro platform comprising:
(i) a nucleic acid replication machinery,
(ii) optionally a nucleic acid transcription machinery, and
(iii) a nucleic acid translation machinery; b) Providing a second composition comprising a library of nucleic acid molecules each comprising at least one candidate sequence encoding at least one candidate polypeptide; c) Mixing said first composition of step a) with said second composition of step b); d) Partitioning the thus obtained mixture into microcompartments, wherein at least one of said microcompartments comprises a single copy of said nucleic acid molecule; e) Allowing said candidate sequence to be amplified by the nucleic acid replication machinery in the microcompartments; f) Optionally allowing the amplified candidate sequence of step e) to be transcribed by the nucleic acid transcription machinery in the microcompartments; g) Allowing the amplified candidate sequence of step e) or the transcribed candidate sequence of step f) to be translated in at least one candidate polypeptide by the nucleic acid translation machinery in the microcompartments; h) Optionally adding at least one compound in the mixture of step c) or in the microcompartments of step g), and/or modifying at least one parameter within the microcompartments of step d) or g); wherein said at least one compound is preferably selected from reagents, substrates, cofactors, coenzymes of the candidate polypeptide encoded by the candidate sequence and any combination thereof; and/or wherein said at least one parameter is preferably selected from concentration, temperature, pH, viscosity, and any combination thereof; i) Detecting the activity of the candidate polypeptide of step g) in the conditions that are optionally modified in step h) within the microcompartments, preferably using an optical method; and j) Using a sorting apparatus, recovering the candidate sequences encoding the candidate polypeptide having the activity of interest from microcompartments displaying specific activity signal.
The method of the invention is particularly adapted for high-throughput screenings (HTS), preferably ultra-high-throughput screenings (uHTS). In particular, the method of the invention allows high or ultrahigh throughput screening of at least 105 nucleic acid molecules comprising at least one candidate sequence, preferably at least 106 nucleic acid molecules comprising at least one candidate sequence.
While the method of screening of the invention is applicable to any activity of interest, an activity of particular interest is the activity of enzymatic polymer degradation. The polymer may notably be a plastic polymer, in particular selected from PLA (polylactic acid), PET (polyethylene terephthalate), PE (polyethylene), PP (polypropylene), PCL (polycaprolactone), and PS (polystyrene). In one embodiment, the polypeptide with an activity of interest is selected from the group consisting of an enzyme, a polypeptide of a protein complex, a nucleic acid binding protein, a transcription factor, and any combination thereof (i.e., the method is preferably for screening a library of nucleic acid molecules encoding each a candidate polypeptide selected from the group consisting of a candidate enzyme, a candidate polypeptide of a protein complex, a candidate nucleic acid binding protein, a transcription factor, and any combination thereof). Step a) - Providing a first composition comprising an in vitro platform
In step a), a first composition comprising an in vitro platform comprising:
(i) a nucleic acid replication machinery,
(ii) optionally a nucleic acid transcription machinery, and
(iii) a nucleic acid translation machinery is provided.
As explained above, the nucleic acid replication machinery comprises or consists essentially of or consists of all the required materials for completing nucleic acid replication. Particularly, it may comprise or consist essentially of or consist of one or more materials selected from: DN polymerases, RN polymerases, DN terminal proteins (TP), Single- Stranded DNA Binding Proteins (SSBPs), Double-Stranded DNA Binding Proteins (DSBPs), nucleic acid sequences coding for DNA polymerases, nucleic acid sequences coding for RNA polymerases, nucleic acid sequences coding for DNA terminal proteins, nucleic acid sequences coding for SSBPs, nucleic acid sequences coding for DSBPs, the necessary cofactors and substrates (such as e.g. nucleosides triphosphates), accessory nucleic acids (such as primers) and any combination thereof.
The combination of materials necessary for completing nucleic acid replication varies depending on the specific type of replication machinery that is selected.
For instance, when the minimal replication machinery of phage Phi29 is used, the nucleic acid replication machinery preferably comprises (or consists essentially of or consists of) a DNA polymerase, a DNA terminal protein (TP), a Single-Stranded DNA Binding Protein (SSBP), and a Double-Stranded DNA Binding Proteins (DSBP). Since the replication reaction happens in an IVTT mixture, one or more of these enzymes can be replaced by its encoding DNA with appropriate regulatory sequence (promotor, Ribosome binding site and terminator). In that case the enzyme(s) required for replication is produced in situ by in vitro transcription and translation before the replication reaction happens. Even more preferably, when the minimal replication machinery of phage Phi29 is used, it comprises (or consists essentially of or consists of):
• a DNA polymerase, a DNA terminal protein (TP), a Single-Stranded DNA Binding Protein (SSBP), and a Double-Stranded DNA Binding Proteins (DSBP); or
• a nucleic acid sequence coding for a DNA polymerase, a nucleic acid sequence coding for a DNA TP, a Single-Stranded DNA Binding Protein (SSBP), and a Double-Stranded DNA Binding Proteins (DSBP). However, other viruses or organisms may have other replication mechanisms, some of which do not use a DNA terminal protein (TP).
As a result, a general minimal replication machinery will preferably comprise (or consist essentially or consist of) a DNA or RNA polymerase, a Single-Stranded DNA Binding Protein (SSBP), and a Double-Stranded DNA Binding Proteins (DSBP). This minimal replication machinery may then be completed by other materials, depending on the specific type of replication machinery that is selected.
The proteins included in the replication machinery (or encoded by a nucleic acid molecule included in the replication machinery) may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages (such as Phi29, T7, QB, MS2, or T4), bacteria (such as Escherichia coli), yeasts (such as Saccharomyces cerevisiae), viruses (such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus), other eucaryotic cells (e.g. rabbit reticulocytes) or any combination thereof.
Preferably, the proteins included in the replication machinery (or encoded by a nucleic acid molecule included in the replication machinery) are selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages, and more preferably from bacteriophage Phi29 (it is then referred to as a “Phi29-based replication machinery”). In this latter case, the replication machinery preferably comprises (or consist essentially or consist of) a DNA polymerase, a DNA terminal protein (TP), a Single-Stranded DNA Binding Protein (SSBP), and a Double-Stranded DNA Binding Proteins (DSBP), wherein the DNA polymerase corresponds to bacteriophage Phi29 protein p2, the DNA terminal protein (TP) to bacteriophage Phi29 protein p3, the Single-Stranded DNA Binding Protein (SSBP) to bacteriophage Phi29 protein p5 and the Double-Stranded DNA Binding Proteins (DSBP) to bacteriophage Phi29 protein p6. In a particularly preferred embodiment, the replication machinery comprises (or consists essentially of or consists of):
• purified Phi29 proteins p2, p3, p5 and p6 (preferably purified recombinant Phi29 proteins p2, p3, p5 and p6 of amino acid sequences SEQ ID NO:45 to 48); or
• nucleic acid sequences coding for Phi29 proteins p2 and p3 (preferably pUC57_OriLR_p2p3 of sequence SEQ ID NO:34), and purified Phi29 proteins p5 and p6 (preferably purified recombinant Phi29 proteins p5 and p6 of amino acid sequences SEQ ID NO:47 and 48). 1
Alternatively, other replication machineries may be used. For instance, a reconstituted plasmidic replication machinery (Ueno H et al. “Amplification of over 100 kbp DNA from Single Template Molecules in Femtoliter Droplets _ACS Synth Biol. 2021 Sep 17; 10(9):2179- 2186), a reconstituted chromosomal replication machinery (Su’etsugu, Masayuki, et al. "Exponential propagation of large circular DNA by reconstitution of a chromosomereplication cycle." Nucleic acids research 45.20 (2017): 11525-11534) or a reconstituted viral replication machinery (Kulczyk AW et al. “The Replication System of Bacteriophage T7” Enzymes. 2016;39:89-136) may be used.
The above-described proteins of the replication machinery may be provided either as a cellular extract or as a reconstituted system. Preferably, a reconstituted system comprising the protein mixtures defined above is used.
Other proteins with enzymatic activity may be included as well. For example, those involved in the replication of DNA in living organisms or in viruses, such as recombinase, topoisomerase, transposases, reverse transcriptase, helicase, nuclease processivity factors, enzyme associated with nucleic acid metabolism, among others, may also be comprised in the nucleic acid replication machinery. Those also may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages (such as Phi29, T7, QB, MS2, or T4), bacteria (such as Escherichia coli), yeasts (such as Saccharomyces cerevisiae), viruses (such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus), other eucaryotic cells (e.g. rabbit reticulocytes) or any combination thereof.
The replication machinery should also contain deoxyribonucleotides triphosphate (dNTP).
The in vitro platform in the first composition optionally comprises a nucleic acid transcription machinery, depending on the type of nucleic acids present in the second composition provided in step b). When the second composition of step b) comprises mRNA, no nucleic acid transcription machinery is needed. In other cases, notably when the second composition of step b) comprises DNA, the in vitro platform in the first composition also comprises a nucleic acid transcription machinery.
As explained above, the nucleic acid transcription machinery comprises or consists essentially of or consists of all the required materials for completing nucleic acid transcription starting from a DNA and ending by an mRNA. Particularly, it comprises one or more materials selected from: RNA polymerases, nucleic acid sequences coding for RNA polymerases, and any combination thereof (and substrates and cofactors and accessory enzymes and proteins such as transcription factors (TF), etc...). Preferably, the nucleic acid transcription machinery comprises an RNA polymerase, for example T7 RNA polymerase.
A nucleic acid translation machinery is also comprised in the in vitro platform of the first composition. This nucleic acid translation machinery comprises or consists essentially of or consists of all the required materials for completing nucleic acid translation starting from an mRNA and ending by a polypeptide. Particularly, it comprises one or more ribosomes, translation factors (e.g., initiation factors, elongation factors, release factors, NTP-recycling enzymes), aminoacyl tRNA synthetases, tRNAs, and the like.
The proteins included in or encoded by a nucleic acid molecule included in the nucleic acid translation machinery and included in or encoded by a nucleic acid molecule included in the optional nucleic acid transcription machinery may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteriophages (such as Phi29, T7, QB, MS2, or T4), bacteria (such as Escherichia coli), yeasts (such as Saccharomyces cerevisiae), viruses (such as Venezuelan equine encephalitis virus, adeno associated virus, norovirus, or influenza virus), other eucaryotic cells (e.g. rabbit reticulocytes) or any combination thereof. They may more particularly may be selected from or derived (e.g., obtained after one or more modifications (e.g., genetic modifications such as genetic mutations), or engineered) from natural or wild-type machineries from bacteria, in particular Escherichia coli.
The in vitro platform comprised in the first composition is preferably a cell-free system, which may be provided as a cellular extract, a reconstituted system and any combination thereof. Preferably, a reconstituted system is used.
When a nucleic acid transcription machinery is provided, the nucleic acid translation machinery and the nucleic acid transcription machinery may preferably be provided in a single reconstituted system, preferably a purified gene expression machinery from E. coli, such as the PURE system (for “Protein synthesis Using Recombinant Elements”), which is a commercially available from various suppliers (for example the “PUREfrex” system from GeneFrontier) reconstituted nucleic acid transcription and translation machinery able to isothermally express any Open Reading Frame under T7 promoter and E. coli Ribosome Binding Site, comprising 32 individually purified components: initiation factors IF1 , IF2, and IF3; elongation factors EF-G, EF-Tu, and EF-Ts; release factors RF1 and RF3; ribosomerecycling factor (RRF), 20 aminoacyl-tRNA synthetases (ARSs), methionyl-tRNA transformylase (MTF), T7 RNA polymerase, and ribosomes. In addition, the PURE system contains 46 tRNAs, nucleosides triphosphate (NTPs), creatine phosphate, 10-formyl-5, 6,7,8- tetrahydrofolic acid, 20 amino acids, creatine kinase, myokinase, nucleoside-diphosphate kinase, and pyrophosphatase (Shimizu, Y. et al. Cell-free translation reconstituted with purified components. Nat. Biotechnol. 19, 751-755 (2001 )). Three PUREfrex versions exist. PUREfrex 2.0 is an upgraded version of PURE 1.0 for protein production, PUREfrex 2.1 is a version of PUREfrex 2.0 lacking any reducing agent, allowing to tune oxidoreductive environment of the solution.
The first composition provided in step a) may further comprise other components useful in the method of the invention. For instance, when the method is applied to the screening of nucleic acid molecules encoding candidate enzymes, a substrate, a cofactor, a coenzyme (prosthetic group or cosubstrate) of the enzyme or any combination thereof may be added in the first composition provided in step a).
Step b) - Providing a second composition comprising a library of nucleic acid molecules each comprising at least one candidate sequence encoding at least one candidate polypeptide
In step b), a second composition comprising a library of nucleic acid molecules each comprising at least one candidate sequence encoding at least one polypeptide is provided.
The nucleic acid molecules are preferably selected from the group consisting of DNAs (such as gDNA, cDNA), RNAs (such as mRNA), and any combination thereof (such as DNA/RNA hybrids). The nucleic acid molecule may be single-stranded, double-stranded, or partially double-stranded. The nucleic acid molecule may be linear or circular. It can originate from in vitro preparation (e.g., chemical synthesis, PCR, isothermal amplification...) or in vivo preparation (e.g., plasmid, viral vectors...). It can contain specific subsequences (in addition to the candidate sequence encoding a polypeptide mentioned) and specific caps, such as terminal proteins. The nucleic acid molecule may additionally contain (non-canonical) modifications, such as non-canonical bases, epigenetic marks, internal modifications or terminal modifications.
Each nucleic acid molecule of the library is preferably a replicator, i.e. a nucleic acid molecule able to replicate under specific conditions in the presence of specific compounds (the “nucleic acid replication machinery”). The replicator acts as a template for its amplification, so that the newly-made nucleic acid molecules are copies or reverse- complemented copies of it. By this replication, the concentration of the replicator will increase. For example, if one starts from a single molecule of the replicator, after some time, 2, then 4, 8 etc molecules may be present, i.e. its concentration will increase exponentially. A replicator can be any type of nucleic acid (DNA, RNA or hybrids) and may contain specific subsequences or non-canonical modifications (see above). For example, a replicator preferably includes one or more subsequence(s) selected from an origin of replication (OR), a primer binding site, one or more genes with appropriate regulatory sequences (e.g. a promoter, terminator and ribosome-binding site (RBS)). More preferably, a replicator comprises an origin of replication (OR), a primer binding site, one or more genes with appropriate regulatory sequences (e.g. a promoter), and a ribosome-binding site (RBS).
When a Phi29-based replication machinery (comprising Phi29 p2, p3, optionally p5 and optionally p6 proteins or nucleic acid sequences encoding them) is used, the replicator is preferably a linear double-stranded DNA comprising at least a Phi29 origin of replication at each extremity, or a derivative thereof (e.g. minimal origin), and one or more sequences encoding a polypeptide with appropriate regulatory sequences (e.g. a promoter, terminator and ribosome-binding site (RBS)), wherein each strand of the replicator is further capped in 5’ by phage Phi29 P3 protein or a phosphate group. Phosphorylated 5’ ends (i.e. capping in 5’ with a phosphate group) may be easily obtained, for instance using 5’phophate primers during PCR, or restriction of adapters, or enzymatic phosphorylation of 5’OH DNA.
In one embodiment, the replicator comprises only one sequence encoding a candidate polypeptide.
However, the inventors have shown that the nucleic acid molecule acting as replicator may contain multiple sequences encoding a polypeptide. For example, they started from a nucleic acid molecule comprising both a reporter gene (such as a gene encoding a fluorescent or a colored gene), inserted as a fusion with the gene encoding the polypeptide of interest. In this case, the colored or fluorescence signal associated with the reporter gene indicates the level of expression of the polypeptide of interest, which can be used to normalize the activity with respect to expression level during the sorting stage. In this case, the sorting process selects nucleic acid sequences encoding polypeptides with high specific activity rather than polypeptides with high apparent activity (where apparent activity is the product of specific activity and polypeptide concentration).
In another embodiment, the replicator comprises several sequences each encoding a distinct polypeptide. In this case, at least one of the sequences encodes a candidate polypeptide to be screened, while other sequence(s) may encode other candidate polypeptide(s) to be screened, a reporter protein (for instance a coloured or fluorescent protein) that may be used to normalize the expression level of the polypeptide(s) to be screened, or a substrate protein or a functional variant or fragment thereof (when the candidate polypeptide is an enzyme with a protein substrate). When several sequences (e.g. two) each encoding a distinct polypeptide are comprised in the replicator, the distinct sequences may be present as one or more fusions, as independent transcriptional units, or as multicistronic (e.g. bicistronic) constructs. A fusion is particularly adapted for normalization as it ensures that the expression level of the fused polypeptides is exactly the same.
Preferably, the candidate polypeptide encoded by the candidate sequence is selected from the group consisting of putative enzymes, putative transcription factors, putative functional derivatives thereof, and putative functional fragments thereof.
In one embodiment, the candidate sequence is preselected based on at least one of: sequence thereof, a portion of sequence thereof, structural properties, functional properties, and any combination thereof.
When the PURE system is used as transcription and translation machinery, the nucleic sequences encoding a candidate (or other additional) polypeptide are preferably under the control of a T7 promoter.
A replicator may further comprise additional subsequences, notably selected from a ribosome-binding site (RBS), barcodes, restriction enzyme binding sites, aptamers.
When the PURE system is used as transcription and translation machinery, the nucleic sequences encoding a candidate (or other additional) polypeptide preferably comprise an E. coli Ribosome Binding Site.
When a Phi29-based replication machinery and a PURE transcription and translation machinery are used, a particularly preferred replicator has 5’ -phosphate capped ends and comprises on one strand the following nucleic acid sequences from 5’ to 3’: oriL (preferably Phi29 oriL), T7 promoter, RBS (ribosome binding site, preferably E. coli RBS), gene encoding the candidate polypeptide, terminator, and oriR (preferably Phi29 oriR).
In addition, the Inventors have applied this novel method in combination with DNA-barcoding techniques by including randomized DNA subsequences called barcodes or Unique Molecular Identifiers (UMI) in the replicator. These barcodes or UMI can be used in combination with sequencing, for example to compute the changes in frequency of a given variant upon cycles of screening, to correct for amplification biases, to compute high quality consensuses from multiple reads, to assemble short reads into a longer complete sequence, or to check for the absence of recombination during the amplification steps or library preparation, or to “call out” one particular variant (for instance by performing a PCR with the barcode sequence as primer(s)). (Schwartz, J. J., Lee, C. & Shendure, J. Accurate gene synthesis with tag-directed retrieval of sequence-verified DNA molecules. Nature methods 9, 913-915 (2012)). All these techniques based on random DNA barcodes or UMI are well-known.
Preferably, the method is a high-throughput screening (HTS) method, more preferably an ultra-high-throughput screening (uHTS) method. In a particularly preferred embodiment, the method allows high or ultrahigh throughput screening of at least 105 nucleic acid molecules comprising at least one candidate sequence (preferably replicators as defined herein), preferably at least 106 nucleic acid molecules comprising at least one candidate sequence (preferably replicators as defined herein).
The second composition provided in step b) may further comprise other components useful in the method of the invention. For instance, when the method is applied to the screening of nucleic acid molecules encoding candidate enzymes, a substrate, a cofactor, a coenzyme (prosthetic group or cosubstrate) of the enzyme or any combination thereof may be added in the second composition provided in step b).
In step c), the first composition provided in step a) and the second composition provided in step b) are mixed. To avoid premature initiation of the replication or expression reactions, this step is preferably performed at cold temperature and/or just before step d).
Step d) - Partitioning the thus obtained mixture into microcompartments
In step d), the mixture obtained in step c) is partitioned into microcompartments, wherein at least one of said microcompartments comprises a single copy of said nucleic acid molecule.
In a preferred embodiment, a substantial fraction of said microcompartments comprises a single copy of said nucleic acid molecule. More preferably, essentially all non-empty microcompartments comprises a single copy of said nucleic acid molecule.
Several techniques for performing such partitioning are well-known in the art, resulting in distinct types of microcompartments. Preferably, the microcompartments are selected from the group consisting of microdroplets (such as water-in-oil microdroplets), microvesicles (such as liposomes), microchambers, and any combination thereof. The desired type(s) of microcompartments may be generated using standard techniques well-known to those skilled in the art. For instance, microdroplets may be obtained by microfluidic techniques or by bulk emulsification techniques. Advantageously, monodisperse water-in-oil microdroplets may be obtained using microfluidic system (Christopher, G. F. & Anna, S. L. Microfluidic methods for generating continuous droplet streams. Journal Of Physics D-Applied Physics 40, R319-R336 (2007)).
Advantageously, in the method of the invention, the average number of nucleic acid molecules per microcompartment (corresponding to the ratio between the sum of the individual nucleic acid molecule contents of the microcompartments and the total number of microcompartments) after partitioning is between 0 and 10, preferably between 0 and 5, yet preferably between 0 and 2, and most preferably between 0.1 and 1.
This can be obtained by using Poissonian partitioning, i.e by adjusting the concentration of the nucleic acid molecules comprising at least one candidate sequence encoding at least one candidate polypeptide of the library in the composition provided in step b) so that few microcompartments receive more than one molecule. In this setting, many microcompartments receive no nucleic acid molecule (they are referred to as empty microcompartments), and most non-empty microcompartments comprise a single nucleic acid molecule. For instance, with a concentration of the nucleic acid molecules (preferably replicators) of the library in the composition provided in step b) set at 0.63 pM and generation of droplets of 10 pm in diameter, the mean number of nucleic acid molecules (preferably replicators) per droplet would be 0.2. In such conditions, a Poissonian partitioning is expected, with 82% of droplets being empty, 16% containing one nucleic acid molecule (preferably a replicator), and approximatively 2% containing two or more nucleic acid molecules (preferably replicators). These conditions are appropriate for uHTS screening experiments, because most non-empty droplets (16/(16+2) = 88%) have received just one nucleic acid molecule (preferably a replicator) and are thus clonal in the sense that they will express many copies of the polypeptide originally encoded by a single nucleic acid variant (and also create many copies of the nucleic acid variant.
In a preferred embodiment, step d), the mixture obtained in step c) is partitioned into microcompartments, using Poissonian partitioning (i.e., poissonian distribution), preferably so that a substantial fraction of said microcompartments comprises a single copy of said nucleic acid molecule, more preferably wherein most of said microcompartments comprises a single copy of said nucleic acid molecule, even more preferably, essentially all non-empty microcompartments comprises a single copy of said nucleic acid molecule. Poissonian partitioning enables clonality, i.e., ensures that only one copy of nucleic acid molecule (i.e., a unique variant of the library) is incorporated in the majority of micro-compartments.
Depending on the type of compartments and production method, the partitioning of nucleic acid molecules may deviate from a Poissonian distribution (e.g., power-law distribution). In such case, experiments can be designed to obtain an empirical validation confirming that few microcompartments contain more than one nucleic acid variant molecule at the time of generation.
Steps e) f) and g) - Allowing said candidate sequence to be amplified by the nucleic acid replication machinery, optionally allowing the amplified candidate sequence of step e) to be transcribed by the nucleic acid transcription machinery, and allowing the amplified candidate sequence of step e) or the transcribed candidate sequence of step f) to be translated in at least one polypeptide by the nucleic acid translation machinery
In step e), optional step f) and step g), the microcompartments are incubated in conditions suitable for the replication machinery, the optional transcription machinery and the translation machinery to perform their function, leading to amplification of both the candidate nucleic sequence encoding the candidate polypeptide and of the candidate polypeptide itself.
Suitable conditions depend on the specific selected replication machinery, optional transcription machinery and translation machinery, and are known in the art.
For instance, when using bacteriophage Phi29 p2, p3, p5, and p6 proteins as replication machinery, conditions used in step e) are preferably between 20 °C and 50 °C, more preferably temperatures ranging from 25 to 37 °C during 1 hour to 16 hours in the presence of 20 mM ammonium sulfate.
When the PURE system is used as transcription-translation machinery in steps f) and g), conditions are preferably those recommended by the PURE system manufacturer. In particular, the reaction temperature can be below 40 °C, more preferably temperatures ranging from 25 to 37 °C during 1 hour to 16 hours. Step h) - Optionally adding at least one compound in the mixture of step c) or in the microcompartments of step g), and/or modifying at least one parameter within the microcompartments of step d) or g)
In step h), at least one compound may be added in the mixture of step c) or in the microcompartments of step g), and/or at least one parameter may be modified within the microcompartments of step d) or g).
For instance, when the candidate polypeptide is a candidate enzyme, at least one compound that may be added in the mixture of step c) or in the microcompartments of step g) may be selected from reagents, substrates, cofactors, coenzymes (including prosthetic groups and cosubstrates) of the candidate polypeptide encoded by the candidate sequence and any combination thereof.
In a particular embodiment, the activity of interest is polymer enzymatic degradation, and the at least one compound added the mixture of step c) or in the microcompartments of step g) is a fluorogenic (e.g. Fluorescein Diacetate substrate (FDA), Resorufin Di-B-D- Galactopyranoside (FDG), etc), chromogenic (e.g. 3, 3’-diaminobenzidine tetrahydrochlorure), solvatochromic (Nile Red) or environment-sensitive (e.g. Bromocresol Green) compound embedded in one or more polymer particles, wherein degradation of the polymer particles catalyzed by the candidate polypeptide leads to release and conversion of the compound into a fluorescent or colored product, or a change in spectral properties in the microcompartment. Sorting may then be done based on the amount of fluorescent or colored product detected in step i), depending whether it is higher or lower a user-defined threshold value.
The polymer may notably be a plastic polymer, in particular selected from PLA (polylactic acid), PET (polyethylene terephthalate), PE (polyethylene), PP (polypropylene), PCL (polycaprolactone), and PS (polystyrene).
When the candidate polypeptide is a candidate transcription factor, at least one compound that may be added in the mixture of step c) or in the microcompartments of step g) may be a nucleic acid molecule comprising a nucleic acid sequence encoding a reporter polypeptide (preferably a fluorescent or colored polypeptide) under the control of said transcription factor.
Alternatively or in combination, at least one parameter may be modified within the microcompartments of step d), in order to put the microcompartments in conditions optimal for replication in step e), optional transcription in step f) and translation in step g). Alternatively or in combination, at least one parameter may be modified within the microcompartments of step g), in order to put the microcompartments in conditions optimal for detecting the candidate polypeptide activity in next step i).
Examples of optimal conditions for steps e), f), g) and i) are disclosed in sections relating corresponding steps, and one or more parameters may thus be adapted in step h) in order to fulfill these optimal conditions.
In particular, at least one parameter that may be modified within the microcompartments of step d) or g) may be selected from concentration, temperature, pH, viscosity, and any combination thereof.
Step i) - Detecting the activity of the candidate polypeptide of step g) in the conditions that are optionally modified in step h) within the microcompartments
In step i), the activity of the candidate polypeptide of step g) is detected in the microcompartments.
Any suitable method known in the art may be used for the detection.
In a preferred embodiment, the activity of the candidate polypeptide encoded by the candidate sequence is determined in step i) by measuring the intensity of one or more signals. This measure can be made using one or more appropriate techniques selected from optical techniques, electrical techniques (such as capacitance, resistance, impedance measurements, and the like), and magnetic techniques.
The detection may be a direct detection assay (preferably a fluorogenic assay or a colorimetric assay), i.e. the measured activity directly generates a detectable optical and/or electrical and/or magnetic changes in the microcompartments. As an example, when the candidate polypeptide is a candidate enzyme, a substrate comprising a part with optical may be used, and the reaction of the candidate enzyme encoded by the candidate sequence on the substrate will generate detectable optical changes in the microcompartments if the candidate enzyme has enzymatic activity. Alternatively, a substrate coupled/fused to a compound with detectable optical properties may be used. More particularly, the activity of the candidate enzyme encoded by the candidate sequence may be determined by detecting a product obtainable from the substrate such as a product that either is converted by the candidate enzyme from the substrate, or is released from the reaction of the candidate enzyme with the substrate. A direct detection assay may also be used when the candidate polypeptide is a candidate reporter protein, such as a fluorescent or colored protein. Alternatively, an indirect activity assay may be used, preferably a fluorogenic assay or a colorimetric assay, wherein the activity of the polypeptide encoded by the candidate sequence is determined by detecting a product of the indirect activity assay. As an example, when the candidate polypeptide is a candidate transcription factor, a nucleic acid molecule comprising a nucleic acid sequence encoding a reporter polypeptide with optical and/or electrical and/or magnetic properties (preferably a fluorescent or colored polypeptide) under the control of said transcription factor may be present in the microcompartments, either due to their presence in the first or second composition provided in step a) or b), or by addition in step h). In this case, detection of the reporter polypeptide expressed only when the candidate transcription factor is active indirectly permits to detect the candidate transcription factor activity. As another example, a product released by the activity of the candidate enzyme may be used as a substrate by a second enzyme, which has been included in the reaction mixture. The reaction of this second enzyme with the product of the candidate enzyme, optionally in presence of a specific cosubstrate, generate a secondary product which may be easily detected for example via spectroscopic method. Many indirect enzyme detection assays, for example with fluorescent readouts, have been described and are well-known.
As another example, a product released by the activity of the candidate enzyme may bind to one (or more) macromolecules and increase its (their) fluorescence by irreversible photoactivation/photoconversion or reversible photoswitching, such as, but not restricted to, fluorescent biosensors proteins or aptamers (Sam Duwe, Peter Dedecker, Optimizing the fluorescent protein toolbox and its use, Current Opinion in Biotechnology, Volume 58, 2019, Pages 183-191 , ISSN 0958-1669; Koveal, D., Rosen, P.C., Meyer, D.J. et al. A high-throughput multiparameter screen for accelerated development and optimization of soluble genetically encoded fluorescent biosensors. Nat Commun 13, 2919 (2022)).
It is preferred to use at least one optical method in step i), which is more preferably selected from an absorption-based detection method (including single or multi-wavelength measurements, turbidometry, etc.), a fluorescence- based detection method (including fluorescence intensity, fluorescence lifetime, fluorescence wavelength, fluorescence quenching/unquenching, BRET, FRET, relaxation etc.), a scattering- based detection method (including Rayleigh or Raman Scattering, nephelometry, laser diffraction, static light scattering, dynamic light scattering, etc.), or a spectroscopy- based method (such as infrared spectroscopy, or Raman spectroscopy). More particularly, the optical method of step i) is preferably selected from the group consisting of absorption-based detection techniques (including single or multi-wavelength measurements, etc.), fluorescence-based detection techniques (including fluorescence intensity, fluorescence lifetime, fluorescence wavelength, FRET, etc.), and any combination thereof.
When the candidate polypeptide is a candidate reporter protein, such as a fluorescent or colored protein, its activity may be directly detected using fluorescence- based or absorption-based detection techniques (see Example 1 below).
When the candidate polypeptide is a candidate enzyme, fluorescence- based or absorptionbased detection techniques may still be used.
For instance, fluorescence- based detection techniques may be used when a fluorogenic substrate is used (see Example 2 below) or when the candidate enzyme is fused to a fluorescent polypeptide (i.e. the candidate nucleic acid molecule comprises at least one candidate sequence encoding a fusion of a candidate polypeptide and a fluorescent polypeptide, see Examples 3-5 below).
Similarly, absorption-based detection techniques may still be used when a chromogenic substrate is used or when the candidate enzyme is fused to a colored polypeptide (i.e. the candidate nucleic acid molecule comprises at least one candidate sequence encoding a fusion of a candidate polypeptide and a colored polypeptide).
For instance, fluorescence- based detection techniques may be also used to detect the product of an enzyme, for example using a fluorogenic probe (or a probe coupled to a fluorogenic compound) is used, wherein the probe is specific to the enzyme’s product (for instance see Example 8 below).
Similarly, absorption-based detection techniques may still be used when a chromogenic probe (probe specific to the product of the enzyme) is used or when the probe (specific to the product of the enzyme) is fused to a colored compound.
When the candidate polypeptide is a candidate regulatory protein such as a transcription factor, fluorescence- based or absorption-based detection techniques may still be used. For instance, a nucleic acid molecule comprising a nucleic acid sequence encoding a reporter polypeptide (preferably a fluorescent or colored polypeptide) under the control of said transcription factor may be present in the microcompartments, either due to their presence in the first or second composition provided in step a) or b), or by addition in step h).
The method of the invention can be a flow-based method (e.g., wherein detection in the microcompartments is sequential) or a non-flow-based method (e.g., wherein detection in the microcompartments is parallel). A flow-based method is preferably performed using a microfluidic device (such as a fluorescence activated droplet sorting (FADS) device (Baret, J.-C. et al. Fluorescence-activated droplet sorting (FADS): efficient microfluidic cell sorting based on enzymatic activity. Lab on a chip 9, 1850-1858 (2009)) or a fluorescence activated cell sorting (FACS) device. As a non-flow-based method, one can cite a method wherein liposomes are immobilized on a flat surface, each on a small electrode, whereby the activity of the protein to be detected changes the electrical property of the droplets, which triggers release of the vesicle in solution.
Step j) - Using a sorting apparatus, recovering the candidate sequence encoding a polypeptide having the activity of interest from microcompartments displaying specific activity signal
In step j), the candidate sequences encoding a polypeptide having the activity of interest (as detected in step i)) are recovered from microcompartments displaying specific activity signal, using a sorting apparatus.
The detecting device used in step i) for detecting the activity of the candidate polypeptides and the sorting apparatus used in step j) will generally be included in a single apparatus comprising several modules (it will then be referred to a detecting/sorting apparatus). In this case, the detecting/sorting apparatus combines a sensor module that detects the activity of interest in each microcompartment, and a sorting module that separates the contents of microcompartments depending on the level of activity measured by the sensor module, based on a user-defined threshold value. The sorting module comprises a software able to compare the level of activity measured by the sensor module with the user-defined threshold value and means for separating the contents of microcompartments depending whether the level of activity measured by the sensor module is above or below the user- defined threshold value.
The threshold value is preferably at least equal to and preferably higher than the level of activity measured for a control polypeptide. Depending on the desired stringency of the screening and the type of polypeptide screened, the threshold value may for instance be at least 1.2, at least 1.3, at least 1.4, at least 1.5, at least 1.6, at least 1.7, at least 1.8, at least 1.9, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 200, at least 300, at least 400, at least 500, at least 103, at least 104, at least 105 or at least 106 times higher than the level of activity measured for a control polypeptide. In the context of directed evolution, the threshold value is preferably at least equal to and preferably higher than (any value mentioned above) the level of activity measured for the original polypeptides from which the candidate polypeptides are derived.
In the context of bioprospection, the threshold value is preferably at least equal to and preferably higher than (any value mentioned above) the level of activity measured for a known polypeptide, more preferably it is higher than the level of activity measured for the polypeptide with the best known level of activity.
Any appropriate sorting apparatus may be used, depending on the particular type of microcompartment used.
For instance, when microdroplets (such as water-in-oil microdroplets) are used as microcompartments, a fluorescence activated droplet sorting (FADS) device may preferably be used. When liposomes are used as microcompartments, a fluorescence activated cell sorting (FACS) device may preferably be used.
Such sorting apparatus are commercially available and should be used as recommended by the manufacturer.
Preferably, the sorting apparatus will pool the contents of all microcompartments for which the level of activity measured by the sensor module is higher than the user-defined threshold value. In this case, the nucleic acid molecules of the pool may be analyzed or, in the context of directed evolution, used as starting library for the next directed evolution cycle.
Uses of the method of screening according to the invention
Two particularly interesting uses of the method of screening according to the invention are directed evolution and bioprospection (also referred to as functional metagenomics).
In the case of directed evolution, the initial library (comprised in the second composition that after mixing with the first composition will be partitioned through all microcompartments) is typically obtained from one or a few candidate sequences (for example from natural or engineered sequences encoding proteins with known properties), which are further diversified with mutagenesis (obtained notably by shuffling, random mutations, targeted mutations, combinatorial mutations, or any combination thereof).
The invention thus also relates to the use of the method of screening according to the invention in a directed evolution method, said directed evolution method preferably comprising multiple cycles of mutagenesis (obtained notably by shuffling, random mutations, targeted mutations, combinatorial mutations, or any combination thereof) followed by screening using the method of screening according to the invention.
For instance, the invention also relates to a method of directed evolution, comprising: a) generating a library of nucleic acid molecules encoding each a candidate polypeptide by mutagenesis (obtained notably by shuffling, random mutations, targeted mutations, combinatorial mutations, or any combination thereof) of one or more nucleic acid molecule(s) encoding each a known polypeptide with an activity of interest, b) screening the library generated in step a) for one or more nucleic acid molecule(s) encoding a polypeptide with the activity of interest using the method of screening according to the invention; c) generating a library of nucleic acid molecules encoding each a candidate polypeptide by mutagenesis (obtained notably by shuffling, random mutations, targeted mutations, combinatorial mutations, or any combination thereof) of the one or more nucleic acid molecule(s) selected in step b); d) screening the library generated in step c) for one or more nucleic acid molecule(s) encoding a polypeptide with the activity of interest using the method of screening according to the invention; and e) repeating steps c) and d) until one or more nucleic acid molecule(s) encoding a polypeptide with the desired activity of interest is obtained.
In the particular case of directed evolution of enzymes able to degrade polymers, the directed evolution method may start from one or more nucleic acid molecule(s) encoding each a known enzyme with the activity of interest. Such enzymes are known in the art. For instance, known enzymes with the ability to degrade PLA (polylactic acid) include cutinases (such as Humicola insolens cutinase (HiC)). Known enzymes with the ability to degrade PET (polyethylene terephthalate) include some proteases and cutinases or IsPETase from Ideonella sakaiensis. Known enzymes with the ability to degrade PE (polyethylene) and PP (polypropylene) include cytochrome P450. Known enzymes with the ability to degrade PCL (polycaprolactone) include Humicola insolens cutinase (HiC) and proteases. Known enzymes with the ability to degrade PS (polystyrene) include hydroquinone peroxidase HM121 from Azotobacter beijerinckii.
In functional metagenomics applications, the initial library is typically a set of natural sequences obtained from an environment where it is likely that genes encoding polypeptides with the desired properties exist. For example, if biomass-degrading enzyme are sought, it may be interesting to use a metagenomic library coming from ruminant gut.
The invention thus also relates to the use of the method of screening according to the invention in a functional metagenomics method.
The invention also relates to a functional metagenomics method comprising: a) generating a library of natural nucleic acid molecules encoding each a candidate polypeptide obtained from an environment where it is likely that genes encoding polypeptides with the desired activity are present; and b) screening the library generated in step a) for one or more nucleic acid molecule(s) encoding a polypeptide with the activity of interest using the method of screening according to the invention.
The method of screening according to the invention may also be used in a method starting by a first step of functional metagenomics and using the selected nucleic acid molecules in a second step of directed evolution. In this case, the method of screening according to the invention may be used for the screening of the functional metagenomics step, for the screening of the directed evolution step, or for the screening of both steps.
Kit
The present invention also relates to a kit, preferably a kit for implementing the method described above, comprising, or consisting essentially of, or consisting of: a) an in vitro system /platform comprising:
(i) a nucleic acid replication machinery,
(ii) optionally a nucleic acid transcription machinery, and
(iii) a nucleic acid translation machinery; b) at least one material for making microcompartments,; and c) optionally instructions of use.
In this kit according to the invention, the in vitro platform may be selected from any general or preferred embodiment disclosed above with respect to the first composition provided in step a) of the method according to the invention.
In particular, a preferred in vitro platform comprised in the kit according to the invention comprises both a bacteriophage Phi29-based replication machinery (as described herein) and a PURE transcription and translation system (as described herein). The kit according to the invention further comprises at least one material for making microcompartments.
The material for making microcompartments may be selected from reagents for making microcompartments and devices for making microcompartments.
As reagent for making microcompartments, the kit according to the invention may comprise at least one partitioning agent, which may notably be selected from suitable continuous phases and surfactants for generating emulsions of microdroplets.
As device for making microcompartments, the kit according to the invention may comprise a microfluidic microarray.
Both a reagent for making microcompartments and a device for making microcompartments may be included in the kit.
DESCRIPTION OF THE FIGURES
Figure 1. Demonstration of IVTTR reaction using PUREfrex 1.0 and the reconstituted replication machinery of phage Phi29 (P2 P3 P5 P6). End-point GFP fluorescence of 10 pL of IVTTR (grey bar) and qPCR quantification of final <gfp> concentration (black stroke) after 5 h incubation at 30 ° C. A, B and C, D: Replication is necessary for high protein production starting from low amounts of template (3 pM or 12 pM of <gfp> gene, respectively). E, F, G: high replication and protein production is possible with purified P2 and P3 proteins, with a balance between protein production and replication.
Figure 2. End-point GFP fluorescence of 10 pL of IVTTR 1.0 (grey bar) and qPCR quantification of final <gfp> concentration (black stroke) after 5 h IVTTR at 30 °C. A-D: increasing amount of pUC57_OriLR_p2p3 plasmid have non monotonous effect on GFP yield, with a maximum at 200 pM plasmid (B). the effect on final <gfp> concentration is positive and saturates above 300 pM plasmid concentration (C-G). The optimal pUC57_OriLR_p2p3 concentration is around 200 pM.
Figure 3. 10X magnified brightfield (A&C) and green fluorescence (B&D) images of IVTTR 1.0 experiment initiated with <gfp>replicator at a concentration of less than one copy per compartment. Condition supplemented with dNTP (A&C) shows a discrete number of fluorescent droplets, indicating replication and expression in droplet for droplet having received one or more copy of <gfp>, whereas condition without dNTP (C&D) does not show any fluorescence. Figure 4. End-point mCherry fluorescence for 5 pL of IVTTR 2.0 containing <rfp> replicator and different amounts of PLA-FDL microparticles after 8 h incubation at 33 °C. Adjacent bars represent the result of two replicates. Little difference can be seen in mCherry yield, proving a low toxicity of PLA-FDL microparticles on IVTTR 2.0.
Figure 5. Steepest slope of green fluorescence emergence for 5 pL of IVTTR 2.0 containing either <hic > (A-D) or <rfp> replicator (E-H) and different amounts of PLA-FDL microparticles after 8 h incubation at 33 °C. Adjacent bars represent the result of two replicates. <rfp> replicator shows low fluorescein hydrolysis whereas <hic> replicator shows increasing fluorescein hydrolysis with substrate concentration.
Figure 6. 4X magnified brightfield (A&D&G) and green fluorescence (B&E&H) and red fluorescence (C&F&I) images of IVTTR 2.0 (A-l) or PUREfrex 2.0 only (G-l) experiment initiated with <rfp-hic> replicator at 1 pM in 40 pm droplets. Condition supplemented with dNTP (A-C) shows a discrete number of fluorescent droplets, indicating replication and expression in droplet for droplet having received one or more copy of <rfp-hic>. Colocalization of green and red fluorescence indicates that cutinase activity is related to protein production. Condition without dNTP (D-F) does not show any fluorescence. Condition with PUREfrex shows a higher background of green fluorescence (H) than IVTTR 2.0 without dNTP (E) but not comparable with IVTTR 2.0 with fluorescence.
Figure 7. 4X magnified green (A), red (B), far-red (C) fluorescence, corresponding respectively to Hie esterase activity, mCherry fluorescence, and a fluorescent dye homogeneously distributed in the initial master mix for droplet tracking, and brightfield (D) images of IVTTR 2.0 mock library experiment after overnight incubation and before sorting. Experiment initiated with a mixture of <hic> and <rfp> replicators at a molar ratio of 10:90 in 47 pm droplets. A majority of non -fluorescent droplets in green and red channel indicates a Poisson distribution with a parameter A << 1. The independence of green and red fluorescence confirms the Poisson distribution and implies clonality of DNA after encapsulation.
Figure 8. Scatter plot of screened droplets according to far-red fluorescence (x-axis) and green fluorescence (y-axis). Horizontal line represents green fluorescence threshold above which droplets are sorted (positive hints). Figure 9. qPCR quantification of hie percentage in gene libraries during mock library screening experiment. Standard deviations are calculated from technical triplicate qPCR quantification.
Figure 10. A. Microscopic image of polyesterase activity (top: red fluorescence; bottom: green fluorescence) in microdroplets containing the wild-type HiC replicator in fusion with mCherry, via IVTTR reaction performed at less than one copy per compartment. B. Histogram of green fluorescence values (polyesterase activity) in non-empty droplets. C. Scatter plot of green (polyesterase activity) versus red (protein concentration) for each droplet, showing a strong correlation. D. Histogram of the ratio of fluorescence signals for all non-empty droplets, showing that a sharp peak for specific activity was obtained by normalizing enzymatic activity with the red expression signal.
Figure 11. A) Flow cytometry analysis of YFP-expressing liposomes in a mock enrichment experiment. The initial library contained a 1 :10 ratio of the yfp and minD genes, except for the 1 nM <yfp> control sample. Starting from 1 pM <yfp>, IVTTR leads to higher levels of expressed YFP compared to IVTT samples. B) Gel electrophoresis analysis of DNA recovered by PCR from FACS-sorted liposomes. Full-length DNA was successfully recovered from IVTTR, but not from IVTT samples, using purified or expressed DNAP and TP proteins. C) Quantitative PCR data showing that both yfp and minD genes are amplified in IVTTR samples, enabling recovery of DNA. D) Fraction of yfp and minD genes in the library, before and after FACS sorting using IVTTR in liposomes, as measured by qPCR. Two gating stringencies were used (sort 1 and sort 2) and two IVTTR configurations (with purified or expressed P2 and P3 from plasmid pUC57_OriLR_p2p3) were tested. The results show that liposomes sorted based on their high YFP fluorescence contain an excess amount of the yfp gene, up to 90%. The results demonstrate that a functional gene variant can be enriched in liposomes after IVTTR when starting from picomolar amounts of DNA.
Figure 12. IVTTR for a non-catalytic and non-fluorescent protein. (A) Schematic of inliposomes ori-p3 DNA amplification and expression via IVTTR. (B) Absolute quantitation of ori-p3 DNA by qPCR in lysed liposomes. Total 10 pM input template DNA concentration was used, which was reduced due to externally supplied DNase I. (C) Populational variation of dsGreen fluorescence in IVTTR-liposomes measured by flow cytometry. (D) Quantitation of the fraction of liposomes showing above-background dsGreen fluorescence estimated by the diagonal gate in C. Figure 13. Application of IVTTR for improving lipid synthesis by a membrane-associated enzyme. (A) Schematic of CDP-DAG conversion to PS by PssA. (B) Schematic of IVTTR- liposomes expressing PssA enzyme and detection of PS-positive liposomes by LactC2-mCherry binding. (C) Absolute DNA quantitation by qPCR of lyzed CADGE liposome samples. ** P < 0.01 ; *** P < 0.001. (D) Box plots and (E) quantitation of PS-positive CADGE liposomes expressing PssA, as assayed by flow cytometry.
Figure 14. GFP fluorescence (grey bar) and final replicator concentration (black stroke) in IVTTR 2.0 with <gfp> replicator and different dNTP compositions in presence (A, C, E) or absence (B, D, F) of DNK gene. A and B conditions have equal concentration dATP, dCTP, dGTP and dTTP. C and D conditions have dTTP replaced by an equal concentration of dTMP, E and F conditions have no dNTP.
Figure 15. Brightfield (left) and fluorescent (right) images showing the detection of DNK activity from single copy of encapsulated DNA fragment containing the DNK gene and a fluorescent protein.
EXAMPLES
Although the present invention herein has been described with reference to particular embodiments, it is to be understood that these embodiments are merely illustrative of the principles and applications of the present invention. It is therefore to be understood that numerous modifications may be made to the illustrative embodiments and that other arrangements may be devised without departing from the spirit and scope of the present invention as defined by the appended claims.
1. EXAMPLE 1 : In vitro coupled replication-expression reactions provide high expression levels even when starting from picomolar gene concentrations
The Inventors first showed that a coupled replication-expression in vitro reaction can produce high level of protein expression even when starting from very low concentration of the encoding gene, whereas a simple in vitro expression reaction cannot, when starting from the same low concentration. Two different ways to introduce the replication machinery were tested: either with purified P2 and P3 proteins, or by in situ production of P2 and P3 from their encoding genes on the plasmid pUC57_OriLR_p2p3. In situ production takes advantage of the protein expression machinery in the mixture to express some of the proteins of the replication machinery from their gene.
A GFP encoding replicator was assembled by combining in order the following nucleic acid sequences: Phi29 oriL, T7 promoter, an E. coli ribosome binding site, the gfp gene, a terminator, and Phi29 oriR. The DNA was capped with 5’ -phosphate introduced by PCR using 5’-phosphate primers. This molecule is indicated as <gfp> (in the following, brackets <...> indicate DNA molecules that are competent for replication via the for reconstituted Phi29 machinery).
A typical coupled replication-expression reaction was performed with PURE/rex 1.0 (GeneFrontier) and a starting replicator DNA concentration of either 3 pM or 12 pM in 10 pL reaction volume. One picomolar concentration of replicator corresponds to a concentration close to one molecule per micro-compartment, in the case where the microcompartment has an inner volume around 1.5 picoliter (which for example is the volume of droplets of 15- 20 pm in diameter, a size typically used in droplet screening applications). Concentration of pUC57_OriLR_p2p3 was fixed at 200 pM in in situ protein production condition. In purified protein condition, P2 protein (NEB) was fixed at 1 U/pL while P3 (homemade) varied between 0.8 mg/mL and 3.2 mg/mL.
Green fluorescence was monitored during 5 h at 30 °C in a standard 96-well plate fluorescence spectrometer (here we use Biorad cfx thermocycler running at constant temperature). Values of final green fluorescence (mature GFP production, grey bars) and final <gfp> concentration (black strokes) are reported in Figure 1. In the absence of dNTP in the mix, <gfp> replication is impossible, and GFP fluorescence is not detectable (A&C). In presence of dNTP, GFP fluorescence is easily detected, and we measure that <gfp> concentration increase by orders of magnitude. This is true using either in situ produced P2 and P3, or purified P2 and P3 (B&D and E&F&G respectively). In the case of purified P3 (F), the existence of an optimal concentration, with a maximal protein production at 1.6 mg/mL P3, suggests a competition effect between replication and protein production.
The optimal concentration of plasmid pUC57_OriLR_p2p3 for GFP production in an IVTTR system was then explored. As the in situ synthesis of P2 and P3 also uses the PURE machinery, it is expected that a too high concentration will negatively affect the final yield of the protein of interest (GFP). A typical IVTTR reaction was performed with PURE/rex 1.0 (GeneFrontier) and a starting GFP-encoding replicator DNA concentration of 12 pM in 10 pL reaction volume. Concentration of pUC57_OriLR_p2p3 was varied between 100 pM and 700 pM by 100 pM steps. Green fluorescence was monitored during 5 h at 30 °C in CFX. Values of final green fluorescence (grey bars) and final <gfp> concentrations (black strokes) are reported in Figure 2. Figure 2 clearly demonstrates an increase followed by a decrease in GFP production with increased pUC57_OriLR_p2p3 concentration, with a maximum fluorescence at 200 pM pUC57_OriLR_p2p3. Replication yield is also increased with increasing pUC57_OriLR_p2p3 concentration but stays high even at the highest concentration (700 pM).
Using these optimal concentrations for IVTTR system, detection of GFP fluorescence in microliter droplets was verified, where the reaction initiated from a unique copy of the DNA replicator. A typical IVTTR reaction was performed with PURE/rex 1.0 (GeneFrontier) and a pUC57_OriLR_p2p3 concentration fixed at 200 pM. Green fluorescence was monitored during 5 h at 30 °C in CFX. <gfp> replicator was set at 0.63 pM and droplets of 10 pm in diameter were generated, so that the mean number of replicator per droplet was 0.2. In such conditions a Poissonian partitioning is expected, with 82 % of droplets are empty, 16 % contain one replicator, and approximatively 2 % contain two or more replicators. These conditions are appropriate for uHTS screening experiments, because most non-empty droplets (16/(16+2)= 88%) have received just one replicator and are thus clonal. In this experiment, Yeast total RNA (Sigma Aldrich, ref. 10109223001 ) and Pluronic F127 (Sigma Aldrich, ref. P2443-250G) were supplemented at a final concentration of 1 ng/pL and 0.4 % (w/v) respectively. The mastermix was split in two: one half was supplemented with 0.3 mM dNTP and the other half with MilliQ water only. Flow focusing device tailored to produce 10 pm diameter droplet was used with fluorinated oil (HFE7500) supplemented with 64 mg/mL (4%) Fluosurf (Emulseo). Emulsions were incubated in PCR machine for 4h at 30° C and then kept at 4°C. Emulsions were spread between hydrophobized slide and coverslip, sealed with epoxy glue to prevent evaporation. Images were taken on epifluorescence microscope (Nikon) with GFP filter at 10X magnification, using the same parameters for both samples. Fluorescence images were inversed for easier interpretation. Brightfield and fluorescence images are gathered Figure 3. The figure shows that only a portion of the droplets exhibit GFP fluorescence in condition with dNTP (C), whereas in condition without dNTP no fluorescence is visible (D). This proves that the replication machinery of the IVTTR system is necessary to produce detectable protein level starting from single DNA templates in microdroplets, while the simple IVTT system is not able to do that. 2. EXAMPLE 2: One-pot in vitro coupled replication-expression and fluorescence monitoring of the activity of a protein with enzymatic activity, starting from pM amount of its gene in replicator form.
Humicola insolens cutinase (HiC) is an enzyme from the esterase family, more precisely a cutinase, with hydrolysis activity toward various esters or polyesters. It is known to degrade cutin and polyesters. Here we use as a model substrate the solid polyester poly-lactic acid (PLA). To assess the capacity of the IVTTR system to replicate and express an enzyme up to detectable activity, and the possibility to detect the activity in situ via a fluorogenic substrate, the fluorogenic compound fluorescein di-laurate (FDL) was used, embedded in PLA microparticles. PLA degradation by HiC cutinase leads to the release of FDL in solution, which is followed by the hydrolysis of FDL into fluorescein, also catalyzed by the esterase activity of HiC. Therefore, the apparition of green fluorescence indicates polyesterase activity and indirectly the presence of HiC. Fluorescein having absorption and emission spectra close to GFP, we used mCherry protein as reporter and negative control. mCherry replicator is noted <rfp> for “red fluorescent protein”, while HiC replicator is noted <hic>. A typical IVTTR reaction was performed with PUREfrex 2.0 (GeneFrontier) and a starting replicator DNA concentration of 10 pM. Reaction volume was set at 5 L. Concentration of pUC57_OriLR_p2p3 was fixed at 200 pM. Green and red fluorescence were monitored during 8 h at 33 °C in CFX. IVTTR experiment was carried with either <hic> or <rfp> replicator. Plastic microparticles were supplemented to a final concentration of 4,3 mM, 2.1 mM, 1.1 mM or 0 mM. Each plastic concentration pipetting was performed in duplicate. Red fluorescence produced by <rfp> replicator was used as a control for plastic toxicity to the replication and expression reaction. Baseline corrected red fluorescence of <rfp> samples are shown in Figure 4. Figure 4 shows that final mCherry fluorescence does not diminishes with increased amount of PLA-FDL microparticles and thus, there is no significant toxicity of the substrate toward IVTTR.
Green fluorescence increases strongly in tubes performing <hic> replication and expression, showing that it is possible to detect the activity of the enzyme encoded in the replicator, that is, to combine replication, expression and enzymatic assay in one-pot condition. The fluorescence still increases after 8 h of incubation, and samples with higher substrate concentration saturate the CFX sensors, so we assess plastic degrading activity by Vmax measurement. The steepest slope in green fluorescence time trace is shown on Figure 5. Figure 5 shows that <rfp> replicator shows low fluorescein hydrolysis whereas <hic> replicator shows increasing fluorescein hydrolysis with substrate concentration. This proves that hydrolysis is performed by HiC cutinase synthesized in situ from the replicator. Figure 5 also shows that that Vmax is approximately doubling when PLA-FDL concentration is doubling, consistent with the HiC enzyme being unsaturated.
3. EXAMPLE 3: IVTTR reactions, but not simple IVTT reactions, yield detectable amounts of enzymatic activity in microcompartments starting from a single copy of the enzyme’s gene
This example shows that a single <hic> gene molecule in an IVTTR microcompartment containing a fluorogenic substrate is able to produce an observable amount of fluorescence, but that it is not the case for IVTT conditions (i.e. in absence of genetic replication). Water- in-oil monodisperse microdroplets were used, with an inner volume around the picoliter, for the micro-compartmentalization step. In other words, this example shows that genetic replication is required to be able to perform uHTS screening in microcompartments directly from a library of linear genes.
For this demonstration, the fact that the IVTTR system can amplify and express longer DNA constructs (for example containing more than one gene) was leveraged. Here, DNA constructs were designed to encode two proteins, mCherry and HiC cutinase, expressed as a single fusion peptide. The link between the two proteins includes a thrombine cleaving site LVPRGS flanked by flexible peptide sequences GGGGS, but many other linkers would be suitable. The colocalization between protein production (red fluorescence) and cutinase activity (green fluorescence) will prove that cutinase activity is dependent on <rfp-hic> presence.
A typical IVTTR reaction was performed with PURE/rex 2.0 (GeneFrontier) and a starting replicator DNA concentration of 1 pM in 40 micrometer diameter droplets. Concentration of pUC57_OriLR_p2p3 was fixed at 200 pM. Droplet. The mastermix was split in two: one half was supplemented with 0.3 mM dNTP and the other half with MilliQ water only. A third sample was produced containing only the three solutions from PUREfrex 2.0 kit and the <rfp- hic> replicator at 1 pM concentration. Flow focusing device tailored to produce 42 pm (39 pL) droplet was used with fluorinated oil (HFE7500) supplemented with 64 mg/mL (4%) Fluosurf (Emulseo). Emulsions were incubated in PCR machine for 16h at 33 °C and then kept at room temperature. Emulsions were spread between hydrophobized slide and coverslip, sealed with epoxy glue to prevent evaporation. Images were taken on epifluorescence microscope (Nikon) with GFP filter at 4X magnification, using the same parameters for both samples. Fluorescence images were inversed for easier interpretation. Brightfield (A, D, G), green (B, E, H) and red (C, F, I) fluorescence images are gathered in Figure 6. Figure 6 shows that in condition with dNTP (A-C) some droplets, but not all, exhibit green and red fluorescence, denoting a sub-Poissonian distribution (i.e. an initial distribution of the replicator DNA molecules in the droplets such that some droplet did not receive any replicator molecule, some got only one and a few more than one, according to the Poisson distribution), and the fact that the reaction initiates from single DNA molecules. Most importantly, one can observe that green and red fluorescence are colocalized, meaning that cutinase activity is correlated with mCherry protein, proving that fusion protein is present in the droplet. Condition without dNTP with PURE 2.0 (D-F) does not show any fluorescence, proving that replication is necessary in an IVTTR process to observe high protein production and enzymatic activity. Condition without dNTP with PUREfrex (G-l) shows a slightly higher background of green fluorescence (H) than IVTTR with PURE 2.0 without dNTP (E). However, this background appears in all droplets, and the level is much lower than positive droplets in IVTTR with PURE 2.0 condition, thus it can be interpreted as a higher background hydrolysis of the polyester particles in PUREfrex conditions.
This proves that the replication machinery of the IVTTR system is necessary to produce detectable enzymatic activity starting from single DNA templates in microdroplets, while the simple IVTT system is not able to do that.
4. EXAMPLE 4: Clonal expression in microdroplets, followed by sorting and enrichment of genes encoding active enzymes using IVTTR of single gene molecules in microdroplets containing a fluorogenic substrate
We then tested the possibility to obtain clonal expression in microcompartments, when starting from a mixture of different genes in replicator form, i.e. a library. First, we prepared a mock library consisting of a mixture of two DNA sequences, one encoding an active polyesterase (cutinase HiC) and one encoding a catalytically inactive protein (the fluorescent protein mCherry). Both sequences were inserted between the replication origins and the expression elements (promoter, terminator and RBS) as explained above, and thus were able to replicate and produce high level of protein expression. The starting library contained 10% <hic> and 90% <rfp> replicators. The reaction mixture was prepared using PUREfrex 2.0 kit and fluorogenic particles to detect esterase activity. The reaction was encapsulated in 47-pm droplets (54 pL volume) in surfactant-containing fluorinated oil. Far- red fluorescent dye was added to the solution to allow detection of all droplets.
After overnight incubation, a drop of emulsion was spread between hydrophobized slide and coverslip, sealed with epoxy glue to prevent evaporation. Images were taken on epifluorescence microscope (Nikon) with GFP filter at 4X magnification. Fluorescence images were inversed for easier interpretation. Green (A), red (B) and far-red fluorescence (C) and Brightfield (D) images are gathered in Figure 7. A majority of non-fluorescent droplets in green and red channel indicates a Poisson distribution with a parameter A « 1 (A corresponds to the average number of replicators per droplet). In addition, we observed a strong green signal in some droplets and a strong red fluorescent signal in other droplets, indicating a high level pf protein expression. As expected, green and red image do not overlap. The statistical independence of green and red fluorescence confirms the Poisson distribution and confirms clonal expression of the genes encoded in the replicators of the library.
We then tested to possibility to submit this emulsion to a microfluidic droplet sorting method, in order to recover genes encoding the active enzyme (and deplete this fraction from genes encoding mCherry, which is catalytically inactive). The droplets were sorted using a microfluidic droplet sorting device (a.k.a. FADS). The sorting gate was set according to green fluorescence in order to send to the sorting gate only the droplets displaying significant green fluorescence (indicative of Hie esterase activity). Figure 8 shows the scatter plot of green droplet fluorescence (y-axis) versus far-red droplet fluorescence (x- axis). Horizontal line represents the green fluorescence threshold, above which droplets are sorted. 900 positive droplets were successfully sorted (“sorted” population) against 1 X106 negative droplets (“waste” population).
Finally, in order to evaluate the sorting efficiency, we used a PCR assay to measure the relative proportion of Hie and mCherry genes in the mock library, sorted bin and waste bin. Sorted droplets, waste droplets and non-screened (mock) droplets were separately extracted in ultrapure water using 1 H,1 H,2H,2H-Perfluoro-1 -octanol to break the emulsion. The amounts of mCherry and HiC coding sequences were quantified in each sample using genespecific qPCR primers. To ensure a high fidelity of HiC over mCherry ratio measurement, a fusion of mCherry and HiC coding sequences was used as a unique standard. Figure 9 summarizes qPCR data of the ratio of the hiC gene in each sample. Starting HiC proportion did not change between mock library DNA mixture and droplets before screening, at around 10%. HiC proportion in the sorted population increased up to 54%, while it decreased in waste population down to 4%. These results confirm the high level of clonal expression illustrate the capacity of the method to discriminate between active and inactive variants, and the possibility to enrich a DNA population with active variants, directly from encapsulation of linear DNA constructs. 5. EXAMPLE 5: Quantification of specific activity level in single-DNA microdroplet-IVTTR assays starting using hic-rfp fusion constructs
The examples above demonstrate that the IVTTR system can amplify and express longer DNA constructs, for example containing more than one gene, and that this can be used to significantly improve the determination of enzyme activity in microdroplets. Here, a rfp-hic fusion gene (where rfp is the fluorescent protein mCherry and hie is the enzyme) was used to measure specific enzymatic activity, instead of apparent enzymatic activity, as in the previous examples. When IVTTR was implemented in droplets from a single DNA template, two fluorescence signals could be observed, one corresponding to the enzymatic activity and one (the red signal) indicating the expression level of the construct. This allowed us to measure in droplets the specific activity. Indeed, the overall enzymatic activity can be calculated as the product of the specific activity and the enzyme concentration. Since the concentration of expressed protein varies between micro-compartments due to stochastic effects in the encapsulation and dynamics of the expression machinery, we observed a distribution of activity values, even when all the droplets contained the same wild-type gene (Figure 10). In addition, the expression level can vary from variant to variant, blurring possible differences in specific activity.
To assess the specific activity from the measurement of the enzymatic activity (A) in a sample, the concentration of enzyme present in the sample (C) was evaluated and the ratio A/C was computed. To measure the enzyme concentration directly in microdroplets, a fluorescent tag was attached as an N terminal fusion. Droplet-to-droplet heterogeneity can thereby be reduced by normalizing activity by protein concentration. A construct was created by the genetic fusion of mCherry fluorescent protein and HiC cutinase, connected by a flexible and cleavable linker. Both proteins being covalently coupled, each cutinase enzyme carries an mCherry tag, which allowed us to measure the cutinase concentration via the red fluorescence level.
An esterase activity assay was carried out in droplets with fluorogenic particles and the fusion <rfp-hic> as the replicating DNA. IVTTR mixture was assembled using PURE 2.0 kit and fluorogenic microparticles. IVTTR solution was encapsulated in 42-pm (39 pL) droplets in surfactant-containing fluorinated oil. The emulsion was incubated at 33 °C overnight. The day after, the emulsion was imaged on a microscope slide with appropriate filters for mCherry and fluorescein detection. A control experiment was performed in the same conditions using a mock library of <hic> and <rfp> replicator DNA. Microscope slides were incubated and imaged three days later to confirm that the green fluorescence was still increasing and thus the reaction was not terminated during the first measurement. Image analysis can extract two signals from each microdroplet, one in the red channel related to the concentration of the expressed fusion protein (Figure 10A bottom), and one in the green channel reporting the polyesterase activity (Figure 10A top). The result in the green channel indicates that, although all active droplets contain the same gene, the level of activity is not identical and displays a large distribution (Figure 10B). This heterogeneity is attributed to differences in the enzyme expression level in each compartment. However, it can be seen that the green signal is strongly correlated to the red one, indicating that a large part of the variability in activity can be explained by droplet-to-droplet differences in enzyme concentration (Figure 10C). When computing the specific activity (i.e. , the green to red ratio), a narrow peak histogram was obtained, which was expected since all droplets express the same wildtype gene (Figure 10D).
6. EXAMPLE 6: In-liposome IVTTR and enrichment of a gene encoding for a fluorescent protein using FACS
The possibility of sorting active variants after direct encapsulation of linear constructs was tested in another enrichment experiment using liposomes as an alternative compartmentalization strategy. In this experiment, we want to enrich the DNA template <yfp> (yellow fluorescent protein) from an excess of the unrelated template <minD> encoding a non -fluorescent protein, based on the fluorescence of expressed YFP, using a commercial fluorescence activated cell sorting apparatus (FACS). The <yfp> DNA template was mixed with 10-fold excess of <minD> template and the IVTTR mixture was encapsulated in liposomes at a total 10 pM DNA concentration (in this case, A = 0.2). At such a low template DNA concentration, expression of YFP is very low compared to higher DNA concentrations typically used in cell-free reactions, leading to a low signal-to-noise ratio. In contrast, liposomes with IVTTR conditions exhibited a higher level of YFP fluorescence (Figure 11A). Two stringency conditions were tested for the sorting gate: the “gate 1 ”, which encompassed the top 1% of all the liposomes (applied towards both IVTT and IVTTR samples), and the “gate 2”, which included only the top (0.2%) of the high-intensity liposomes (applied to IVTTR samples only). It was reproducibly difficult to recover the full-length DNA by PCR from the non-amplified liposome samples, while full-length DNA from liposomes with implemented IVTTR was easily recovered (Figure 11 B). This finding can be explained by higher DNA titers in the sorted liposomes from IVTTR samples. Indeed, as assayed by qPCR, <yfp>/<minD> mixtures in IVTTR liposomes were comparably (more than a 100-fold) and uniformly amplified (Figure 11C). Furthermore, qPCR quantification of sorted liposome samples suggests that using the more stringent condition of “gate 2” in IVTTR samples results in improved purity of YFP sorting compared to “gate 1 ” in both IVTTR configurations (Figure 11 D). These findings show that IVTTR enables efficient enrichment and DNA recovery of gene encoding for a functional protein, in a single round of FACS sorting of liposomes.
7. EXAMPLE 7: Compartmentalized IVTTR for a non-catalytic and non-fluorescent protein (in liposomes)
This example shows that partitioned IVTTR from single genes can be applied to non-catalytic and non-fluorescent proteins, by making use of an indirect fluorescent readout method. Here, we exemplify this with the Terminal Protein (TP) of phage phi29, which is encoded by p3 gene. The p3 gene was inserted between the replication origins downstream a T7 promotor. This template was introduced at 10 pM concentration in the IVTTR mix, supplemented with an excess amount of plasmid encoding the DNA polymerase, and encapsulated in liposomes at Poisson parameter 2 = 0.2 (Figure 12A). After incubation, quantitative PCR showed that the p3 gene was amplified inside liposomes by three orders of magnitude in the presence of dNTPs compared to the -dNTPs control (Figure 12B). The DNA intercalating dye dsGreen was used as a fluorescent marker to assess DNA amplification in single vesicles by flow cytometry. A fraction of liposomes with increased dsGreen fluorescence compared to the background was detected in the presence of dNTPs, which correspond to liposome initially containing a copy of the ori-p3 replicator, where DNA amplification took place (Figure 12C,D). The liposomes with background level fluorescence correspond to liposomes that did not initially receive a copy of the p3 gene.
8. EXAMPLE 8: Compartmentalized IVTTR for lipid synthesis by a membrane-associated enzyme (in liposomes)
In this example, IVTTR and detection of the encoded enzyme activity is demonstrated in liposomes containing on average less than a copy of the replicator per compartment. We chose PssA from the E. coli Kennedy phospholipid biosynthesis pathway. This enzyme conjugates cytidine diphosphate-diacylglycerol (CDP-DAG) with L-serine to produce cytidine monophosphate and phosphatydilserine (PS), a precursor of phosphatidylethanolamine (Figure 13A). The activity of in-vesiculo synthesized PssA enzyme in PURE system from the pssA gene was assayed by preparing liposomes containing 5 mol% CDP-DAG (Figure 13B). Production of PS lipid was detected by externally staining the liposomes with a PS-specific probe which consists in the C2-domain of lactadherin protein (LactC2) fused to a fluorescent protein like mCherry or eGFP (Figure 13B) (Blanken et al., 2019). First, using qPCR, we confirmed 10- to 100-fold amplification of the pssA gene compared to -dNTP controls with an input ori-pssA concentrations of 10 pM (Poisson parameter 2 = 0.2), in both IVTTR configurations, i.e., using purified or expressed replication proteins (Figure 13C). By flow cytometric analysis of LactC2-eGFP-stained liposomes, we observed that the mean intensity of the fluorescence signal (from recruited LactC2-eGFP) increased with functional IVTTR (+dNTPs) (Figure 13D), even though low levels of PS synthesis was detectable in -dNTP samples. The number of liposomes exhibiting a PS-positive phenotype was also higher when IVTTR was applied, i.e., when gene expression was coupled to clonal amplification (Figure 13E).
Reference: D. Blanken, D. Foschepoth, A. Cala a Serrao, and C. Danelon. Genetically controlled membrane synthesis in liposomes. Nat. Commun. 2020, 11 (1 ):4317.
9. EXAMPLE 9: Compartmentalized IVTTR for deoxyribonucleoside monophosphate kinase enzyme (in micro-droplets)
This example demonstrates the possibility to apply the invention to the screening of yet another enzymatic activity and it shows that this screen can be done in microcompartments much larger than liposomes (picoliter scale). This picoliter scale is typical of water-in-oil droplets that are often used for directed evolution protocols based on microfluidic dropletsorting chips. Here, we selected the deoxyribonucleoside monophosphate kinase from T5 phage (Uniprot: Q6QGP4). Its activity is the phosphorylation of deoxythymidine monophosphate (dTMP). In an IVTTR reaction where a nucleoside triphosphate (here dTTP) is replaced by its monophosphate counterpart (dTMP), the DNK enzyme is able to transfer the terminal phosphate of ATP molecule on dTMP to create a deoxynucleoside diphosphate (dTDP). This dTMP is converted to dTTP by Nucleoside Diphosphate Kinase (NDK) and Adenylate kinase (ADK, myokinase) using ATP as phosphate donor. Both enzymes are present in PURE system, so it is not necessary to add them. Newly produced dTTP can finally be used by Phi29 polymerase to amplify DNA flanked by origins of replication. The replicating DNA includes the gene of a fluorescent reporter protein, which gets expressed in high yield and provide a fluorescent readout.
The bulk experiment reported in Figure 14 shows that in situ expression of DNK gene can trigger DNA amplification in IVTTR when dTTP was replaced by dTMP, and that no DNA amplification happens in absence of DNK gene. The T5 phage wildtype DNK protein coding sequence was inserted downstream a T7 promoter and synthesized as a double stranded DNA fragment (GeneStrand, Eurofins), diluted in ultrapure water (Merck Milli-Q) and used without further purification. This fragment is not flanked by origins of replication so it is not recognized by Phi29 replication machinery. A IVTTR reaction was prepared using PURE/rex 2.0 (GeneFrontier) and a starting concentration of 12 pM of <gfp> as the only replicator. P2 and p3 are introduced via an expression plasmid (concentration of pUC57_oriLR_p2p3 plasmid was fixed at 100 pM) while p5 and p6 purified protein were kept at final concentration of 0.375 mg/mL and 0.105 mg/mL respectively and supplemented with 20 mM ammonium sulfate. Six conditions were produced, crossing between presence or absence of 25 pM of DNK gene fragment and composition of dNTP mixture. dNTP mixture was either 0.3 mM of each four dNTP, or 0.3 mM of dATP, dGTP, dCTP and dTMP, or absence of dNTP. Fluorescence of 5 pL reaction mixture was monitored in CFX thermocycler for twenty hours at 33 °C. As green fluorescence of positive conditions saturated detector for FAM channel (>65000 RFU), Cal Gold 540 channel was used. Final fluorescence of reaction mixtures containing all dNTP irrespectively of presence of DNK gene (A and B in Figure 14), and in condition where dTTP was replaced to dTMP in presence of DNK gene (C in Figure 14), increased their fluorescence to around 8.103 RFU. Conditions lacking all dNTP and where dTTP was replaced to dTMP in absence of DNK gene (D in Figure 14) had a final fluorescence of around 100 RFU. This result leads to conclude that presence of DNK gene is necessary for <gfp> replication and high fluorescence production when dTTP is replaced by dTMP in IVTTR. The result shown in Figure 15 shows that single DNK genes distributed in microcompartments (here 15 pm diameter droplets) can trigger the replication of encoded DNA and express easily detectable levels of a fluorescent protein. To ensure co-encapsulation of the DNK gene and the fluorescent protein gene, the gene encoding DNK enzyme in C-terminal fusion with mCherry red fluorescent protein, downstream T7 promoter and followed by T7 terminator and flanked by origins of replication. This construction was obtained as a clonal plasmid was isolated and the sequence was verified via Sanger sequencing, the replicator was then amplified from the plasmid using PCR with phosphorylated primers, giving linear fragment competent for replication. A typical IVTTR reaction was prepared using PURE/rex 2.0 (GeneFrontier). The concentration of pUC57_oriLR_p2p3 plasmid was fixed at 200 pM while p5 and p6 purified protein were kept at final concentration of 0.375 mg/mL and 0.105 mg/mL respectively, and the solution was supplemented with 20 mM ammonium sulfate. 2 pM Dextran, Alexa Fluor™ 647; 10 000 MW (Invitrogen) and 2 pM Dextran Alexa Fluor™ 680; 10 000 MW (Invitrogen) were added for pipetting control purposes and do not interfere with the reaction. dNTP mixture consisted of dATP, dGTP, dCTP and dTMP, each at a final concentration of 0.3 mM.
The replicator, containing a single copy of DNK gene and a single copy of mCherry gene, was added at a final concentration of 1 pM, giving a theoretical Poisson parameter . - 1 in 15 pm diameter droplets. This mean that, on average, each droplet contains one copy of the replicator. Flow focusing device tailored to produce 15 pm diameter droplets was used with fluorinated oil (HFE7500) supplemented with 64 mg/mL (4%) Fluosurf (Emulseo). Emulsions were incubated in PCR machine overnight at 33 °C. Emulsions were spread between hydrophobized slide and coverslip, sealed with epoxy glue to prevent evaporation. Images were taken on epifluorescence microscope (Nikon) with mCherry configuration (625 nm emission filter) at 4X magnification. Fluorescence images were inversed for easier interpretation. Brightfield and fluorescence images are gathered in Figure 15. The observation of discrete droplet showing fluorescence indicates that the presence of a single copy a DNK gene in a picoliter droplet can trigger genetic replication and expression of a fluorescent protein in sufficient amounts to produce a bright fluorescence.
10. EXAMPLE 10: MATERIALS AND METHODS OF EXAMPLES 1 TO 9
10.1. Bacterial strains and enzymes for molecular biology
Enzymes and bacterial strains were purchased from New England Biolabs (NEB) unless specified otherwise. Strain for plasmid isolation is E. coli NEB® 5-alpha.
10.2. Chemical products
Chemical products such as Ammonium Sulfate (ref A4418-1 KG) were purchased from Sigma- Aldrich unless specified otherwise.
10.3. qPCR machines and consumables
All bulk fluorescence kinetics were monitored in CFX96 machines (BioRad) using white “Low- Profile 0.2 ml 8-Tube Strips” with “Optical Flat 8-Cap Strips” (BioRad #TLS0851 and #TCS0803 respectively).
10.4. Oligonucleotides
All oligonucleotides were purchased from Eurofins with high purity salt free grade purification, except for T22 and T23 which were purified with HPLC grade purification. Sequences of said oligonucleotides are shown in paragraph 7.14 below (Table 1 ).
10.5. Double stranded DNA
Starting dsDNA material (OriLR, T7 promoter and T7 terminator), referred in Table 2 below, were submitted to Eurofins (Germany) for standard GeneStrand synthesis. Plasmids pUC57_OriLR_p2p3 was isolated from E. coui by midipreparation using Plasmid Midi kit (Qiagen) and pR045-plVEX-mCherry, pUC57_OriLR_gfp, pUC57_OriLR_rfp-hic_AC and pEX_A128_pANobar were was isolated from E. coli using Plasmid Mini Kit (Macherey Nagel), eluted in MilliQ water. All plasmids were Sanger sequenced for the region of interest. They are reported in Table 4 below. 10.6. Preparation of DNA constructs
10.6.1. PCR for phosphorylated DNA replicators
All replicators are linear and phosphorylated to be able to trigger the P2P3 replication system. To obtain replicators, all DNA constructs were amplified by PCR using Q5® HotStart polymerase and T22/T23 phosphorylated primers listed in Table 1 below. Typical PCR conditions are as follow: In a PCR microtube were mixed a given amount of DNA template (typically 1 fmol), 2 units of Q5® Hot Start DNA polymerase (NEB), 0.2 mM dNTPs and 500 nM forward and reverse primers in a total volume of 50 pL of 1X Q5 High Fidelity buffer. After an initial denaturation at 98 °C for 30 s, 25 PCR reaction cycles consisted in a denaturing step at 98 °C for 10 s, followed by an annealing step at 69 °C for 10 s, followed by an elongation step at 72 °C for 30 s per kb of replicator DNA. After cycling, a completion step at 72 °C for 2 min was performed.
10.6.2. <sfp> construct
Replicating DNA encoding GFP protein was obtained by typical PCR described above, using 0.2 fmol pUC57_OriLR_gfp and an elongation time of 37s, yielding a single 1229 bp product visible on agarose gel. Replicator was purified using PCR cleanup kit (Macherey Nagel) following manufacturer’s protocol, eluting DNA with 30 pL MilliQ water. Nucleic acid sequences of gfp gene used is shown in Table 5 below.
10.6.3. <hic> and <rfp> constructs
Replicating DNA encoding HiC and mCherry proteins were built by Golden Gate Assembly (GGA). Nucleic acid sequences of HiC (cutinase from soft-rot fungus Humicola insolens (HiC), known to possess esterase and polyesterase activity (A.M. Ronkvist, 2009)), rfp, and mCherry genes used are shown in Table 5 below. GGA fragments were obtained by PCR reaction with using primers listed in Table 1 below. Regular PCR reactions were performed with 1 fmol of DNA template, with 2 units of Q5® Hot Start DNA polymerase (NEB), 0.2 mM dNTPs and 500 nM forward and reverse primers in a total volume of 50 pL. After an initial denaturation at 98 °C for 30 s, PCR reaction consisted of a first stage of 2 cycles at low annealing temperature, then annealing temperature was raised at 72 °C for 20 cycles. Cycles consisted of a denaturing step at 98 °C for 10 s, followed by an annealing step for 10 s, followed by an elongation step at 72 °C for 30 s per kb of amplicon product. Low annealing temperatures and elongation times are reported in Table 2 below. After cycling, a completion step at 72 °C for 2 min was performed. Reaction products were purified with PCR cleanup kit (Macherey Nagel) following manufacturer’s protocol, eluting DNA with 30 pL MilliQ water. Golden Gate Assemblies were performed with 50 fmol of each purified PCR products in a final volume of 20 pL using Bsal-HF®v2 Golden Gate Assembly Kit (NEB). GGA were incubated at 37 °C for 15 min, followed by denaturation at 60 °C for 5 min, yielding circular assembly products called pLR-mCherry or pLR-HiC. One microliter of GGA product was directly used for PCR amplification of DNA constructs using phosphorylated primers T22 and T23 as described above, using 45 s elongation time, yielding a single product of 1408 bp and 1507 bp for <hic> and <rfp> respectively. PCR products were purified with PCR cleanup kit (Macherey Nagel) following manufacturer’s protocol, eluting DNA with 30 pL MilliQ water.
10.6.4. Barcoded XOriLbar and XOriRbar DNA fragments
Barcoded OriL and OriR DNA fragments were used for <rfp-hic> replicator construction. First, 1 fmol of plasmid pEX_A128_pAnobar was amplified by PCR using T93 and T360 primers to linearize it. In a PCR microtube were mixed 1 fmol of plasmid DNA, 2 units of Q5® Hot Start DNA polymerase (NEB), 0.2 mM dNTPs and 500 nM forward and reverse primers in a total volume of 50 pL of 1X Q5 High Fidelity buffer. After an initial denaturation at 98 °C for 30 s, 18 PCR cycles, consisting in a denaturing step at 98 °C for 10 s, followed by an annealing step at 68 °C for 10 s and an elongation step at 72 ° C for 60 s. After cycling, a completion step at 72 °C for 2 min was performed, yielding a single product of 2072 bp visible on agarose gel. PCR was purified using Zymo -5 kit (Zymo research) following manufacturer’s protocol, eluting DNA with 30 pL of water. This PCR product was then used for a second PCR amplification, using primers T22/M75 or M74/T23 for OriL and OriR respectively. Primers M74 and M75 bear a barcode of degenerated sequence NNNNNWNNNNNWNNNNN, where N means an equal amount of ACGT and W means an equal amount of AT during chemical synthesis (sequences of M74 and M75 in Table 1 below). In a PCR microtube were mixed 0.3 fmol of plasmid DNA, 2 units of Q5® Hot Start DNA polymerase (NEB), 0.2 mM dNTPs and 500 nM forward and reverse primers in a total volume of 50 pL of 1X Q5 High Fidelity buffer. After an initial denaturation at 98 °C for 30 s, 17 PCR cycles, consisting of a denaturing step at 98 °C for 10 s, followed by an annealing step at 55 °C for 10 s and an elongation step at 72 ° C for 20 s. After cycling, a completion step at 72 °C for 2 min was performed, yielding a single product of 346 bp and 343 bp for OriL and OriR respectively. PCR products were purified using Zymo -5 kit (Zymo research) following manufacturer’s protocol, eluting DNA with 30 pL of water. At this step, half of DNA molecules bear a mismatched barcode. We performed only one round of PCR using primers T22/M81 or T23/M80 for OriL or OriL respectively. Primers M80 and M80 and M81 are ahead of degenerated sequence, so polymerase will complete the matching sequence during this cycle. In a PCR microtube were mixed 16 fmol (1010 molecules) of plasmid DNA, 2 units of Q5® Hot Start DNA polymerase (NEB), 0.2 mM dNTPs and 500 nM forward and reverse primers in a total volume of 100 pL of 1X Q5 High Fidelity buffer. Thermocycler parameters were 40 s at 98 °C, 10 s at 70 °C, 2 min 10 s at 72 °C. After cycling, a completion step at 72 °C for 2 min was performed, yielding a single product of 346 bp and 343 bp for OriL and OriR respectively. PCR products were purified using Zymo -5 kit (Zymo research) following manufacturer’s protocol, eluting DNA with 30 pL of water. These barcoded products were called XOriLbar and XoriRbar.
10.6.5. <rfp-hic> fusion construct
Barcoded fusion protein replicators <rfp-hic> were obtained by Golden Gate Assembly from four different DNA fragments called XOriLbar, XOriRbar, prom_rfp and hic_ter. Construction of XoriLbar and XoriRbar is described above. Promrfp and hicter were obtained with PCR on plasmid pUC57_OriLR_rfp-hic_AC using primers M105/M106 or M104/M84 for prom_rfp and hic_ter respectively. In a PCR microtube were mixed 1 fmol of pUC57_OriLR_rfp-hic_AC plasmid, 2 units of Q5® Hot Start DNA polymerase (NEB), 0.2 mM dNTPs and 500 nM forward and reverse primers in a total volume of 50 pL of 1X Q5 High Fidelity buffer. After an initial denaturation at 98 °C for 30 s, PCR reaction consisted of a first stage of 4 cycles 59 °C annealing temperature, then annealing temperature was raised at 72 °C for 18 cycles. Cycles consisted of a denaturing step at 98 °C for 10 s, followed by an annealing step for 10 s, followed by an elongation step at 72 °C for 30 s. After cycling, a completion step at 72 °C for 2 min was performed, yielding a single amplicon product of 813 bp ad 756 bp for prom_rfp and hic_ter respectively. Golden Gate Assembly (GGA) was performed as follows: in a PCR microtube, 2.5 fmol of each of the four parts were mixed with 0.5 pL of BsmBI-v2 Golden Gate Enzyme Mix (NEB) in 10 pL of 1X ligase buffer. GGA mix were incubated in thermocycler for 30 cycles of 1 min at 42 ° C and 1 min at 16 °C, followed by denaturation step at 65 °C for 5 min. One microliter of GGA mix was directly used for typical PCR amplification using phosphorylated T22/T23 primers as described above, with 66 s elongation time, yielding a single product of 2199 bp. PCR product was purified using Zymo -5 kit (Zymo research) following manufacturer’s protocol, eluting DNA with 30 pL of water.
10.7. Purified P5 and P6 proteins
The DNA binding proteins P5 and P6 were prepared according to (Soengas, Gutierrez and Salas, 1995) and (Mencia et al. , “Terminal protein-primed amplification of heterologous DNA with a minimal replication system based on phage <D29”. PNAS 108, 18655-18660 (2011 )) respectively, and have the following stock concentrations and storage buffers: P5 (10 mg/mL in 50 mM Tris pH 7.5, 60 mM ammonium sulfate, 1 mM EDTA, 7 mM BME, 50% glycerol), P6 (10 mg/ml in 50 mM Tris pH 7.5, 0.1 M ammonium sulfate, 1 mM EDTA, 7 mM BME, 50% glycerol). The proteins were aliquoted and stored at -80 °C.
10.8. In Vitro Transcription Translation of genes (IVTT)
The in vitro transcription and translation systems (IVTT) PURE/rex 1.0 and PURE/rex 2.0 were purchased from Euromedex (France). Three main constituents of the kit are: energy solution, enzymes solution and ribosome solution. They were added to the final mixture as specified by the manufacturer: For PURE/rex 1.0 the proportions are one 1 /2 of energy solution, 1 /20 of enzymes solution and 1 /20 of ribosome solution. For PUREfrex 2.0, the proportions are 1 /2 of energy solution, 1 /20 of enzymes solution and 1 /10 of ribosome solution. For high protein production mimicking a successful replication, the linear DNA was added at a final concentration in the nanomolar range, typically 2 nM.
10.9. In Vitro Transcription Translation and Replication (IVTTR)
The basis of IVTTR is the same as for IVTT. The same kits were used in the same proportions as previously depicted. For replication to occurs, the solution was typically supplemented with 10 mM ammonium sulfate, 100 pM of pUC57_OriLR_p2p3, 38 pg /mL of purified P5 protein, 112 pg/mL of purified P6 protein and 300 pM of dNTP solution (New England Biolabs). The remaining of the reaction mixture was composed of replicator DNA and substrate particles at different concentrations. In some cases, purified P2 and P3 were used instead of pUC57_OriLR_p2p3.
Any piece of DNA competent for replication in the IVTTR system is herein called a “replicator” and is noted between two outward pointing chevrons in small letters (e.g. <rfp> for red fluorescent protein reporter). To be recognized by Phi29 replication machinery, the DNA construct must be double stranded, linear, and flanked by active origins of replication. The origins are named OriL and OriR, conventionally put on the left side and the right side of the construct, respectively. To be expressed in PURE system the coding sequence must be placed downstream of a T7 promoter. A T7 terminator is generally positioned downstream of the gene. We used the promoter and terminator previously described in IVTTR experiment (Van Nies et al. , “Self-replication of DNA by its encoded proteins in liposome- based synthetic cells”. Nature Communications 9, 1583 (2018)) with a slight modification: in order to be used in Golden Gate Assembly, the Bsal recognition site has been removed from the downstream of T7 promoter sequence. We opted +1 to +8 sequence of high transcription rate according to (Conrad et al., “Maximizing transcription of nucleic acids with efficient T7 promoters”. Commun Biol 3, 439 (2020)): GGGATAAT. Three replicators were designed: <rfp>, <hic> and <rfp-hic> that stands for the red fluorescent protein mCherry, the cutinase from Humicola insolens and a fusion of the two. Peptide link between the two proteins was made of thrombine cleaving site LVPRGS (SEQ ID NO:49) flanked by flexible peptides GGGGS (SEQ ID N0:50).
10.10. Fluoroqenic PLA-FDL Particle production
100 mg of Poly-lactic acid (PLA) (high molecular weight PLA LX175, Total carbion) and 4 mg of Fluorescein Diacetate substrate (FDA) (ThermoFisher Scientific) were dissolved in 1 mL of Dichloromethane (DCM, Sigma-Aldrich). This solution was added to 9 mL of water containing 5% (w/v) of polyvinyl acid (PVA) (Mw 13,000; Sigma-Aldrich). This mixture was then emulsified by sonication during 20 s in a 15 mL glass vial, on ice. The solid PCL particles were formed by overnight evaporation of the DCM solvent at RT with magnetic stirring at 400 rpm. After DCM evaporation, solid or PLA-FDL (4%) particles were washed with 3 cycles of: 10000 g, 20 min centrifugation to pellet the particles, removal and addition of 10 mL of clean MilliQ water to remove the residual PVA. Optionally a wash with ethanol can be added to remove non-encapsulated FDA. The particles suspension was then filtered with a 5-pm filter to remove potential aggregates.
10.11. Quantification of PLA-FDL plastic preparations
Polyester concentrations were quantified through the fluorescein signal obtained after total degradation by proteinase K. In a Low-Profile 0.2 ml 8-Tube Strips (BioRad) was mixed serial dilutions of diluted plastic solution in activity buffer (100 mM Tris-HCl pH 7.8) with 0.06 U of Proteinase K (New England Biolabs) in a total volume of 10 pL. Fluorescence was monitored in CFX machine at 33 °C (50° C lid temperature) until plateau was reached. Final fluorescein concentration was determined using a calibration curve of Dextran FITC 2 MDa (Invitrogen D7137) of known concentration in the same buffer. Considering a total degradation by Proteinase K of available PLA and fluorescein di-laurate, and a ratio of 4% w/w of fluorescein di-laurate over PLA, polyester microparticle concentration was calculated according to Lactic acid monomer equivalent: C3H4O2 (MW 72.06 g/mol).
10.12. Liposome sorting
<yfp> (encoding YFP, a yellow-fluorescent protein) and <minD> (encoding MinD, a non- fluorescent protein) constructs were mixed at 1 :10 molar ratio and IVTTR reactions were assembled, at a final and total DNA concentration of 10 pM in either expression buffer (PURE/rex 2.0: 50% V/V solution I, 5 % V/V solution II, and 10% V/V solution III and 0.6 units/ul of Superase- In RNase inhibitor) or in replication buffer (PURE/rex 2.0 with an addition of 20 mM ammonium sulfate, 300 pM dNTPs, 375 pg/ml purified P5 protein, 52.5 pg/ml purified P6 protein, 3 ng/pl purified P2 protein, 3 ng/pl purified P3 protein, and 0.6 units/pl of Superase- In RNase inhibitor). The well-mixed solution was encapsulated in liposomes by adding 10 mg lipid-coated beads and rotating on an automatic tube rotator (VWR) at 4 °C for 30 minutes. The mixtures were then subjected to four freeze/thaw cycles. Then, 10 pl of bead-free liposome suspension was transferred to a PCR tube, where it was mixed with 0.5 units of Proteinase K (Thermo Scientific), and incubated at 30 °C for 16 h. Three microliter of liposome suspension was mixed with 497 pl buffer and filtered through the a 35 pm nylon mesh of the cell-strainer cap from the 5 ml round-bottom polystyrene test tubes (Falcon).
Fluorescence-activated cell sorting was conducted on FACSMelody (BD Biosciences). Lasers PE-CF594(YG) and FITC-BB515, 100-micron nozzle, 23.14 PSI pressure and 34.2 kHz drop frequency were used. Photon multiplier tube voltages applied were 320 V for forward scatter, 455 for side scatter, 337 V for Texas Red, and 673 V for GFP, and a threshold of 359 V at the side scatter was applied. Liposomes with 1% highest YFP signal were sorted out from liposomes prepared in expression buffer (gate 1 ), and the same gate was applied to the liposomes prepared in replication buffer (sort 1 ) or an adjusted gate including only 0.2% highest YFP signal (sort 2). Around 50,000 (sort 1 ) or 10,000 (sort 2) liposomes were sorted into a 1.5 ml Eppendorf tube. Liposomes from gate 1 were further concentrated by centrifugation at 12000 g for 3 min, and removing 3/4 of the supernatant volume. The proteinase K was heat inactivated at 95 °C for 5 min.
Liposome suspension was next used as a template for PCR amplification using phosphorylated primers (ChD 491 /ChD 492). Reactions were set up in 100 pl volume, 300 nM each primer, 400 pM dNTP, 10 pl diluted liposome suspension, and 2 units of KOD Xtreme Hotstart DNA polymerase in Xtreme buffer, and thermal cycling was performed as follows: 2 min at 94 °C for polymerase activation, and thermal cycling at (98 °C for 10 s, 65 C for 20 s, 68 °C for 1 .5 min)x30. The amplified PCR fragments were purified using QIAquick PCR purification buffers (Qiagen) and RNeasy MinElute Cleanup columns (Qiagen) using the manufacturer’s guidelines for QIAquick PCR purification, except for longer pre-elution column drying step (4 min at 10000 g with open columns), and elution with 14 pl ultrapure water (Merck Milli-Q) in the final step. The purified DNA was quantified by the Nanodrop 2000c spectrophotometer (Isogen Life Science).
10.13. Sequences
All sequences used in the above examples are listed in the tables below.
Table 1 : Sequences of primers used in the Examples
Table 2: List of primers parameters used for PCR amplification of Golden Gate Assembly parts for <hic> and <rfp> construction (sequences are in Table 1 above) Table 3 : Sequences of GeneStrands (Eurofins) used in the examples
Table 4 : Sequences of isolated clonal plasmids used in the examples. Uppercase letters are regions verified by Sanger sequencing
Table 5: Sequences of gene-encoding replicator DNA used in the examples.
Table 6. Amino acid sequences of Phi29 p2, p3, p5 and p6 proteins recombinantly produced by recombinant expression of Phi29 phage genes gp2, gp3, gp5 and gp6 in E. coli:
Table 7. Other sequences

Claims

1. A method for screening a library of nucleic acid molecules encoding each a candidate polypeptide for one or more nucleic acid molecule(s) encoding a polypeptide with an activity of interest, comprising: a) Providing a first composition comprising an in vitro platform comprising:
(i) a nucleic acid replication machinery,
(ii) optionally a nucleic acid transcription machinery, and
(iii) a nucleic acid translation machinery; b) Providing a second composition comprising a library of nucleic acid molecules each comprising at least one candidate sequence encoding at least one candidate polypeptide; c) Mixing said first composition of step a) with said second composition of step b); d) Partitioning the thus obtained mixture into microcompartments, wherein at least one of said microcompartments comprises a single copy of said nucleic acid molecule; e) Allowing said candidate sequence to be amplified by the nucleic acid replication machinery in the microcompartments; f) Optionally allowing the amplified candidate sequence of step e) to be transcribed by the nucleic acid transcription machinery in the microcompartments; g) Allowing the amplified candidate sequence of step e) or the transcribed candidate sequence of step f) to be translated in at least one candidate polypeptide by the nucleic acid translation machinery in the microcompartments; h) Optionally adding at least one compound to the mixture of step c) or in the microcompartments of step g), and/or modifying at least one parameter within the microcompartments of step d) or g); wherein said at least one compound is preferably selected from reagents, substrates, cofactors, coenzymes of the candidate polypeptide encoded by the candidate sequence and any combination thereof; and/or wherein said at least one parameter is preferably selected from concentration, temperature, pH, viscosity, and any combination thereof; i) Detecting the activity of the candidate polypeptide of step g) in the conditions that are optionally modified in step h) within the microcompartments, preferably using an optical method; and j) Using a sorting apparatus, recovering the candidate sequences encoding the candidate polypeptide having the activity of interest from microcompartments displaying specific activity signal.
2. The method of claim 1 , wherein the microcompartments are selected from the group consisting of microdroplets, microvesicles, microchambers, and any combination thereof.
3. The method of claim 1 or 2, wherein the nucleic acid molecules are selected from the group consisting of DNAs, RNAs, and any combination thereof.
4. The method of any one of the preceding claims, wherein the nucleic acid replication machinery comprises one or more materials selected from: DNA polymerases, RNA polymerases, DNA terminal proteins, Single-Stranded DNA Binding Proteins (SSBPs), Double- Stranded DNA Binding Proteins (DSBPs), nucleic acid sequences coding for DNA polymerases, nucleic acid sequences coding for RNA polymerases, nucleic acid sequences coding for DNA terminal proteins, nucleic acid sequences coding for SSBPs, nucleic acid sequences coding for DSBPs, and any combination thereof.
5. The method of any one of the preceding claims, wherein the nucleic acid transcription machinery comprises an RNA polymerase.
6. The method of any one of the preceding claims, wherein the nucleic acid translation machinery comprises all the required materials for completing nucleic acid translation, such as ribosomes, translation factors, aminoacyl tRNA synthetases, and tRNAs.
7. The method of any one of the preceding claims, wherein the in vitro platform is a cell- free system, such as a cell extract and a reconstituted in vitro system, and any combination thereof.
8. The method of any one of the preceding claims, wherein the proteins included in or encoded by the nucleic acid molecules included in the nucleic acid replication machinery, optionally the nucleic acid transcription machinery, the nucleic acid translation machinery, or any combination thereof, are selected or derived from a bacteriophage, a bacterium, a yeast, a virus, or any combination thereof.
9. The method of any one of claims 1 to 8, wherein the activity of interest is polymer enzymatic degradation, and the at least one compound added the mixture of step c) or in the microcompartments of step g) is a fluorogenic or chromogenic compound embedded in one or more polymer particles, wherein degradation of the polymer particles catalyzed by the candidate polypeptide leads to release and conversion of the compound into a fluorescent or colored product.
10. The method of any one of claims 1 to 9, wherein the activity of the candidate polypeptide encoded by the candidate sequence is determined by measuring the intensity of one or more signals, preferably using an optical technique selected from the group consisting of absorption-based detection techniques, fluorescence- based detection techniques, and any combination thereof.
11. The method of any one of claims 1 to 10, wherein the method is a high-throughput screening (HTS) method, preferably an ultra-high-throughput screening (uHTS) method.
12. The method of any one of claims 1 to 11 , wherein the method allows high or ultrahigh throughput screening of at least 105 nucleic acid molecules comprising at least one candidate sequence, preferably at least 106 nucleic acid molecules comprising at least one candidate sequence.
13. The method of any one of claims 1 to 12, wherein a candidate polypeptide encoded by the candidate sequence is selected from the group consisting of putative enzymes, putative transcription factors, putative functional derivatives thereof, and putative functional fragments thereof.
14. The method of any one of claims 1 to 13, where said compound of step h) is selected from: substrates, cofactors, and coenzymes (including prosthetic groups and cosubstrates) of the candidate polypeptide encoded by the candidate sequence, or wherein the nucleic acid molecule further comprises a sequence encoding a protein substrate, wherein the substrate is selected from the group consisting of a substrate of an enzyme, a substrate of a cascade of enzymes, a functional derivative thereof, and a functional fragment thereof.
15. The method of any one of the preceding claims, wherein, in step d), the mixture is partitioned into microcompartments using poissonian partitioning, preferably so that a substantial fraction of said microcompartments comprises a single copy of said nucleic acid molecule.
16. A kit for implementing the method of any one of claims 1 to 15, comprising: a) an in vitro system /platform comprising:
(i) a nucleic acid replication machinery,
(ii) optionally a nucleic acid transcription machinery, and
(iii) a nucleic acid translation machinery; b) at least one material for making microcompartments, preferably at least reagent or device for making microcompartments, more preferably at least one partitioning agent or at least one microfluidic microarray.
17. Use of the method of any one of claims 1 to 15 in a directed evolution method, said directed evolution method preferably comprising multiple cycles of mutagenesis followed by screening with the method of any one of claims 1 to 15.
EP23804695.7A 2022-11-10 2023-11-10 Micro-compartmentalized ultra-high throughput screening from single copy gene libraries Pending EP4615970A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
EP22315279 2022-11-10
PCT/EP2023/081530 WO2024100292A1 (en) 2022-11-10 2023-11-10 Micro-compartmentalized ultra-high throughput screening from single copy gene libraries

Publications (1)

Publication Number Publication Date
EP4615970A1 true EP4615970A1 (en) 2025-09-17

Family

ID=85036080

Family Applications (1)

Application Number Title Priority Date Filing Date
EP23804695.7A Pending EP4615970A1 (en) 2022-11-10 2023-11-10 Micro-compartmentalized ultra-high throughput screening from single copy gene libraries

Country Status (3)

Country Link
EP (1) EP4615970A1 (en)
JP (1) JP2025539740A (en)
WO (1) WO2024100292A1 (en)

Also Published As

Publication number Publication date
JP2025539740A (en) 2025-12-09
WO2024100292A1 (en) 2024-05-16

Similar Documents

Publication Publication Date Title
US10308931B2 (en) Methods for screening proteins using DNA encoded chemical libraries as templates for enzyme catalysis
Katzen et al. The past, present and future of cell-free protein synthesis
Albayrak et al. Cell-free co-production of an orthogonal transfer RNA activates efficient site-specific non-natural amino acid incorporation
Shimizu et al. Cell‐free translation systems for protein engineering
US7883848B2 (en) Regulation analysis by cis reactivity, RACR
US20260056184A1 (en) Methods for generating and screening compartmentalised peptide libraries
KR20020059370A (en) Methods and compositions for the construction and use of fusion libraries
EP3665286B1 (en) Cell-free protein expression using double-stranded concatameric dna
Napiorkowska et al. High‐throughput optimization of recombinant protein production in microfluidic gel beads
US12571017B2 (en) Devices and methods for producing nucleic acids and proteins
EP3914704B1 (en) A method for screening of an in vitro display library within a cell
US20060084136A1 (en) Production of fusion proteins by cell-free protein synthesis
WO2024100292A1 (en) Micro-compartmentalized ultra-high throughput screening from single copy gene libraries
US20210214791A1 (en) Absolute quantification of target molecules at single-entity resolution using tandem barcoding
EP3545104B1 (en) Tandem barcoding of target molecules for their absolute quantification at single-entity resolution
US20220002719A1 (en) Oligonucleotide-mediated sense codon reassignment
CN112538469A (en) Restriction endonuclease DpnI preparation and preparation method thereof
Cirino et al. Protein engineering as an enabling tool for synthetic biology
WO2023048291A1 (en) tRNA, AMINOACYL-tRNA, POLYPEPTIDE SYNTHESIS REAGENT, UNNATURAL AMINO ACID INCORPORATION METHOD, POLYPEPTIDE PRODUCTION METHOD, NUCLEIC ACID DISPLAY LIBRARY PRODUCTION METHOD, NUCLEIC ACID/POLYPEPTIDE CONJUGATE, AND SCREENING METHOD
EP3519571B1 (en) Compositions, methods and systems for identifying candidate nucleic acid agent
Shozen et al. Amber codon-mediated expanded saturation mutagenesis of proteins using a cell-free translation system
Terasaka et al. Efficient cell-free evolution of RNA polymerases by droplet microfluidics
Botte et al. Cell-free synthesis of macromolecular complexes
Bouhedda et al. Compartmentalization‐Based Technologies for In Vitro Selection and Evolution of Ribozymes and Light‐Up RNA Aptamers
Restrepo et al. Efficient multi-gene expression in cell-free droplet microreactors

Legal Events

Date Code Title Description
STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: UNKNOWN

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE

PUAI Public reference made under article 153(3) epc to a published international application that has entered the european phase

Free format text: ORIGINAL CODE: 0009012

STAA Information on the status of an ep patent application or granted ep patent

Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE

17P Request for examination filed

Effective date: 20250610

AK Designated contracting states

Kind code of ref document: A1

Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR

DAV Request for validation of the european patent (deleted)
DAX Request for extension of the european patent (deleted)