EP4681203A1 - System and method for protein sequence screening - Google Patents
System and method for protein sequence screeningInfo
- Publication number
- EP4681203A1 EP4681203A1 EP24714556.8A EP24714556A EP4681203A1 EP 4681203 A1 EP4681203 A1 EP 4681203A1 EP 24714556 A EP24714556 A EP 24714556A EP 4681203 A1 EP4681203 A1 EP 4681203A1
- Authority
- EP
- European Patent Office
- Prior art keywords
- protein
- sequences
- input
- expression
- sequence
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B15/00—ICT specially adapted for analysing two-dimensional [2D] or three-dimensional [3D] molecular structures, e.g. structural or functional relations or structure alignment
- G16B15/20—Protein or domain folding
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B35/00—ICT specially adapted for in silico combinatorial libraries of nucleic acids, proteins or peptides
- G16B35/20—Screening of libraries
-
- G—PHYSICS
- G16—INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
- G16B—BIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
- G16B40/00—ICT specially adapted for biostatistics; ICT specially adapted for bioinformatics-related machine learning or data mining, e.g. knowledge discovery or pattern finding
- G16B40/20—Supervised data analysis
Definitions
- systems and methods for protein sequence screening are provided herein.
- systems and methods are provided that determine one or more preferred candidates from a list of protein sequences, express the one or more preferred candidates to yield one or more candidate proteins, and determine an optimal protein of the one or more candidate proteins.
- Determining the degree of similarity between the protein variants and the target protein can be time-consuming and expensive, particularly if every candidate protein variant is to be expressed and assayed to determine its structure.
- Prior art protein structure prediction software includes the AlphaFold R TM software.
- AlphaFold employs a deep learning architecture, specifically a deep residual convolutional neural network (CNN), to predict protein structures.
- CNN convolutional neural network
- Deep learning models, particularly CNNs, have shown great promise in various domains due to their ability to learn hierarchical features from data.
- the model is trained on a large dataset of known protein structures, which serve as examples of how amino acid sequences fold into three-dimensional structures. These structures are obtained from publicly available databases like the Protein Data Bank (PDB). During training, the model learns to predict the distance and orientation between pairs of amino acids in a protein sequence.
- PDB Protein Data Bank
- AlphaFold learns to predict the distance and orientation between pairs of amino acids in a protein sequence.
- One of the key innovations of AlphaFold is its method for predicting inter-residue distances between pairs of amino acids in a protein sequence. Instead of directly predicting distances, AlphaFold predicts a probability distribution over the distances between pairs of residues. This approach allows the model to capture the uncertainty inherent in distance predictions. Using the predicted distance distributions, AlphaFold employs optimization algorithms, such as gradient descent, to infer the most probable spatial arrangement of amino acids in the protein sequence.
- AlphaFold refines the predicted structures using optimization techniques to improve their accuracy. This refinement step helps to correct any errors or inaccuracies in the initial predictions. AlphaFold also provides estimates of the uncertainty associated with its predictions. This is crucial because protein structure prediction inherently involves uncertainty due to the complexity of protein folding and the limitations of available data.
- the final output of AlphaFold is a predicted three-dimensional structure of the protein, represented in terms of atomic coordinates. This structure can be visualized and analyzed to gain insights into the protein's function and interactions.
- Goverde of identifying preferred protein sequences from a plurality of input sequences based on determined structural properties, nor of expressing and characterising the preferred sequences to determine one or more optimal sequences.
- the method of Goverde further differs from the present invention in that a single sequence, which does not code for a variant of a target protein, is input into the AF2 model as an initialisation step.
- Embodiments of the invention seek to provide a method for analysing protein variants (for example length variants, isoforms, polymorphisms, orthologs, homologs or the structure of binding sites), the method comprising receiving input data comprising a plurality of amino acid sequences or their corresponding DNA sequences, each of the DNA sequences coding for a variant of a target protein; for each of the input sequences in the plurality of sequences, predicting a folded protein structure based on the sequence of the expressed protein of the input sequence, determining one or more structural properties of the predicted folded protein structure; identifying one or more preferred input sequences based on the determined structural properties; expressing the coded variant of each of the identified one or more preferred input sequences in a cell-free or cell-based system; and characterizing the protein variant sequences based on the expressed one or more coded variants.
- protein variants for example length variants, isoforms, polymorphisms, orthologs, homologs or the structure of binding sites
- Protein expression usually benefits if the expressed protein can be purified from the expression system. Most purification systems rely on affinity binding to a tag attached to the expressed protein. Protein structure prediction can be used to determine if a particular expressed protein has an available binding tag, or whether the binding tag is likely to be folded internally within the protein structure. Proteins where the binding tag is not externally available are not suitable candidates for expression. Embodiments of the invention seek to use protein structure prediction to triage input sequences for protein expression and characterisation.
- FIG. 1 For embodiments of the invention, seek to address the above problem by providing a system for analysing protein variants, the system comprising an input module configured to receive input data comprising a plurality of input sequences, each of the sequences coding for a variant of a target protein; a processing module configured to receive the input data from the input module, and provide the input data to a sequence analysis module; a sequence analysis module configured to, for each of the input sequences in the plurality of sequences predict a folded protein structure of the coded protein of the input sequence, and determine a structural similarity between the predicted folded protein structure and the target protein; the processing module being further configured to receive the determined structural similarities from the sequence analysis module, and identify one or more preferred input sequences based on the determined structural similarities; a protein expression module configured to express the coded variant of each of the identified one or more preferred input sequences in a cell-based or cell- free system; the processing module being further configured to determine one or more optimal input sequences based on the expressed one or more coded variants.
- Further embodiments of the invention seek to express a protein having the shortest functionally equivalent domain.
- Expression of proteins, particularly large proteins can be limited by the length of the nucleic acid template and the amounts of reagents required for expression.
- Structure prediction can be used to identify regions of the amino acid sequence which can be removed without impacting the function of the expressed protein. The undesired regions may include disordered regions or regions which may promote aggregation.
- the reagents such as amino acids are resource limited, expression of shorter chains gives a higher concentration (for example half as long may be twice the number of molecules).
- structure prediction may be used to design nucleic acid or protein sequence inputs for protein expression.
- Figure 1 shows a method for screening input sequences in accordance with an embodiment of the present invention.
- Figure 2 shows a schematic of a system for screening input sequences in accordance with an embodiment of the present invention.
- Figure 3 shows a protein of interest with a peptide tag (detection tag) in the presence of a protein that binds to the peptide tag (detector protein).
- the detection tag and detector protein generate a detectable signal upon binding, then the quantity of the protein of interest can be measured.
- Figure 4 shows a cartoon representation of protein expression of different proteins having varying soluble yields and levels of aggregation. Some conditions give no expression (dark). Some conditions give a high level of soluble protein (uniform intensity white squares). Some conditions give protein aggregates, meaning the expressed proteins are insoluble and clump together (white spots in dark background). Some give a mix of soluble and aggregated proteins (white spots in a grey background). Thus the level of soluble expression and level of aggregation can be determined by adding a detector species to the expressed proteins.
- Figure 5 shows the experimental result from 24 different proteins expressed in a reconstituted cell-free protein synthesis system in droplets on an electrowetting on dielectric (EWoD) device.
- Each construct contains a ccGFPn tag.
- the ccGFPi-io detector species is present from the start of expression.
- the rows marked Endpoint shows the fluorescence signal from 10 hours expression in the absence of ccGFPi-io detector species followed by 5 hours complementation with the ccGFPi-io detector species.
- This experiment showed significant differences between expression/complementation in Screen BioInk compared to Endpoint detection.
- the detected protein clusters formed after expression only with endpoint detection mean that the protein aggregated after expression, thereby lowering the soluble yield.
- Protein structure prediction can be used to screen sequences prior to expression, for example to identify those which may aggregate, or for which the detector tag is less accessible for binding.
- FIG. 6 shows AlphaFold structures for 8 terminal transferase (TdT) enzymes.
- TdT terminal transferase
- Figure 7 shows the removal of the unstructured regions flanking the BRCT domain and the BRCT domain itself increased expression and yields of the TdT ortholog candidate TdT proteins.
- Figure 8 shows structure predictions for a variety of length truncated 9°N polymerases. The structure of the shortened chains having the functional core unchanged were selected for expression.
- Figure 9 shows date from an expression screening report on various length variants of a DNA polymerase. Structure prediction was used to guide the choice of constructs for expression.
- the present invention solves the problems set out above by providing a system and method that receives input data comprising a plurality of input sequences that each code for a protein variant of a target protein and/or an indication of the target protein having an accessible binding tag or detection tag, predicts the folded protein structure of each variant sequences, determines a structural property of the folded protein structure of the variant sequence, expresses the chosen variant input sequences in a cell-free or cell-based system, and determines the optimum variant sequence or sequences based on the expressed input sequences and optionally the availability of the binding or detection tags.
- the system and method of the present invention may be fully automated, and thereby operate substantially without human intervention.
- the present invention therefore allows a user to determine which variant, or variants, of a target protein are candidates for protein expression and exhibit one or more preferred characteristics, including optionally the ability to detect or purify the protein, without requiring the user to express and assay every possible input protein variant.
- a method for analysing protein sequences comprising: receiving input data comprising a plurality of input sequences, each of the input sequences coding for a variant of a target protein; for each of the input sequences in the plurality of input sequences: predicting a folded protein structure of the coded protein of the input sequence; and determining one or more structural properties of the predicted folded protein structure; identifying one or more preferred input protein sequences based on the determined structural properties; taking nucleic acid input sequences for expressing the input protein sequences; expressing the coded variant of each of the identified one or more preferred input protein sequences; characterizing the protein variant sequences based on the expressed one or more coded variants; and determining one or more optimal input sequences based on the expressed one or more coded variants.
- Protein variants may include for example length variants, isoforms, polymorphisms, orthologs, homologs or the structure of binding sites.
- Variants of a protein may include N-terminal truncations, C-terminal truncations, internal deletions, one or more point mutations, orthologs and homologs. Truncations and deletions may be used to remove regions of a protein expected to negatively impact protein expression, folding or stability. For instance, highly disordered regions may be truncated from the termini of a protein.
- Orthologs are proteins that diverged as a result of speciation.
- the protein TYK2 Non-receptor tyrosine-protein kinase TYK2
- mice Q.9R117
- the two proteins are said to be orthologs.
- Expressing orthologs as protein variants may help obtaining proteins and is also an essential component of the drug discovery process due to the use of ortholog animal models.
- Commonly used orthologs include mice, dogs, zebrafish, rats, and non-human primates.
- the invention may choose similar and dissimilar isoforms. For example from a list of isoforms, some are predicted to be similar and some dissimilar. Similar and dissimilar inputs may be chosen in order to screen for diversity.
- the invention may be used for lead compound toxicity screening. For example to screen a diversity of protein targets for binding to particular compounds. Thus a range of predicted input structures may be chosen for expression and binding, which a view to seeing if a particular lead compound is target selective or binds to a wide variety of protein structures, a likely indicator of toxicity.
- the invention may be used to identify or characterise molecular binding sites.
- the ability to predict structures allows identification of mutations which alter a particular binding site.
- protein sequences can be expressed to perform a binding site screen, thus making inactive mutants.
- the binding sites may be small molecule binding sites, antibody binding sites or sites involved in protein-protein interactions.
- An example of binding site determination is Wang et. al. (Identification of Drug Binding Sites and Action Mechanisms with Molecular Dynamics Simulations Curr Top Med Chem. 2018;18(27):2268-2277).
- Use of the invention may involve for example 6 isoform structures predicted through this pipeline. Through public or proprietary prior knowledge (i.e., structurally determined binding site), the binding pocket is identified. The structural similarity of the binding sites are determined. The user may use programs like autodock or GoLD to dock ligands. The user then chooses all 6 proteins to be expressed/characterized/purified to empirically validate if their prediction of binding strength between the 6 targets held true.
- the user if the user is trying a multi-target small molecule drug discovery approach, the user understands that global RMSD is quite a crude comparison and wants to compare 6 isoform drug binding sites based on prior knowledge.
- the user wants to be able to target 5 out of 6 is- forms but does not want to target 1 of them.
- the user needs to know if the isoforms are structurally different at that binding pocket. Can they use 1 isoform to target the 5 similar isoforms? Is their hypothesis correct that they can counter select against one of the isoforms through binding to a specific site? Are the LEC flanks they're appending going to influence the site that differentiates 1 isoform from the other 5?
- the selection process helps the user design which isoforms to make and how the construct looks.
- protein folding refers to the process by which a sequence of amino acids, such as a sequence of amino acids of a particular protein, folds into a specific, typically complex, three-dimensional structure.
- the particular folded three-dimensional structure of the amino acid sequence effects its biological characteristics, including the existence of, and/or location of molecular binding sites.
- the molecular binding sites may be binding pockets for therapeutic agents.
- Therapeutic agents may, non-exhaustively, include small molecule drugs, proteins, and nucleic acids. It is important that the existence, location, and structure of molecule binding sites reflects the native, naturally occurring target protein. As such, when determining the structural similarity between variants, there may be more weight given to the similarity of certain regions of the protein, for instance known molecular binding sites may be weighted highly compared to less structured or unstructured domains of the target protein.
- the molecular binding sites may be part of the protein of interest or may be flank sequences appended to a protein of interest (POI).
- the POI may have a known structure or an unknown structure.
- the POI having flank sequences where the flanks change the structure of the POI may not be ideal candidates to obtain the POI.
- the flank sequences may cause the POI to unfold.
- the flank sequences may cause the POI to adopt a different, unnatural, or non-native structure.
- the flank sequences may have a binding or detection tag which is not available for binding when the protein is expressed. Protein structure prediction can be used to decide the optimal candidates for a particular set of conditions, and the conditions where a protein is most likely to be obtained.
- the input sequences and expression conditions can be pre-selected and the sequences and conditions where protein is unlikely to be obtained in measurable form, due to for example lack of stability or binding affinity may be down-selected.
- optimise linker length If linker is too short, then there may be negative effects on fusion tag - POI interaction. If linker is too short, the affinity tag may be inaccessible.
- flank sequences may include elements required for transcription and translation, such as for example a promoter region such as a T7 promoter biding site, a ribosome binding site, start codons and stop codons.
- the flank sequences may optionally also include elements such as solubility tags, purification tags or detection tags.
- Figure 1 shows a schematic functional component diagram of the system 100 according to an embodiment of the present invention.
- the system 1 comprises an input/output module 101, user interface 102, processing module 103, sequence analysis module 105, and protein expression module 106. It will be appreciated that the system 100 according to the present invention need not necessarily include every module shown in Figure 1, and may include other modules that are not illustrated in Figure 1.
- the input/output module 101 is configured to receive input data from a user.
- the input data comprises a plurality of input sequences, either as amino acid sequences or DNA sequences, where each DNA sequence codes for a variant of a target protein.
- DNA sequences are familiar to the skilled person, and may be in the form of a character string where each character signifies a specific nucleotide.
- An example of a DNA sequence is: AAACAAGTGGGT.
- Such a sequence codes for four amino acids (KQVG) following triplet codon translation.
- the input sequence may be a DNA or protein sequence.
- the DNA sequence may have any suitable length to code for a protein variant, the protein variant optionally having additional flanks such as those needed for protein expression.
- a protein variant formed of 1000 amino acids may have a DNA sequence of 3000 bases. It will be appreciated that the DNA sequences need not necessarily be represented as a character string and may be represented in any other suitable format, including vector format, alpha-numeric, binary, hexadecimal etc.
- the input/output module 101 may be configured to receive an indication of a target protein, where the plurality of received input sequences code for protein variants of said target protein.
- the input/output module 101 may be configured to receive the indication of the target protein in addition to the plurality of input sequences.
- the input/output module 101 may be configured to receive an indication of a target protein instead of the plurality of input sequences.
- the system 100 may calculate, using the processing module 103, a plurality of input sequences that code for protein variants of the received target protein.
- the input/output module 101 may be configured to receive a plurality of amino acid sequences.
- Amino acid sequences are also known to the skilled person and are signified by a character string where each character represents a particular amino acid.
- DNA and amino acid sequences can be easily translated back and forth through the use of genetic code and codon tables. It should be noted that while a DNA sequence typically codes for a single amino acid sequence (if in frame from a single start codon to a single stop codon), an amino acid sequence could be reverse translated into many different DNA sequences.
- Alanine (A) can be encoded as GCU, GCC, GCA, or GCG.
- the input/output module 101 may receive input data via the user interface 102.
- the user interface 102 may comprise a Graphical User Interface.
- a Graphical User Interface may be provided in the form of a widget embedded in web site, as an application for a device, or on a dedicated landing web page.
- Computer readable program instructions for implementing the Graphical User Interface may be provided to a user device from a remote computer readable storage medium via a network connection, for example, the Internet, a local area network, a wide area network and/or a wireless network.
- the input/output module 101 may receive input data via a network connection.
- One possible implementation of the user interface is a web-based application accessible via a web browser which communicates inputs and outputs with a cloud based server via a REST API.
- the processing module 103 is configured to receive the input data from the input/output module 101.
- the processing module is communicatively coupled to the input/output module 101, the sequence analysis module 105, and the protein expression module 106.
- the processing module 103 receives input data from the input/output module 101 comprising a plurality of input sequences, each input sequence coding a protein variant of a target protein.
- the processing module 103 may further receive input data comprising an indication of the target protein, for example a further input sequence that codes for the target protein.
- the processing module 103 receives a DNA sequence that codes for a target protein, and is configured to generate a plurality of DNA sequences that each code for a different variant of the target protein.
- the processing module 103 receives a protein sequence the codes for a target protein, and is configured to generate a plurality of protein sequences that each code for a different variant of the target protein.
- the processing module 103 may convert the desired protein sequences into nucleic acid sequences suitable for expression.
- the sequences may be codon optimised for the expression system.
- a possible implementation of the input/output process is a cloud based publish-subscribe messaging service.
- the processing module 103 outputs the received or generated plurality of input sequences to the sequence analysis module 105.
- a possible implementation of the processing module is a managed cloud-based pipeline service which schedules and runs containerized computation jobs in a direct acyclic graph manner.
- the sequence analysis module 105 is configured to receive a plurality of DNA sequences, or amino acid sequences and predict the folded protein structure of the coded protein of each DNA sequence or amino acid sequence.
- a possible implementation of the sequence analysis module is a multiple sequence alignment module detecting similarities between multiple sequences and their subsequences.
- the sequence analysis module 105 receives as input data representing a particular input sequence of the plurality of input sequences, and outputs an estimate of the folded protein structure of the protein that the particular input sequence codes for.
- the folded protein structure is output in the form of a three-dimensional configuration of the atoms of the coded protein of the DNA sequence, when the coded protein has undergone protein folding.
- the three-dimensional configuration may be output in the form of a coordinate vector, the vector having a length equal to the number of amino acids in the coded protein of the input sequence, and each element of the vector corresponding to the three-dimensional cartesian coordinates of the corresponding amino acid of the input sequence.
- the folded protein structure may be output in any other suitable format, for example a parametric model, or an alternative coordinate system.
- the folded protein structure may be output in the format known as the Protein Data Bank (PDB) format, which is a standardised form for files containing atomic coordinates.
- PDB Protein Data Bank
- Identification of global versus local similarity represents two orthogonal directions in comparison of protein structures, i.e. structures that are most similar globally may not be the best in terms of local similarity.
- Flexible or disordered fragments such as long loops and/or termini are often poorly predicted and may significantly compromise the otherwise good similarity between structures.
- Relative domain movements observed in multi-domain proteins can also contribute to the poor global similarity scores. Focusing on local similarity helps to avoid these issues. Local similarity can be interpreted as a cumulative similarity score for all regions of the protein or, otherwise, can focus on a specific region such as, for example, ligand binding pocket, while ignoring the remaining parts of the protein.
- Root Mean Square Deviation is the most commonly used quantitative measure of the similarity between two superimposed atomic coordinates. RMSD values are presented in A and calculated by where the averaging is performed over the n pairs of equivalent atoms and d, is the distance between the two atoms in the /-th pair. RMSD can be calculated for any type and subset of atoms; for example, Ca atoms of the entire protein, Ca atoms of all residues in a specific subset (e.g. the transmembrane helices, binding pocket, or a loop), all heavy atoms of a specific subset of residues, or all heavy atoms in a small-molecule ligands. The most common way to evaluate the correctness of the docking geometry is to measure the Root Mean Square Deviation (RMSD) of the ligand from its reference position in the answer complex after the optimal superimposition of the receptor molecules.
- RMSD Root Mean Square Deviation
- the sequence analysis module 105 may utilize one or more machine learning models, including but not limited to one or more neural network models.
- Such neural network models are well known to the skilled person and comprise a plurality of interconnected nodes.
- the machine learning models may be deep models that employ multiple levels of machine learning models, such as combinations of multiple neural networks or combinations of neural networks and/or other machine learning models.
- neural network based system that may form a part of or all of the sequence analysis module 105 is the AlphaFold Al system developed by DeepMind and EMBL's European Bioinformatics Institute. Implementational details of the AlphaFold Al system can be found in US 2021/304847 Al, the contents of which are incorporated herein by reference in its entirety.
- the processing module 103 receives from the sequence analysis module 105 data indicating the structural similarity between the folded protein structures of each of the variant input sequences and the folded protein structure of the target protein. Based on the received data, the processing module determines a set of preferred input sequences that correspond to the most structurally-similar variants to the target protein. The number of preferred input sequences may be selected by the user. Merely by way of example, if 10 input sequences are passed to the sequence analysis module 105, the processing module 103 may select the top 4 most structurally-similar input sequences as preferred input sequences.
- the processing module 103 may select the top 4 most likely input sequences to express a protein having an available binding tag as preferred input sequences.
- the processing module 103 may select the top 4 most likely DNA sequences to express a protein having an available detection tag as preferred DNA sequences.
- the processing module 103 may select the top 2 most likely sequences to express a structurally-similar protein that has an available purification tag and an available detection tag.
- the processing module 103 receives the folded protein structure of each of the input sequences from the sequence analysis module 105, and determines the structural similarity between the folded protein structure of the target protein and that of each input sequence.
- the protein expression module 106 is configured to receive data from the processing module 103 indicating the preferred input sequences based on their determined structural similarity to the target protein.
- the protein expression module 106 is configured to express each of the preferred DNA sequences in a cell-free system. Possible implementations of the Protein Expression module is the Nuclera eProtein Discovery”TM system, an automated liquid handling protein expression system, or a manual protein expression protocol. Any suitable platform for protein expression may be used, for example liquid handling in microtitre plates.
- the chosen protein may be expressed in cells or using a cell-free expression system.
- the cell-free system may be a cell lysate or a reconstituted system.
- Cell-free protein synthesis also known as in-vitro protein synthesis or CFPS, is the production of peptides or proteins using biological machinery in a cell-free system, that is, without the use of living cells.
- the in-vitro protein synthesis environment is not constrained within a cell wall or limited by conditions necessary to maintain cell viability, and enables the rapid production of any desired protein from a nucleic acid template, usually plasmid DNA or RNA from an in-vitro transcription.
- CFPS has been known for decades, and many commercial systems are available.
- Cell-free protein synthesis encompasses systems based on crude lysate (Cold Spring Harb Perspect Biol.
- CFPS requires significant concentrations of biomacromolecules, including DNA, RNA, proteins, polysaccharides, molecular crowding agents, and more (Febs Letters 2013, 2, 58, 261- 268).
- the protein expression module 106 performs the cell-free expression of the preferred DNA sequences in a digital microfluidic device.
- An example implementation of cell-free expression in digital microfluidic devices is described in WO 2022/038353, the entire contents of which are incorporated herein by reference.
- Microfluidic devices for manipulating droplets or magnetic beads based on electrowetting have been extensively described. Electrowetting is the modification of the wetting properties of a surface (which is typically hydrophobic) with an applied electric field. In the case of droplets in channels the manipulation of droplets can be achieved by causing the droplets, for example in the presence of an immiscible carrier fluid, to travel through a microfluidic channel defined by the walls of a cartridge or microfluidic tubing.
- DMF digital microfluidics
- DMF utilizes alternating currents on an electrode array for moving fluid on the surface of the array. Liquids can thus be moved on an open-plan device by electrowetting. Digital microfluidics allows precise control over the droplet movements including droplet fusion and separation.
- the protein expression module 106 is configured to determine one or more optimal DNA sequences.
- the optimal DNA sequences may be determined by calculating one or more assay metrics, the one or more assay metrics being calculated by performing one or more protein assays.
- Protein assays are well known to the skilled person, and can be used to measure a variety of characteristics of the expressed protein including the amount of protein expressed, the solubility of the protein, protein stability or the purifiability of the protein.
- the protein expression module 106 may perform a known assay such as fluorescence complementation, the Bradford assay, Folin- Lowry assay, or Bicinchoninic Acid (BCA) assay to determine the amount of each variant protein expressed in the cell-free system. The protein expression module 106 may then determine one or more optimal input sequences corresponding to the highest expressing variant proteins.
- a known assay such as fluorescence complementation, the Bradford assay, Folin- Lowry assay, or Bicinchoninic Acid (BCA) assay to determine the amount of each variant protein expressed in the cell-free system.
- BCA Bicinchoninic Acid
- the determination of the one or more optimum input sequences is performed by the processing module 103, based on assay metric data output from the protein expression module 106.
- the determination of structure may be based on computed features such as surface charge, surface hydrophobicity, the presence of unstructured regions, solubility predictions or the predicted melting temperature. For example the elimination of regions having a high surface charge may be beneficial for use in certain expression devices. The removal of unstructured or hydrophobic regions may help to prevent aggregation or improve the level of soluble expression.
- Tm melting temperature
- PROTHERM is a web server that predicts the thermal stability of proteins and their mutants. It uses an empirical approach based on experimental data from the ProTherm database. Users can input the amino acid sequence or the PDB ID of the protein and obtain predictions for various parameters, including melting temperature.
- FoldX is a molecular modeling software package that includes modules for predicting protein stability and melting temperature. It utilizes an empirical force field to estimate the free energy change upon protein unfolding. Users can input protein structures in PDB format and calculate various thermodynamic parameters, including melting temperature.
- PoPMuSiC Prediction of Protein Mutant Stability Changes
- PoPMuSiC is a web server for predicting the effects of mutations on protein stability, including melting temperature changes. It employs a statistical potential-based approach to estimate the change in protein stability upon mutation. Users can input protein sequences or structures along with mutations and obtain predictions for stability changes and melting temperature shifts.
- ThermoProt is a web server for predicting the thermal stability of proteins. It integrates various computational methods, including machine learning algorithms and biophysical models, to predict melting temperatures. Users can input protein sequences or structures, and the server provides predictions along with confidence scores.
- DUET Dynamic Undocking and Energy Transfer
- the protein expression module 106 can be an automated protein expression system configured to take one or more protein sequences as input, express each of the one or more protein sequences in a cell or cell-free system, perform one or more protein assays on the expressed one or more proteins to calculate one or more assay metrics, and determine one or more optimal protein sequences based on the calculated assay metrics (or output the calculated assay metrics to the processing module 103). The automated protein expression module then outputs the determined one or more optimal protein sequences, and/or calculated assay metrics, to the processing module. As an automated system the automated protein expression module performs these steps substantially or entirely without human intervention.
- an automated protein expression module may be implemented as a digital microfluidics system communicatively coupled to, or comprised within, the system 100.
- Automated systems can handle the transformation and cultivation of bacterial and yeast systems for protein expression.
- Escherichia coli E. coli
- Bacillus subtilis is another bacterium commonly used for protein expression, especially for secreted proteins.
- Automated systems can handle the transformation and cultivation of B. subtilis strains.
- automated systems can handle yeast transformation, growth, and induction for eukaryotic protein expression.
- Pichia pastoris is a methylotrophic yeast system often used for expressing eukaryotic proteins at high levels. Automation can assist in handling the transformation and cultivation of P. pastoris strains.
- Baculovirus-lnsect Cell System (BEVS): Insect cells such as Sf9 or Sf21 infected with recombinant baculovirus are used for expressing complex proteins. Automation can facilitate the handling of insect cell cultures, infection with baculovirus, and protein expression.
- CHO Cells Choinese Hamster Ovary are widely used for the production of recombinant proteins due to their ability to perform post-translational modifications similar to human cells.
- Automated systems can manage the culture of CHO cells in bioreactors, transfection with expression vectors, and subsequent protein purification steps.
- the reconstituted PURE System utilizes purified components of the translation machinery extracted from cells, allowing protein synthesis without the need for intact cells. Automation can streamline the preparation of reaction mixes and monitoring of protein synthesis.
- Automated Liquid Handling Systems Automated liquid handling platforms equipped with robotic arms and various modules can assist in dispensing reagents, conducting transformations, setting up expression cultures, and performing purification steps.
- Automated systems integrated with high-throughput screening assays can accelerate the identification and optimization of protein expression conditions, such as media composition, induction conditions, and temperature regimes.
- Automated systems equipped with sensors and monitoring devices can continuously track parameters such as cell growth, protein expression levels, and culture conditions. This real-time data enables precise control and optimization of the expression process.
- the system 100 may be implemented as an integrated system comprising each of the modules described above and shown in Figure 1.
- the input/output module 101, processing module 103, and sequence analysis module 105 may be implemented as separate software modules in a standalone computing device comprising one or more processors and one or more memory modules, the computing device being communicatively coupled to a separate standalone protein expression module 103.
- the separate software modules may be configured to communicate via one or more software interfaces such as Application Programming Interfaces (APIs).
- the one or more processors may be multi-threaded processors, and may comprise one or more Computer Processing Units (CPUs) and/or Graphics Processing Units (GPUs).
- the sequence analysis module 105 may be a separate sequence analysis module implemented on a separate computing device such as a cloud computing server, and communicatively coupled with the processing module 103 via a software interface such as an API.
- each of the input/output module 101, processing module 103, and sequence analysis module 105 are implemented on separate computing devices communicatively coupled together over a network such as a Local Area Network (LAN), or Wireless Local Area Network (WLAN).
- Figure 2 illustrates a method 200 for analysing protein variants in accordance with an embodiment of the present invention. The steps of method 200 correspond to the functions described above performed by components of the system 100 of Figure 1.
- the method 200 comprises receiving a plurality of input sequences.
- the receiving a plurality of input sequences comprises receiving a plurality of input sequences, each sequence coding for a variant of a target protein as described above in relation to the input/output module 101.
- the method 200 comprises predicting 202a the folded protein structure of the coded protein for each received input sequence, and determining 202b a structural similarity between the predicted folded protein structure of each received input sequence and the target protein.
- the predicting 202a and determining 202b are performed as described above in relation to the processing module 103 and sequence analysis module 105.
- the method 200 comprises identifying one or more preferred input sequences.
- the one or more preferred input sequences may be identified as those corresponding to the protein variants whose predicted folded protein structure is most structurally similar to that of the target protein, as described above in relation to the sequence analysis module 106 and processing module 103.
- the preferred input amino acid sequence may be converted to a nucleic acid sequence suitable for expression.
- the nucleic acid sequence may be codon optimised for the desired expression system.
- the method 200 comprises expressing the coded variants of each identified preferred input sequence.
- the expressing 204 is performed in a cell-free system, and may optionally be performed in a digital microfluidic device, as described above in relation to the protein expression module 106.
- the method 200 comprises determining one or more optimal input sequences based on the expressed coded variants.
- the determining 205 may be performed by calculating one or more assay metrics, by performing one or more protein assays on the expressed preferred input sequences, as described above in relation to the protein expression module 106.
- a protein "variant” may include known isoforms or orthologs, length variants derived from the original sequence with a number of amino acids truncated at either or both of the ends or at some predetermined position in the sequence.
- the location of the truncation position may depend on various factors including, but not limited to, conserving regions, domains and other positions (e.g. catalytic residues important for enzymatic function), and other computed conservation or disorder scores.
- DNA sequences refers to a succession of bases (nucleotides) signified by a set of four different letters that indicate the order of nucleotides forming alleles within a DNA molecule or an RNA molecule. DNA sequences may also be known to the skilled person as, and referred to as "nucleic acid sequences”.
- the input/output module 101, processing module 103, and sequence analysis module may be implemented as functional components within a single device.
- the sequence analysis module 105 may be provided as a separate functional module, for example being stored on a separate device such as a cloud computing server.
- the invention described herein may be embodied in whole or in part as a method, a data processing system, or a computer program product including computer readable instructions. Accordingly, the invention may take the form of an entirely hardware embodiment or an embodiment combining software, hardware and any other suitable approach or apparatus.
- the computer readable program instructions may be stored on a non-transitory, tangible computer readable medium.
- the computer readable storage medium may include one or more of an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk.
- RAM random access memory
- ROM read-only memory
- EPROM or Flash memory erasable programmable read-only memory
- SRAM static random access memory
- CD-ROM compact disc read-only memory
- DVD digital versatile disk
- memory stick a floppy disk.
- Exemplary embodiments of the invention may be implemented as a circuit board which may include a CPU, a bus, RAM, flash memory, one or more ports for operation of connected I/O apparatus such as printers, display, keypads, sensors and cameras, ROM, a communications subsystem such as a modem, and communications media.
- a circuit board which may include a CPU, a bus, RAM, flash memory, one or more ports for operation of connected I/O apparatus such as printers, display, keypads, sensors and cameras, ROM, a communications subsystem such as a modem, and communications media.
- processing means may correspond to any suitable processing device.
- the processing means may be any of a computer processor, graphics processor, programmable logic device, microprocessor, or any other suitable device. It will be appreciated that the processing means may comprise a plurality of processing devices, and may be a combination of different processing devices such as those described above.
- Expression of the protein of interest may be monitored, for example using a split fluorescent protein.
- a method for the real-time monitoring of in-vitro protein synthesis comprising a. In-vitro transcription and translation of a protein of interest fused to a peptide tag; and b. monitoring the presence of the peptide tag using a further polypeptide which in the presence of the peptide tag produces a detectable signal.
- Disclosed herein is a method for the monitoring of cell-free protein synthesis in a droplet on a digital microfluidic device comprising a. cell-free transcription and translation of a protein of interest fused to a peptide tag; and b. monitoring the presence of the peptide tag using a further polypeptide which in the presence of the peptide tag produces a detectable signal.
- Disclosed herein is a method for the monitoring of cell-free protein synthesis in a droplet on a digital microfluidic device comprising a. cell-free transcription and translation of a protein of interest fused to a ccGFPn peptide tag; and b. monitoring the presence of the peptide tag using a further ccGFPi-io polypeptide which in the presence of the ccGFPn peptide tag produces a detectable signal.
- the detectable signal may be for example fluorescence or luminescence.
- the detectable signal may also be caused by the binding or docking of a ligand to the complemented oligopeptide, peptide, or polypeptide tag fused to the protein of interest.
- the detectable signal may also be caused by the binding of the polypeptide to the protein of interest fused to a His-tag.
- Any in-vitro transcription and translation may be used, for example extract-based systems derived from rabbit reticulocyte lysate, human lysate, Chinese Hamster Ovary lysate, a wheat germ, HEK293 lysate, E. coli lysate, yeast lysate.
- the in-vitro transcription and translation may be assembled from purified components, for example a system of purified recombinant elements (PURE).
- purified components for example a system of purified recombinant elements (PURE).
- the in-vitro transcription and translation may be coupled or uncoupled.
- the peptide tag may be one component of a fluorescent protein and the further polypeptide a complementary portion of the fluorescent protein.
- the fluorescent protein could include sfGFP, GFP, eGFP, deGFP, frGFP, eYFP, eBFP, eCFP, Citrine, Venus, Cerulean, Dronpa, DsRED, mKate, mCherry, mRFP, FAST, SmURFP, miRFP670nano.
- the peptide tag may be GFPn and the further polypeptide GFPi-io.
- the peptide tag may be one component of sfCherry.
- the peptide tag may be sfCherryn and the further polypeptide sfCherryi-io.
- the peptide tag may be CFASTn or CFASTio and the further polypeptide NFAST in the presence of a hydroxybenzylidene rhodanine analog.
- the peptide tag may be ccGFPn and the further polypeptide ccGFPi-io.
- GFPi-io polypeptide amino acid sequence could be derived from sfGFP:
- the GFPi-io polypeptide amino acid sequence could be further mutated from the sequence above to become brighter more quickly upon complementation.
- the sequence may have a greater than 90 % homology to any sequence mentioned herein.
- the sequence may have a greater than 95 % homology to any sequence mentioned herein.
- the GFPi-io polypeptide amino acid sequence could also be derived from ccGFP, having a greater than 90 or 95% homology to:
- the complementary GFPn peptide amino acid sequence could be the following:
- KRDHMVLLEFVTAAGITGT (SEQ. ID NO: 9)
- Truncations may involve a shortening of up to 5 amino acids from the N terminus, the C terminus or a combination thereof.
- GFPn or GFPi-io can be fused to the protein of interest through an amino acid linker.
- the oligopeptide, peptide, or polypeptide linker can be 0 - 50 amino acids.
- nucleic acid sequences for expressing particular tags.
- Nucleic acid sequences include
- sequences may be repeated one or more times to produce a protein having multiple GFPn domains.
- the sfCherryi-io polypeptide amino acid sequence could be:
- the complementary sfCherryll peptide amino acid sequence could be:
- YTIVEQYERAEGRHSTGG sfCherryll or sfCherryi-io can be fused to the protein of interest through an amino acid linker.
- the oligopeptide, peptide, or polypeptide linker can be 0 - 50 amino acids.
- NFAST polypeptide amino acid sequence could be:
- the complementary CFAST11 peptide amino acid sequence could be:
- NFAST, CFAST11, and/or CFAST10 can be fused to the protein of interest through an amino acid linker.
- the oligopeptide, peptide, or polypeptide linker can be 0 - 50 amino acids.
- the peptide tag may also be one component of a protein that forms a detectable substrate, such as a luminescent or colorigenic substrate.
- the protein could include beta-galactosidase, betalactamase, or luciferase.
- the protein may be fused to multiple tags. For example the protein may be fused to multiple GFPn peptide tags and the synthesis occurs in the presence of multiple GFPi-io polypeptides.
- the GFP may be ccGFP or sfGFP.
- the protein may be fused to multiple ccGFPn peptide tags and the synthesis occurs in the presence of multiple ccGFPi-io polypeptides.
- the protein of interest may be fused to one or more sfCherryn peptide tags and one or more GFPn peptide tags and the synthesis occurs in the presence of one or more GFPi-io polypeptides and one or more sfCherryno polypeptides.
- the protein may be an enzyme, for example a terminal deoxynucleotidyl transferase (TdT) enzyme or a truncated version thereof or the homologous amino acid sequence of a terminal deoxynucleotidyl transferase (TdT) enzyme in other species or the homologous amino acid sequence of Polp, Poip, PoIX, and Pol0 of any species or the homologous amino acid sequence of X family polymerases of any species.
- TdT terminal deoxynucleotidyl transferase
- TdT terminal deoxynucleotidyl transferase
- Protein sequences disclosed herein may be attached to further elements to improve solubility.
- the variant may be attached to one or more solubility enhancing sequences.
- the solubility enhancing sequence may be a peptide sequence or a naturally occurring sequence.
- the solubility enhancing sequence may be selected from for example maltose binding protein (MBP), Small Ubiquitin-like Modifier (SUMO), Glutathione S-transferase (GST) or thioredoxin (TRX).
- MBP maltose binding protein
- SUMO Small Ubiquitin-like Modifier
- GST Glutathione S-transferase
- TRX thioredoxin
- the tags may be attached to either the C or N terminus. Any example of a solubility enhancer may be used.
- a list of possible proteins is shown below. Any sequence selected from the list below may be chosen:
- the purification tag can be a region of amino acid/peptide sequence.
- the affinity binding site can be a region of amino acid/peptide sequences specific to a particular antibody.
- the tag can be attached to the N or C terminus.
- the purification tag can be selected from the list of exemplary peptide affinity binding sites below:
- Isopeptag (TDKDMTITFTNKKDAE) (SEQ ID NO: 33) lanthanide binding tag (LBT) (FIDTNNDGWIEGDELLLEEG) (SEQ ID NO: 34) Myc (EQKLISEEDL) (SEQ ID NO: 35)
- T7tag (MASMTGGQQMG) (SEQ ID NO: 51)
- VSV-tag (YTDIEMNRLGK) (SEQ ID NO: 54)
- the binding moiety tag can be a sub-component of a fluorescent protein.
- the fully assembled protein becomes fluorescent.
- the detector protein contains GFPi-io and the expressed protein tag contains a GFPn peptide, complementation forms fluorescent GFP, allowing simultaneous monitoring and stability evaluation.
- the expressed material can be monitored over time or conditions by monitoring changes in the level of material generating a fluorescent signal.
- Electrokinesis occurs as result of a non-uniform electric field that influences the hydrostatic equilibrium of a dielectric liquid (dielectrophoresis or DEP) or a change in the contact angle of the liquid on solid surface (electrowetting-on-dielectric or EWoD).
- DEP can also be used to create forces on polarizable particles to induce their movement.
- the electrical signal can be transmitted to a discrete electrode, a transistor, an array of transistors, or a sheet of semiconductor film whose electrical properties can be modulated by an optical signal.
- EWoD phenomena occur when droplets are actuated between two parallel electrodes covered with a hydrophobic insulator or dielectric.
- the electric field at the electrode-electrolyte interface induces a change in the surface tension, which results in droplet motion as a result of a change in droplet contact angle.
- an electrowetting force induced by electric field and resistant forces that include the drag forces resulting from the interaction of the droplet with filler medium and the contact line friction (ref).
- the minimum voltage applied to balance the electrowetting force with the sum of all drag forces is variably determined by the thickness-to-dielectric contact ratio of the insulator/dielectric, (t/e) 1/2 .
- t/e 1/2 thickness-to-dielectric contact ratio of the insulator/dielectric
- High voltage EWoD-based devices with thick dielectric films have limited industrial applicability largely due to their limited droplet multiplexing capability.
- the use of low voltage devices including thin-film transistors (TFT) and optically-activated amorphous silicon layers (a- Si) have paved the way for the industrial adoption of EWoD-based devices due to their greater flexibility in addressing electrical signals in a highly multiplex fashion.
- the driving voltage for TFTs or optically-activated a-Si are low (typically ⁇ 15 V).
- the bottleneck for fabrication and thus adoption of low voltage devices has been the technical challenge of depositing high quality, thin film insulators/dielectrics. Hence there has been a particular need for improving the fabrication and composition of thin film insulator/dielectric devices.
- the electrodes (or the array elements) used for EWoD are covered with (i) a hydrophilic insulator/dielectric and a hydrophobic coating or (ii) a hydrophobic insulator/dielectric.
- a hydrophilic insulator/dielectric and a hydrophobic coating or (ii) a hydrophobic insulator/dielectric.
- Commonly used hydrophobic coatings comprise of fluoropolymers such as Teflon AF 1600 or CYTOP.
- the thickness of this material as a hydrophobic coating on the dielectric is typically ⁇ 100 nm and can have defects in the form of pinholes or a porous structure; hence, it is particularly important that the insulator/dielectric is pinhole free to avoid electrical shorting.
- Teflon has also been used as an insulator/dielectric, but it has higher voltage requirements due to its low dielectric constant and the thickness required to make it pinhole free.
- Other hydrophobic insulator/dielectric materials can include polymer-based dielectrics such as those based on siloxane, epoxy (e.g. SU-8), or parylene (e.g., parylene N, parylene C, parylene D, or parylene HT). Due to minimal contact angle hysteresis and a higher contact angle with aqueous solutions, Teflon is still used as a hydrophobic topcoat on these insulator/dielectric polymers.
- EWoD devices suffers from contact angle saturation and hysteresis, which is believed to be brought about by either one or combination of these phenomena: (1) entrapment of charges in the hydrophobic film or insulator/dielectric interface, (2) adsorption of ions, (3) thermodynamic contact angle instabilities, (4) dielectric breakdown of dielectric layer, (5) the electrode-electrode-insulator interface capacitance (arising from the double layer effect), and (6) fouling of the surface (such as by biomacromolecules).
- contact angle saturation and hysteresis which is believed to be brought about by either one or combination of these phenomena: (1) entrapment of charges in the hydrophobic film or insulator/dielectric interface, (2) adsorption of ions, (3) thermodynamic contact angle instabilities, (4) dielectric breakdown of dielectric layer, (5) the electrode-electrode-insulator interface capacitance (arising from the double layer effect), and (6) fouling of the surface (such as by biomacromolecules).
- An electrokinetic device includes a first substrate having a matrix of electrodes, wherein each of the matrix electrodes is coupled to a thin film transistor, and wherein the matrix electrodes are overcoated with a functional coating comprising: a dielectric layer in contact with the matrix electrodes, a conformal layer in contact with the dielectric layer, and a hydrophobic layer in contact with the conformal layer; a second substrate comprising a top electrode; a spacer disposed between the first substrate and the second substrate and defining an electrokinetic workspace; and a voltage source operatively coupled to the matrix electrodes.
- the dielectric layer may comprise silicon dioxide, silicon oxynitride, silicon nitride, hafnium oxide, yttrium oxide, lanthanum oxide, titanium dioxide, aluminium oxide, tantalum oxide, hafnium silicate, zirconium oxide, zirconium silicate, barium titanate, lead zirconate titanate, strontium titanate, or barium strontium titanate.
- the dielectric layer may be between 10 nm and 100 pm thick. Combinations of more than one material may be used, and the dielectric layer may comprise more than one sublayer that may be of different materials.
- the conformal layer may comprise a parylene, a siloxane, or an epoxy. It may be a thin protective parylene coating in between the insulating dielectric and the hydrophobic coating. Typically, parylene is used as a dielectric layer on simple devices. In this invention, the rationale for deposition of parylene is not to improve insulation/dielectric properties such as reduction in pinholes, but rather to act as a conformal layer between the dielectric and hydrophobic layers. The inventors find that parylene, as opposed to other similar insulating coatings of the same thickness such as PDMS (polydimethylsiloxane), prevent contact angle hysteresis caused by high conductivity solutions or solutions deviating from neutral pH for extended hours.
- the conformal layer may be between 10 nm and 100 pm thick.
- the hydrophobic layer may comprise a fluoropolymer coating, fluorinated silane coating, manganese oxide polystyrene nanocomposite, zinc oxide polystyrene nanocomposite, precipitated calcium carbonate, carbon nanotube structure, silica nanocoating, or slippery liquid-infused porous coating.
- the elements may comprise one or more of a plurality of array elements, each element containing an element circuit; discrete electrodes; a thin film semiconductor in which the electrical properties can be modulated by incident light; and a thin film photoconductor whose properties can be modulated by incident light.
- the functional coating may include a dielectric layer comprising silicon nitride, a conformal layer comprising parylene, and a hydrophobic layer comprising an amorphous fluoropolymer. This has been found to be a particularly advantageous combination.
- the electrokinetic device may include a controller to regulate a voltage provided to the individual matrix electrodes.
- the electrokinetic device may include a plurality of scan lines and a plurality of gate lines, wherein each of the thin film transistors is coupled to a scan line and a gate line, and the plurality of gate lines are operatively connected to the controller. This allows all the individual elements to be individually controlled.
- the second substrate may also comprise a second hydrophobic layer disposed on the second electrode.
- the first and second substrates may be disposed so that the hydrophobic layer and the second hydrophobic layer face each other, thereby defining the electrokinetic workspace between the hydrophobic layers.
- the method is particularly suitable for aqueous droplets with a volume of 1 pL or smaller.
- EWoD-based devices shown and described below are active matrix thin film transistor devices containing a thin film dielectric coating with a Teflon hydrophobic top coat. These devices are based on devices described in the E Ink Corp patent filing on "Digital microfluidic devices including dual substrate with thin-film transistors and capacitive sensing", US patent application no 2019/0111433, incorporated herein by reference.
- electrokinetic devices including: a first substrate having a matrix of electrodes, wherein each of the matrix electrodes is coupled to a thin film transistor, and wherein the matrix electrodes are overcoated with a functional coating comprising: a dielectric layer in contact with the matrix electrodes, a conformal layer in contact with the dielectric layer, and a hydrophobic layer in contact with the conformal layer; a second substrate comprising a top electrode; a spacer disposed between the first substrate and the second substrate and defining an electrokinetic workspace; and a voltage source operatively coupled to the matrix electrodes;
- an electrokinetic device including: a first substrate having a matrix of electrodes, wherein each of the matrix electrodes is coupled to a thin film transistor, and wherein the matrix electrodes are overcoated with a functional coating comprising: one or more dielectric layer(s) comprising silicon nitride, hafnium oxide or aluminum oxide in contact with the matrix electrodes, a conformal layer comprising parylene in contact with the dielectric layer, and a hydrophobic layer in contact with the conformal layer; a second substrate comprising a top electrode; a spacer disposed between the first substrate and the second substrate and defining an electrokinetic workspace; and a voltage source operatively coupled to the matrix electrodes;
- electrokinetic devices as described may be used with other elements, such as for example devices for heating and cooling the device or reagent cartridges for the introduction of reagents as needed.
- the device can be an active-matrix thin film transistor (AM-TFT) based device.
- AM-TFT active-matrix thin film transistor
- an active-matrix thin film transistor (AM-TFT) device having a substrate bearing a plurality of electrodes, the device comprising multiple fluidic inlet ports on at least two sides of the device, wherein the inlet ports on each side of the device are evenly spaced and wherein the device is connected to a syringe pump.
- A-TFT active-matrix thin film transistor
- the device may comprise two substrates, wherein at least one substrate has a plurality of electrodes, and the two substrates define parallel plates that are separated by a spacer to define a volume.
- the fluidic entry may come via holes in the upper plate or through the spacer.
- the entry holes may be in the top substrate.
- the plurality of electrodes may be on a bottom substrate.
- the top substrate be of glass or polymer and may have a thickness ranging from 0.5 mm to 20 mm.
- the spacer may comprise an adhesive with beads of a defined size distribution.
- the spacer may comprise a polymer material of a defined thickness.
- the spacer may comprise glass, in which case the layers can be fused together.
- the spacer gap and therefore height of fluid in the device may be between 50 microns and 250 microns.
- the spacer gap and therefore height of fluid in the device may be between 100 microns and 150 microns.
- the filler liquid may be moved via an automated manner, or may be moved under gravity. A hydrostatic head of pressure can be used to move the liquid within the device.
- the wells are at least partially filled with filler fluid before the aqueous reagents are loaded.
- the filler fluid may be less dense than the aqueous phase such that the aqueous phase sinks in the wells. Alternatively the aqueous phase may sit above the filler fluid, in which case all the filler fluid must be withdrawn from the wells in order to enable entry of the aqueous fluid.
- the device may be connected to a pump, for example a syringe pump, a peristaltic pump, a disc pump, a diaphragm pump, or a pneumatic pump.
- the pump enables filling of the device with filler liquid in an automated manner. Once filled, the pump enables partial withdrawal of the filler fluid to create a negative pressure in the device which draws in reagents from the wells.
- the filling and withdrawal of fluid may be performed in an automated manner to allow largely 'hands-free' loading of the aqueous reagents.
- An automated filler liquid filling and withdrawal method may be integrated into an instrument that provides other functions relating to the digital microfluidic device, including heating, cooling, optical, sensing, mechanical, and magnetic functions.
- the wells of the device may be at 90 degrees to each other.
- the inlets may be at 180 degrees to each other.
- the inlets may be on 4 sides of the device. Each side may have at least 4, 8 or 12 ports. Each side may have 8 ports.
- the device may have 4 sets of 8 ports. The number of ports may vary on different sides of the device, for example one side may have 8 ports and one side 4 ports.
- the device may have 8 ports on 3 sides and 16 ports on a fourth side.
- the ports may be offset to give multiple rows of linear ports on one side, for example a first and second row where the second row is behind by offset from the first row such that the source liquid can flow between the ports of the first row.
- the rows may be a zig-zag fashion.
- the pitch between inlet ports may be 9 mm.
- the pitch between inlet ports may be 4.5 mm.
- the inlet ports have a pitch of 4.5 mm or a multiple of thereof. This would cover 24 well, 48 well, 96 well, 384 well ports.
- the pitch of the ports may be the same on each side of the device, or may be different sized. In this context the pitch refers to the distance between the centre of each inlet.
- the volume of aqueous reagents loaded per inlet port may be between 1 microlitre and 50 microlitres.
- the volume may be between 1 microlitre and 20 microlitres.
- the aqueous liquid may be introduced to the wells by a pipette, a multichannel pipette, a syringe, a blister pack, an acoustic dispenser, or a robotic liquid handler.
- the aqueous liquids may be loaded simultaneously from multiple wells, which may be on the same side or multiple sides of the devices.
- Each well is a separate liquid, and can be the same or different to the contents of the aqueous volume in other wells.
- the volume of aqueous liquid loaded in each port can be the same or can be different.
- the automated filling and/or withdrawing of filler fluid may be controlled by software.
- the device may be part of a larger instrument system that provides environmental control such as temperature control or light control and may have analytical capabilities such as optical systems for fluorescence or luminescence assay detection.
- the location of the aqueous layer is controlled by the actuation of electrodes to form reservoirs in defined areas.
- a plurality of electrodes is actuated to control the location of the aqueous liquid once it has been drawn onto the substrate bearing a plurality of electrodes. Multiple reservoirs may be formed on the device.
- FIG. 5 shows the experimental result from 24 different proteins expressed in a reconstituted cell-free protein synthesis system in droplets on an electrowetting on dielectric (EWoD) device.
- Each construct contains a GFPn tag.
- the GFPi-io detector species is present from the start of expression.
- the rows marked Endpoint shows the fluorescence signal from 10 hours expression in the absence of GFPi-io detector species followed by 5 hours complementation with the GFPi-io detector species.
- This experiment showed significant differences between expression/complementation in Screen BioInk compared to Endpoint detection.
- the detected protein clusters formed after expression only with endpoint detection mean that the protein aggregated after expression, thereby lowering the soluble yield.
- BRCT wild type terminal transferase
- the "BRCA-1 C-terminal (BRCT) domain” refers to the C-terminal domain of a breast cancer susceptibility protein. This domain is found predominantly in proteins involved in cell cycle checkpoint functions responsive to DNA damage, for example as found in the breast cancer DNA-repair protein BRCA1.
- the domain is an approximately 100 amino acid tandem repeat, which appears to act as a phospho-protein binding domain.
- the BRCT domain is present in all TdT sequences in the region from amino acid residues 1-130.
- Each of the TdT structures shown in the left column of Figure 6 has a long unstructured region at the C terminus corresponding to the BRCT domain. Expression of the wild type proteins ( Figure 7) gives a low level of soluble expression.
- FIG. 6 Removal of the BRCT domain gives the structures shown in the right column of Figure 6. The long unstructured region has been removed. The structure of the central active region is not altered by the deletion.
- Figure 7 shows expression levels of the BRCT deleted constructs. Removal of the unstructured regions flanking the BRCT domain and the BRCT domain itself increased expression and yields of the TdT ortholog candidate TdT proteins.
- Aim generate length variants of 9°N DNA polymerase based on domain prediction, structure data and AlphaFold predictions for expression screening.
- Suitable nucleic acid templates were prepared to allow expression using an E. coli derived expression system.
- the full length nucleic acid was transformed into a vector suitable for growth in NEB 5-alpha Competent E. coli cells, which were grown overnight at 37 °C.
- the full length nucleic acid material was amplified using colony PCR to obtain the full length template.
- Varying length nucleic acid templates were prepared by further PCR amplifications using suitable primers.
- the PCR products corresponded to the expected MW of the templates and showed as single gel bands.
- the samples were purified using the GeneJet PCR purification kit. Further nucleic acid amplification reactions were performed in order to add the Nuclera adaptor sequences required for protein expression.
- Templates for expression were prepared and purified using Nuclera's commercially available eGeneTM preparation kit and instructions therein. Adapter sequences add an optional solubility tag or no solubility tag depending on the sequence of the adapter.
- the 192 datapoints are shown in table 1 and plotted as a heatmap in Figure 9.
- the data shows a clear correlation between molecular weight and expression, as expected.
- expression of full length transcripts are typically less than 3 pM, whereas the truncations are greater than 4 pM.
- the best truncation conditions using manganese give expression greater than 5 pM (5.21 pM) for the truncation with FH8 tags, compared to the full length sequence having FH8 tags expressing at 1.4 pM (1.44 pM).
Landscapes
- Engineering & Computer Science (AREA)
- Physics & Mathematics (AREA)
- Life Sciences & Earth Sciences (AREA)
- Spectroscopy & Molecular Physics (AREA)
- Health & Medical Sciences (AREA)
- Evolutionary Biology (AREA)
- Chemical & Material Sciences (AREA)
- Biophysics (AREA)
- Theoretical Computer Science (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Biotechnology (AREA)
- Library & Information Science (AREA)
- General Health & Medical Sciences (AREA)
- Medical Informatics (AREA)
- Biochemistry (AREA)
- Molecular Biology (AREA)
- Crystallography & Structural Chemistry (AREA)
- Peptides Or Proteins (AREA)
Abstract
Provided herein are methods for analysing protein sequences, the method comprising receiving (201) input data comprising a plurality of input sequences, each of the input sequences coding for a target protein; for each of the input sequences in the plurality of input sequences, predicting (202a) a folded protein structure of the coded protein of the DNA sequence, and determining (202b) a structural similarity between the predicted folded protein structure and the target protein, or availability of a binding or detection tag; identifying (203) one or more preferred input sequences based on the determined structural similarities; expressing (204) the coded variant of each of the identified one or more preferred input sequences in a cell-free system; and determining (205) one or more optimal DNA sequences based on the expressed one or more coded variants.
Description
SYSTEM AND METHOD FOR PROTEIN SEQUENCE SCREENING
FIELD OF THE INVENTION
Provided herein are systems and methods for protein sequence screening. In particular, systems and methods are provided that determine one or more preferred candidates from a list of protein sequences, express the one or more preferred candidates to yield one or more candidate proteins, and determine an optimal protein of the one or more candidate proteins.
BACKGROUND TO THE INVENTION
It is common practise in the field of protein synthesis to search for and try to synthesise variants of a target protein. It may be desirable to produce a large number of variants of a target protein. However, it is important to ensure that each variant appears similar to the native conformation of the target protein to have confidence that the edits made to the target protein are not impacting the native conformation, or at least to understand the nature of the variation if there is a variation.
Determining the degree of similarity between the protein variants and the target protein can be time-consuming and expensive, particularly if every candidate protein variant is to be expressed and assayed to determine its structure.
Prior art protein structure prediction software includes the AlphaFoldR™ software. AlphaFold employs a deep learning architecture, specifically a deep residual convolutional neural network (CNN), to predict protein structures. Deep learning models, particularly CNNs, have shown great promise in various domains due to their ability to learn hierarchical features from data.
The model is trained on a large dataset of known protein structures, which serve as examples of how amino acid sequences fold into three-dimensional structures. These structures are obtained from publicly available databases like the Protein Data Bank (PDB). During training, the model learns to predict the distance and orientation between pairs of amino acids in a protein sequence. One of the key innovations of AlphaFold is its method for predicting inter-residue distances between pairs of amino acids in a protein sequence. Instead of directly predicting distances, AlphaFold predicts a probability distribution over the distances between pairs of residues. This approach allows the model to capture the uncertainty inherent in distance predictions.
Using the predicted distance distributions, AlphaFold employs optimization algorithms, such as gradient descent, to infer the most probable spatial arrangement of amino acids in the protein sequence. This process involves determining which pairs of residues are in close proximity and assembling them into a three-dimensional structure. After the initial structure assembly, AlphaFold refines the predicted structures using optimization techniques to improve their accuracy. This refinement step helps to correct any errors or inaccuracies in the initial predictions. AlphaFold also provides estimates of the uncertainty associated with its predictions. This is crucial because protein structure prediction inherently involves uncertainty due to the complexity of protein folding and the limitations of available data. The final output of AlphaFold is a predicted three-dimensional structure of the protein, represented in terms of atomic coordinates. This structure can be visualized and analyzed to gain insights into the protein's function and interactions.
BioRxiv, 2022, Goverde et al., "De novo protein design by inversion of the AlphaFold structure prediction network" presents a research study in which the AlphaFold v2 (AF2) neural network is inverted, with the aim of testing whether the AF2 model has learned the principles of protein folding sufficiently for de novo design of protein structures. With an inverted AF2 network, a structure is input to the inverted model as a structural loss, and backpropagated through the network to produce as output amino acid sequences compatible with a target structure. Accordingly, the aim of Goverde at al is to produce a protein sequence based on a target structure, rather than to determine an optimal sequence, or subset of sequences, for a variant of a target protein. Accordingly, there is no disclosure in Goverde of identifying preferred protein sequences from a plurality of input sequences based on determined structural properties, nor of expressing and characterising the preferred sequences to determine one or more optimal sequences. The method of Goverde further differs from the present invention in that a single sequence, which does not code for a variant of a target protein, is input into the AF2 model as an initialisation step.
BMC biology, vol. 15, no. 100, 2017, Osterle et al., "Sequence-based prediction of permissive stretches for internal protein tagging and knockdown" describes a proposed approach for predicting permissive stretches (PSs) in proteins based on the identification of length-variable regions in homologous proteins. Oesterle describes testing the predicted PSs by expressing proteins in a cell-free system.
US 2021/043272 Al describes a method of predicting protein structure and properties based on an input protein sequence. The method comprises determining correlations between differences of protein sequences and changes to structural features and biophysical properties and using the determined correlations to generate models that can utilize protein sequences to predict structural features and biophysical properties of proteins. However, there is no disclosure of identifying preferred sequences based on determined structural properties, nor of expressing and characterising protein variant sequences of identified preferred sequences to determine optimal protein sequences.
SUMMARY
Described herein is the use of protein structure prediction software to pre-select a subset of input sequences for protein expression. Embodiments of the invention seek to provide a method for analysing protein variants (for example length variants, isoforms, polymorphisms, orthologs, homologs or the structure of binding sites), the method comprising receiving input data comprising a plurality of amino acid sequences or their corresponding DNA sequences, each of the DNA sequences coding for a variant of a target protein; for each of the input sequences in the plurality of sequences, predicting a folded protein structure based on the sequence of the expressed protein of the input sequence, determining one or more structural properties of the predicted folded protein structure; identifying one or more preferred input sequences based on the determined structural properties; expressing the coded variant of each of the identified one or more preferred input sequences in a cell-free or cell-based system; and characterizing the protein variant sequences based on the expressed one or more coded variants.
Protein expression usually benefits if the expressed protein can be purified from the expression system. Most purification systems rely on affinity binding to a tag attached to the expressed protein. Protein structure prediction can be used to determine if a particular expressed protein has an available binding tag, or whether the binding tag is likely to be folded internally within the protein structure. Proteins where the binding tag is not externally available are not suitable candidates for expression. Embodiments of the invention seek to use protein structure prediction to triage input sequences for protein expression and characterisation.
Further embodiments of the invention seek to address the above problem by providing a system for analysing protein variants, the system comprising an input module configured to receive input data comprising a plurality of input sequences, each of the sequences coding for a variant of a target protein; a processing module configured to receive the input data from the input
module, and provide the input data to a sequence analysis module; a sequence analysis module configured to, for each of the input sequences in the plurality of sequences predict a folded protein structure of the coded protein of the input sequence, and determine a structural similarity between the predicted folded protein structure and the target protein; the processing module being further configured to receive the determined structural similarities from the sequence analysis module, and identify one or more preferred input sequences based on the determined structural similarities; a protein expression module configured to express the coded variant of each of the identified one or more preferred input sequences in a cell-based or cell- free system; the processing module being further configured to determine one or more optimal input sequences based on the expressed one or more coded variants.
Further embodiments of the invention seek to express a protein having the shortest functionally equivalent domain. Expression of proteins, particularly large proteins can be limited by the length of the nucleic acid template and the amounts of reagents required for expression. Structure prediction can be used to identify regions of the amino acid sequence which can be removed without impacting the function of the expressed protein. The undesired regions may include disordered regions or regions which may promote aggregation. Where the reagents such as amino acids are resource limited, expression of shorter chains gives a higher concentration (for example half as long may be twice the number of molecules). Thus structure prediction may be used to design nucleic acid or protein sequence inputs for protein expression.
BRIEF DESCRIPTION OF THE DRAWINGS
An embodiment of the invention will now be described, by way of example only, and with reference to the accompanying drawings, in which:
Figure 1 shows a method for screening input sequences in accordance with an embodiment of the present invention.
Figure 2 shows a schematic of a system for screening input sequences in accordance with an embodiment of the present invention.
Figure 3 shows a protein of interest with a peptide tag (detection tag) in the presence of a protein that binds to the peptide tag (detector protein). In the case that the detection tag and detector protein generate a detectable signal upon binding, then the quantity of the protein of interest can be measured.
Figure 4 shows a cartoon representation of protein expression of different proteins having varying soluble yields and levels of aggregation. Some conditions give no expression (dark). Some conditions give a high level of soluble protein (uniform intensity white squares). Some conditions give protein aggregates, meaning the expressed proteins are insoluble and clump together (white spots in dark background). Some give a mix of soluble and aggregated proteins (white spots in a grey background). Thus the level of soluble expression and level of aggregation can be determined by adding a detector species to the expressed proteins.
Figure 5 shows the experimental result from 24 different proteins expressed in a reconstituted cell-free protein synthesis system in droplets on an electrowetting on dielectric (EWoD) device. Each construct contains a ccGFPn tag. In the rows marked Screen, the ccGFPi-io detector species is present from the start of expression. The rows marked Endpoint shows the fluorescence signal from 10 hours expression in the absence of ccGFPi-io detector species followed by 5 hours complementation with the ccGFPi-io detector species. This experiment showed significant differences between expression/complementation in Screen BioInk compared to Endpoint detection. The detected protein clusters formed after expression only with endpoint detection mean that the protein aggregated after expression, thereby lowering the soluble yield. It is evident from the image the presence of speckles for several constructs, indicating the likely presence of protein aggregates. The level of aggregation enables identification of conditions which are worthy of further testing, and identification of conditions having high level of aggregated protein from which further purification is unlikely to give material. Protein structure prediction can be used to screen sequences prior to expression, for example to identify those which may aggregate, or for which the detector tag is less accessible for binding.
Figure 6 shows AlphaFold structures for 8 terminal transferase (TdT) enzymes. Four wild type sequences (Coho Salmon, Comon Starling, Tasmanian Devil, California Sea Lion) each have multiple unstructured regions (left side structures), including the long BRCT domain. The right side structures show the predictions with the terminal 'BRCT' domain removed.
Figure 7 shows the removal of the unstructured regions flanking the BRCT domain and the BRCT domain itself increased expression and yields of the TdT ortholog candidate TdT proteins.
Figure 8 shows structure predictions for a variety of length truncated 9°N polymerases. The structure of the shortened chains having the functional core unchanged were selected for expression.
Figure 9 shows date from an expression screening report on various length variants of a DNA polymerase. Structure prediction was used to guide the choice of constructs for expression.
DETAILED DESCRIPTION OF THE INVENTION
The present invention solves the problems set out above by providing a system and method that receives input data comprising a plurality of input sequences that each code for a protein variant of a target protein and/or an indication of the target protein having an accessible binding tag or detection tag, predicts the folded protein structure of each variant sequences, determines a structural property of the folded protein structure of the variant sequence, expresses the chosen variant input sequences in a cell-free or cell-based system, and determines the optimum variant sequence or sequences based on the expressed input sequences and optionally the availability of the binding or detection tags. The system and method of the present invention may be fully automated, and thereby operate substantially without human intervention. The present invention therefore allows a user to determine which variant, or variants, of a target protein are candidates for protein expression and exhibit one or more preferred characteristics, including optionally the ability to detect or purify the protein, without requiring the user to express and assay every possible input protein variant.
Disclosed is a method for analysing protein sequences, the method comprising: receiving input data comprising a plurality of input sequences, each of the input sequences coding for a variant of a target protein; for each of the input sequences in the plurality of input sequences: predicting a folded protein structure of the coded protein of the input sequence; and determining one or more structural properties of the predicted folded protein structure; identifying one or more preferred input protein sequences based on the determined structural properties; taking nucleic acid input sequences for expressing the input protein sequences; expressing the coded variant of each of the identified one or more preferred input protein sequences; characterizing the protein variant sequences based on the expressed one or more coded variants; and determining one or more optimal input sequences based on the expressed one or more coded variants.
Protein variants may include for example length variants, isoforms, polymorphisms, orthologs, homologs or the structure of binding sites. Variants of a protein may include N-terminal truncations, C-terminal truncations, internal deletions, one or more point mutations, orthologs and homologs. Truncations and deletions may be used to remove regions of a protein expected to negatively impact protein expression, folding or stability. For instance, highly disordered regions may be truncated from the termini of a protein. Orthologs are proteins that diverged as a result of speciation. For example, the protein TYK2 (Non-receptor tyrosine-protein kinase TYK2) exists in humans (Uniprot P29597) and mice (Q.9R117). The two proteins are said to be orthologs. Expressing orthologs as protein variants may help obtaining proteins and is also an essential component of the drug discovery process due to the use of ortholog animal models. Commonly used orthologs include mice, dogs, zebrafish, rats, and non-human primates.
The invention may choose similar and dissimilar isoforms. For example from a list of isoforms, some are predicted to be similar and some dissimilar. Similar and dissimilar inputs may be chosen in order to screen for diversity.
The invention may be used for lead compound toxicity screening. For example to screen a diversity of protein targets for binding to particular compounds. Thus a range of predicted input structures may be chosen for expression and binding, which a view to seeing if a particular lead compound is target selective or binds to a wide variety of protein structures, a likely indicator of toxicity.
The invention may be used to identify or characterise molecular binding sites. The ability to predict structures allows identification of mutations which alter a particular binding site. Thus protein sequences can be expressed to perform a binding site screen, thus making inactive mutants. The binding sites may be small molecule binding sites, antibody binding sites or sites involved in protein-protein interactions. An example of binding site determination is Wang et. al. (Identification of Drug Binding Sites and Action Mechanisms with Molecular Dynamics Simulations Curr Top Med Chem. 2018;18(27):2268-2277).
Use of the invention may involve for example 6 isoform structures predicted through this pipeline. Through public or proprietary prior knowledge (i.e., structurally determined binding site), the binding pocket is identified. The structural similarity of the binding sites are determined. The user may use programs like autodock or GoLD to dock ligands. The user then
chooses all 6 proteins to be expressed/characterized/purified to empirically validate if their prediction of binding strength between the 6 targets held true.
Alternatively, if the user is trying a multi-target small molecule drug discovery approach, the user understands that global RMSD is quite a crude comparison and wants to compare 6 isoform drug binding sites based on prior knowledge. The user wants to be able to target 5 out of 6 is- forms but does not want to target 1 of them. The user needs to know if the isoforms are structurally different at that binding pocket. Can they use 1 isoform to target the 5 similar isoforms? Is their hypothesis correct that they can counter select against one of the isoforms through binding to a specific site? Are the LEC flanks they're appending going to influence the site that differentiates 1 isoform from the other 5? The selection process helps the user design which isoforms to make and how the construct looks.
As described herein, "protein folding" refers to the process by which a sequence of amino acids, such as a sequence of amino acids of a particular protein, folds into a specific, typically complex, three-dimensional structure. The particular folded three-dimensional structure of the amino acid sequence (the protein) effects its biological characteristics, including the existence of, and/or location of molecular binding sites.
The molecular binding sites may be binding pockets for therapeutic agents. Therapeutic agents may, non-exhaustively, include small molecule drugs, proteins, and nucleic acids. It is important that the existence, location, and structure of molecule binding sites reflects the native, naturally occurring target protein. As such, when determining the structural similarity between variants, there may be more weight given to the similarity of certain regions of the protein, for instance known molecular binding sites may be weighted highly compared to less structured or unstructured domains of the target protein.
The molecular binding sites may be part of the protein of interest or may be flank sequences appended to a protein of interest (POI). The POI may have a known structure or an unknown structure. The POI having flank sequences where the flanks change the structure of the POI may not be ideal candidates to obtain the POI. For example the flank sequences may cause the POI to unfold. For example the flank sequences may cause the POI to adopt a different, unnatural, or non-native structure. Alternatively the flank sequences may have a binding or detection tag which is not available for binding when the protein is expressed. Protein structure prediction can be used to decide the optimal candidates for a particular set of conditions, and the
conditions where a protein is most likely to be obtained. Thus the input sequences and expression conditions can be pre-selected and the sequences and conditions where protein is unlikely to be obtained in measurable form, due to for example lack of stability or binding affinity may be down-selected. Alternatively one could use this to optimise linker length. If linker is too short, then there may be negative effects on fusion tag - POI interaction. If linker is too short, the affinity tag may be inaccessible.
The flank sequences may include elements required for transcription and translation, such as for example a promoter region such as a T7 promoter biding site, a ribosome binding site, start codons and stop codons. The flank sequences may optionally also include elements such as solubility tags, purification tags or detection tags.
Figure 1 shows a schematic functional component diagram of the system 100 according to an embodiment of the present invention. The system 1 comprises an input/output module 101, user interface 102, processing module 103, sequence analysis module 105, and protein expression module 106. It will be appreciated that the system 100 according to the present invention need not necessarily include every module shown in Figure 1, and may include other modules that are not illustrated in Figure 1.
The input/output module 101 is configured to receive input data from a user. The input data comprises a plurality of input sequences, either as amino acid sequences or DNA sequences, where each DNA sequence codes for a variant of a target protein. DNA sequences are familiar to the skilled person, and may be in the form of a character string where each character signifies a specific nucleotide. An example of a DNA sequence is: AAACAAGTGGGT. Such a sequence codes for four amino acids (KQVG) following triplet codon translation.
The input sequence may be a DNA or protein sequence. The DNA sequence may have any suitable length to code for a protein variant, the protein variant optionally having additional flanks such as those needed for protein expression. For example, a protein variant formed of 1000 amino acids may have a DNA sequence of 3000 bases. It will be appreciated that the DNA sequences need not necessarily be represented as a character string and may be represented in any other suitable format, including vector format, alpha-numeric, binary, hexadecimal etc.
In some embodiments, the input/output module 101 may be configured to receive an indication of a target protein, where the plurality of received input sequences code for protein variants of
said target protein. The input/output module 101 may be configured to receive the indication of the target protein in addition to the plurality of input sequences. In alternative embodiments, the input/output module 101 may be configured to receive an indication of a target protein instead of the plurality of input sequences. In such embodiments, the system 100 may calculate, using the processing module 103, a plurality of input sequences that code for protein variants of the received target protein.
In alternative embodiments, the input/output module 101 may be configured to receive a plurality of amino acid sequences. Amino acid sequences are also known to the skilled person and are signified by a character string where each character represents a particular amino acid. DNA and amino acid sequences can be easily translated back and forth through the use of genetic code and codon tables. It should be noted that while a DNA sequence typically codes for a single amino acid sequence (if in frame from a single start codon to a single stop codon), an amino acid sequence could be reverse translated into many different DNA sequences. For example, Alanine (A) can be encoded as GCU, GCC, GCA, or GCG.
The input/output module 101 may receive input data via the user interface 102. In some embodiments the user interface 102 may comprise a Graphical User Interface. A Graphical User Interface may be provided in the form of a widget embedded in web site, as an application for a device, or on a dedicated landing web page. Computer readable program instructions for implementing the Graphical User Interface may be provided to a user device from a remote computer readable storage medium via a network connection, for example, the Internet, a local area network, a wide area network and/or a wireless network. In other embodiments, the input/output module 101 may receive input data via a network connection.
One possible implementation of the user interface is a web-based application accessible via a web browser which communicates inputs and outputs with a cloud based server via a REST API.
The processing module 103 is configured to receive the input data from the input/output module 101. The processing module is communicatively coupled to the input/output module 101, the sequence analysis module 105, and the protein expression module 106. In some embodiments, the processing module 103 receives input data from the input/output module 101 comprising a plurality of input sequences, each input sequence coding a protein variant of a target protein. The processing module 103 may further receive input data comprising an
indication of the target protein, for example a further input sequence that codes for the target protein.
In other embodiments, the processing module 103 receives a DNA sequence that codes for a target protein, and is configured to generate a plurality of DNA sequences that each code for a different variant of the target protein. Alternatively the processing module 103 receives a protein sequence the codes for a target protein, and is configured to generate a plurality of protein sequences that each code for a different variant of the target protein. The processing module 103 may convert the desired protein sequences into nucleic acid sequences suitable for expression. The sequences may be codon optimised for the expression system.
A possible implementation of the input/output process is a cloud based publish-subscribe messaging service.
The processing module 103 outputs the received or generated plurality of input sequences to the sequence analysis module 105. A possible implementation of the processing module is a managed cloud-based pipeline service which schedules and runs containerized computation jobs in a direct acyclic graph manner.
The sequence analysis module 105 is configured to receive a plurality of DNA sequences, or amino acid sequences and predict the folded protein structure of the coded protein of each DNA sequence or amino acid sequence. A possible implementation of the sequence analysis module is a multiple sequence alignment module detecting similarities between multiple sequences and their subsequences.
In certain embodiments, the sequence analysis module 105 receives as input data representing a particular input sequence of the plurality of input sequences, and outputs an estimate of the folded protein structure of the protein that the particular input sequence codes for. The folded protein structure is output in the form of a three-dimensional configuration of the atoms of the coded protein of the DNA sequence, when the coded protein has undergone protein folding. The three-dimensional configuration may be output in the form of a coordinate vector, the vector having a length equal to the number of amino acids in the coded protein of the input sequence, and each element of the vector corresponding to the three-dimensional cartesian coordinates of the corresponding amino acid of the input sequence. It will be appreciated that the folded protein structure may be output in any other suitable format, for example a
parametric model, or an alternative coordinate system. The folded protein structure may be output in the format known as the Protein Data Bank (PDB) format, which is a standardised form for files containing atomic coordinates.
The sequence analysis module 105 may be configured to receive data signifying the target protein, such as a DNA sequence that codes for the target protein. In such embodiments, the sequence analysis module 105 predicts the folded protein structure of the target protein (known as the natural conformation) and then optionally determines the structural similarity between the target protein and each of the variant proteins coded for by the plurality of input sequences. The structural similarity may be calculated using any known measure of protein structure similarity. In some embodiments, the sequence analysis module calculates the root mean square difference metric (RMSD). If two structures have a small RMSD they may be similar, while if two structures have a large RMSD they may be less similar. The sequence analysis module 105 outputs data indicating, for each input sequence representing a particular protein variant, a value representing the structural similarity between the folded protein coded by that input sequence and the folded target protein.
Methods of protein structure comparison are described for example by Kufareva and Abagyan (Methods Mol Biol. 2012; 857: 231-257), the contents of which are incorporated herein by reference in their entirety. Protein structures may be determined using sequence-dependent vs. sequence-independent methods. Sequence-dependent methods of protein structure comparison assume strict one-to-one correspondence between target and model residues. In sequence-independent methods, structural superimposition is performed independently, followed by the evaluation of residue correspondence obtained from such superimposition. A variety of sequence-independent structural alignment methods have been developed in the field: CE (Shindyalov IN, Bourne PE. Protein Engineering. 1998;11:739-747), DALI (Holm L, Sander C. Journal of Molecular Biology. 1993;233:123-138), DejaVu (Kleywegt GJ, Jones AT. Methods in Enzymology. Academic Press; 1997. pp. 525-545), MAMMOTH (Ortiz AR, Strauss CEM, Olmea O. Protein Science. 2002;11:2606-2621), Structal (Levitt M, Gerstein M. Proceedings of the National Academy of Sciences of the United States of America. 1998;95:5913-5920), FOLDMINER (Shapiro J, Brutlag D. Nucleic Acids Research. 2004;32:W536-W541), KENOBI/K2 (Szustakowski JD, Weng Z. Proteins: Structure, Function, and Bioinformatics. 2000;38:428-440), LSQMAN (Kleywegt GJ. Acta Crystallogr D Biol Crystallogr. 1996;52:842-857), Matras (Kawabata T, Nishikawa K. Proteins. 2000;41:108-122), PrISM (Yang A-S, Honig B. Journal of Molecular Biology. 2000;301:665-678), ProSup (Lackner P,
Koppensteiner WA, Sippl MJ, Domingues FS. Protein Engineering. 2000;13:745-752), SSM (Krissinel E, Henrick K. Acta Crystallographica Section D. 2004;60:2256-2268), and others.
Identification of global versus local similarity represents two orthogonal directions in comparison of protein structures, i.e. structures that are most similar globally may not be the best in terms of local similarity. Flexible or disordered fragments such as long loops and/or termini are often poorly predicted and may significantly compromise the otherwise good similarity between structures. Relative domain movements observed in multi-domain proteins can also contribute to the poor global similarity scores. Focusing on local similarity helps to avoid these issues. Local similarity can be interpreted as a cumulative similarity score for all regions of the protein or, otherwise, can focus on a specific region such as, for example, ligand binding pocket, while ignoring the remaining parts of the protein.
Root Mean Square Deviation (RMSD) is the most commonly used quantitative measure of the similarity between two superimposed atomic coordinates. RMSD values are presented in A and calculated by
where the averaging is performed over the n pairs of equivalent atoms and d, is the distance between the two atoms in the /-th pair. RMSD can be calculated for any type and subset of atoms; for example, Ca atoms of the entire protein, Ca atoms of all residues in a specific subset (e.g. the transmembrane helices, binding pocket, or a loop), all heavy atoms of a specific subset of residues, or all heavy atoms in a small-molecule ligands. The most common way to evaluate the correctness of the docking geometry is to measure the Root Mean Square Deviation (RMSD) of the ligand from its reference position in the answer complex after the optimal superimposition of the receptor molecules.
The sequence analysis module 105 may utilize one or more machine learning models, including but not limited to one or more neural network models. Such neural network models are well
known to the skilled person and comprise a plurality of interconnected nodes. The machine learning models may be deep models that employ multiple levels of machine learning models, such as combinations of multiple neural networks or combinations of neural networks and/or other machine learning models.
An example of a neural network based system that may form a part of or all of the sequence analysis module 105 is the AlphaFold Al system developed by DeepMind and EMBL's European Bioinformatics Institute. Implementational details of the AlphaFold Al system can be found in US 2021/304847 Al, the contents of which are incorporated herein by reference in its entirety.
In certain embodiments, the processing module 103 receives from the sequence analysis module 105 data indicating the structural similarity between the folded protein structures of each of the variant input sequences and the folded protein structure of the target protein. Based on the received data, the processing module determines a set of preferred input sequences that correspond to the most structurally-similar variants to the target protein. The number of preferred input sequences may be selected by the user. Merely by way of example, if 10 input sequences are passed to the sequence analysis module 105, the processing module 103 may select the top 4 most structurally-similar input sequences as preferred input sequences. By way of further example, if 10 input sequences expressing a protein of interest attached to a binding tag are passed to the sequence analysis module 105, the processing module 103 may select the top 4 most likely input sequences to express a protein having an available binding tag as preferred input sequences. By way of further example, if 10 DNA sequences expressing a protein of interest attached to a detection tag are passed to the sequence analysis module 105, the processing module 103 may select the top 4 most likely DNA sequences to express a protein having an available detection tag as preferred DNA sequences. By way of further example, if 10 DNA sequences expressing a protein of interest bearing a detection and purification tag are passed to the sequence analysis module 105, the processing module 103 may select the top 2 most likely sequences to express a structurally-similar protein that has an available purification tag and an available detection tag.
In alternative embodiments, the processing module 103 receives the folded protein structure of each of the input sequences from the sequence analysis module 105, and determines the structural similarity between the folded protein structure of the target protein and that of each input sequence.
The protein expression module 106 is configured to receive data from the processing module 103 indicating the preferred input sequences based on their determined structural similarity to the target protein. The protein expression module 106 is configured to express each of the preferred DNA sequences in a cell-free system. Possible implementations of the Protein Expression module is the Nuclera eProtein Discovery”™ system, an automated liquid handling protein expression system, or a manual protein expression protocol. Any suitable platform for protein expression may be used, for example liquid handling in microtitre plates.
The chosen protein may be expressed in cells or using a cell-free expression system. The cell- free system may be a cell lysate or a reconstituted system.
Cell-free protein synthesis, also known as in-vitro protein synthesis or CFPS, is the production of peptides or proteins using biological machinery in a cell-free system, that is, without the use of living cells. The in-vitro protein synthesis environment is not constrained within a cell wall or limited by conditions necessary to maintain cell viability, and enables the rapid production of any desired protein from a nucleic acid template, usually plasmid DNA or RNA from an in-vitro transcription. CFPS has been known for decades, and many commercial systems are available. Cell-free protein synthesis encompasses systems based on crude lysate (Cold Spring Harb Perspect Biol. 2016 Dec; 8(12): a023853) and systems based on reconstituted, purified molecular reagents, such as the PURE system for protein production (Methods Mol Biol. 2014; 1118: 275- 284). CFPS requires significant concentrations of biomacromolecules, including DNA, RNA, proteins, polysaccharides, molecular crowding agents, and more (Febs Letters 2013, 2, 58, 261- 268).
In some embodiments, the protein expression module 106 performs the cell-free expression of the preferred DNA sequences in a digital microfluidic device. An example implementation of cell-free expression in digital microfluidic devices is described in WO 2022/038353, the entire contents of which are incorporated herein by reference. Microfluidic devices for manipulating droplets or magnetic beads based on electrowetting have been extensively described. Electrowetting is the modification of the wetting properties of a surface (which is typically hydrophobic) with an applied electric field. In the case of droplets in channels the manipulation of droplets can be achieved by causing the droplets, for example in the presence of an immiscible carrier fluid, to travel through a microfluidic channel defined by the walls of a cartridge or microfluidic tubing. Embedded in the walls of the cartridge or tubing are electrodes covered with a dielectric layer each of which are connected to an A/C biasing circuit capable of
being switched on and off rapidly at intervals to modify the electrowetting field characteristics of the layer. This gives rise to the ability to steer the droplet along a given path. As an alternative to microfluidic channel systems, droplets can also be generated and manipulated on planar surfaces using digital microfluidics (DMF). In contrast to channel based microfluidics, DMF utilizes alternating currents on an electrode array for moving fluid on the surface of the array. Liquids can thus be moved on an open-plan device by electrowetting. Digital microfluidics allows precise control over the droplet movements including droplet fusion and separation.
Once the preferred input DNA sequences have been expressed by the protein expression module 106, the protein expression module 106 is configured to determine one or more optimal DNA sequences. The optimal DNA sequences may be determined by calculating one or more assay metrics, the one or more assay metrics being calculated by performing one or more protein assays. Protein assays are well known to the skilled person, and can be used to measure a variety of characteristics of the expressed protein including the amount of protein expressed, the solubility of the protein, protein stability or the purifiability of the protein. For example, where the assay metric is the amount of protein expressed, the protein expression module 106 may perform a known assay such as fluorescence complementation, the Bradford assay, Folin- Lowry assay, or Bicinchoninic Acid (BCA) assay to determine the amount of each variant protein expressed in the cell-free system. The protein expression module 106 may then determine one or more optimal input sequences corresponding to the highest expressing variant proteins.
In alternative embodiments, the determination of the one or more optimum input sequences is performed by the processing module 103, based on assay metric data output from the protein expression module 106.
The determination of structure may be based on computed features such as surface charge, surface hydrophobicity, the presence of unstructured regions, solubility predictions or the predicted melting temperature. For example the elimination of regions having a high surface charge may be beneficial for use in certain expression devices. The removal of unstructured or hydrophobic regions may help to prevent aggregation or improve the level of soluble expression.
Predicting the melting temperature (Tm) of a protein, which represents the temperature at which a protein denatures or unfolds, can be crucial for various biochemical and biophysical studies. Several software tools are available for predicting protein melting temperatures, often based on empirical or computational models, for example:
PROTHERM: PROTHERM is a web server that predicts the thermal stability of proteins and their mutants. It uses an empirical approach based on experimental data from the ProTherm database. Users can input the amino acid sequence or the PDB ID of the protein and obtain predictions for various parameters, including melting temperature.
FoldX: FoldX is a molecular modeling software package that includes modules for predicting protein stability and melting temperature. It utilizes an empirical force field to estimate the free energy change upon protein unfolding. Users can input protein structures in PDB format and calculate various thermodynamic parameters, including melting temperature.
PoPMuSiC: PoPMuSiC (Prediction of Protein Mutant Stability Changes) is a web server for predicting the effects of mutations on protein stability, including melting temperature changes. It employs a statistical potential-based approach to estimate the change in protein stability upon mutation. Users can input protein sequences or structures along with mutations and obtain predictions for stability changes and melting temperature shifts.
ThermoProt: ThermoProt is a web server for predicting the thermal stability of proteins. It integrates various computational methods, including machine learning algorithms and biophysical models, to predict melting temperatures. Users can input protein sequences or structures, and the server provides predictions along with confidence scores.
DUET: DUET (Dynamic Undocking and Energy Transfer) is a computational tool that predicts protein stability changes upon mutations. While its primary focus is on predicting binding affinity changes, it also provides estimates of protein melting temperature changes. DUET combines molecular dynamics simulations with statistical potentials to predict the effects of mutations on protein stability.
In some embodiments, the protein expression module 106 can be an automated protein expression system configured to take one or more protein sequences as input, express each of the one or more protein sequences in a cell or cell-free system, perform one or more protein assays on the expressed one or more proteins to calculate one or more assay metrics, and determine one or more optimal protein sequences based on the calculated assay metrics (or output the calculated assay metrics to the processing module 103). The automated protein expression module then outputs the determined one or more optimal protein sequences, and/or calculated assay metrics, to the processing module. As an automated system the
automated protein expression module performs these steps substantially or entirely without human intervention. As an example, an automated protein expression module may be implemented as a digital microfluidics system communicatively coupled to, or comprised within, the system 100.
Automated Expression Systems:
Automated systems can handle the transformation and cultivation of bacterial and yeast systems for protein expression. For example Escherichia coli (E. coli) is one of the most common organisms used for protein expression. It offers rapid growth and high expression levels. Automated systems can handle the transformation of plasmids into E. coli, induction of protein expression, and subsequent purification steps. Similarly Bacillus subtilis is another bacterium commonly used for protein expression, especially for secreted proteins. Automated systems can handle the transformation and cultivation of B. subtilis strains. Similarly automated systems can handle yeast transformation, growth, and induction for eukaryotic protein expression. Pichia pastoris is a methylotrophic yeast system often used for expressing eukaryotic proteins at high levels. Automation can assist in handling the transformation and cultivation of P. pastoris strains.
Insect Cell Expression Systems:
Baculovirus-lnsect Cell System (BEVS): Insect cells such as Sf9 or Sf21 infected with recombinant baculovirus are used for expressing complex proteins. Automation can facilitate the handling of insect cell cultures, infection with baculovirus, and protein expression.
Mammalian Cell Expression Systems:
CHO Cells (Chinese Hamster Ovary) are widely used for the production of recombinant proteins due to their ability to perform post-translational modifications similar to human cells. Automated systems can manage the culture of CHO cells in bioreactors, transfection with expression vectors, and subsequent protein purification steps.
Cell-Free Protein Synthesis Systems:
The reconstituted PURE System utilizes purified components of the translation machinery extracted from cells, allowing protein synthesis without the need for intact cells. Automation can streamline the preparation of reaction mixes and monitoring of protein synthesis.
Automated Liquid Handling Systems:
Automated liquid handling platforms equipped with robotic arms and various modules can assist in dispensing reagents, conducting transformations, setting up expression cultures, and performing purification steps.
High-Throughput Screening Platforms:
Automated systems integrated with high-throughput screening assays can accelerate the identification and optimization of protein expression conditions, such as media composition, induction conditions, and temperature regimes.
Online Monitoring and Control Systems:
Automated systems equipped with sensors and monitoring devices can continuously track parameters such as cell growth, protein expression levels, and culture conditions. This real-time data enables precise control and optimization of the expression process.
By integrating these automated approaches, researchers can efficiently express and purify proteins for various applications in fields such as drug discovery, structural biology, and biotechnology.
The system 100 may be implemented as an integrated system comprising each of the modules described above and shown in Figure 1. In some embodiments the input/output module 101, processing module 103, and sequence analysis module 105 may be implemented as separate software modules in a standalone computing device comprising one or more processors and one or more memory modules, the computing device being communicatively coupled to a separate standalone protein expression module 103. The separate software modules may be configured to communicate via one or more software interfaces such as Application Programming Interfaces (APIs). The one or more processors may be multi-threaded processors, and may comprise one or more Computer Processing Units (CPUs) and/or Graphics Processing Units (GPUs). In some embodiments, the sequence analysis module 105 may be a separate sequence analysis module implemented on a separate computing device such as a cloud computing server, and communicatively coupled with the processing module 103 via a software interface such as an API. In alternative embodiments, each of the input/output module 101, processing module 103, and sequence analysis module 105 are implemented on separate computing devices communicatively coupled together over a network such as a Local Area Network (LAN), or Wireless Local Area Network (WLAN).
Figure 2 illustrates a method 200 for analysing protein variants in accordance with an embodiment of the present invention. The steps of method 200 correspond to the functions described above performed by components of the system 100 of Figure 1. In particular, at step 201, the method 200 comprises receiving a plurality of input sequences. The receiving a plurality of input sequences comprises receiving a plurality of input sequences, each sequence coding for a variant of a target protein as described above in relation to the input/output module 101. At step 202, the method 200 comprises predicting 202a the folded protein structure of the coded protein for each received input sequence, and determining 202b a structural similarity between the predicted folded protein structure of each received input sequence and the target protein. The predicting 202a and determining 202b are performed as described above in relation to the processing module 103 and sequence analysis module 105. At step 203, the method 200 comprises identifying one or more preferred input sequences. The one or more preferred input sequences may be identified as those corresponding to the protein variants whose predicted folded protein structure is most structurally similar to that of the target protein, as described above in relation to the sequence analysis module 106 and processing module 103. The preferred input amino acid sequence may be converted to a nucleic acid sequence suitable for expression. The nucleic acid sequence may be codon optimised for the desired expression system. At step 204, the method 200 comprises expressing the coded variants of each identified preferred input sequence. The expressing 204 is performed in a cell-free system, and may optionally be performed in a digital microfluidic device, as described above in relation to the protein expression module 106. At step 205, the method 200 comprises determining one or more optimal input sequences based on the expressed coded variants. The determining 205 may be performed by calculating one or more assay metrics, by performing one or more protein assays on the expressed preferred input sequences, as described above in relation to the protein expression module 106.
As described herein, a protein "variant" may include known isoforms or orthologs, length variants derived from the original sequence with a number of amino acids truncated at either or both of the ends or at some predetermined position in the sequence. The location of the truncation position may depend on various factors including, but not limited to, conserving regions, domains and other positions (e.g. catalytic residues important for enzymatic function), and other computed conservation or disorder scores.
The term "DNA sequences" used herein refers to a succession of bases (nucleotides) signified by a set of four different letters that indicate the order of nucleotides forming alleles within a DNA
molecule or an RNA molecule. DNA sequences may also be known to the skilled person as, and referred to as "nucleic acid sequences".
It should be appreciated that one or more of the modules of system 1 may be implemented as distinct functional components. The input/output module 101, processing module 103, and sequence analysis module may be implemented as functional components within a single device. In other embodiments, the sequence analysis module 105 may be provided as a separate functional module, for example being stored on a separate device such as a cloud computing server.
As will be appreciated by one of skill in the art, the invention described herein may be embodied in whole or in part as a method, a data processing system, or a computer program product including computer readable instructions. Accordingly, the invention may take the form of an entirely hardware embodiment or an embodiment combining software, hardware and any other suitable approach or apparatus.
The computer readable program instructions may be stored on a non-transitory, tangible computer readable medium. The computer readable storage medium may include one or more of an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk.
Exemplary embodiments of the invention may be implemented as a circuit board which may include a CPU, a bus, RAM, flash memory, one or more ports for operation of connected I/O apparatus such as printers, display, keypads, sensors and cameras, ROM, a communications subsystem such as a modem, and communications media.
As will be appreciated by one of skill in the art, the term "processing means" may correspond to any suitable processing device. For example, the processing means may be any of a computer processor, graphics processor, programmable logic device, microprocessor, or any other suitable device. It will be appreciated that the processing means may comprise a plurality of
processing devices, and may be a combination of different processing devices such as those described above.
Expression of the protein of interest may be monitored, for example using a split fluorescent protein. Disclosed herein is a method for the real-time monitoring of in-vitro protein synthesis comprising a. In-vitro transcription and translation of a protein of interest fused to a peptide tag; and b. monitoring the presence of the peptide tag using a further polypeptide which in the presence of the peptide tag produces a detectable signal.
Disclosed herein is a method for the monitoring of cell-free protein synthesis in a droplet on a digital microfluidic device comprising a. cell-free transcription and translation of a protein of interest fused to a peptide tag; and b. monitoring the presence of the peptide tag using a further polypeptide which in the presence of the peptide tag produces a detectable signal.
Disclosed herein is a method for the monitoring of cell-free protein synthesis in a droplet on a digital microfluidic device comprising a. cell-free transcription and translation of a protein of interest fused to a ccGFPn peptide tag; and b. monitoring the presence of the peptide tag using a further ccGFPi-io polypeptide which in the presence of the ccGFPn peptide tag produces a detectable signal.
The use of the terms "in-vitro" and "cell-free" may be used interchangeably herein.
The detectable signal may be for example fluorescence or luminescence. The detectable signal may also be caused by the binding or docking of a ligand to the complemented oligopeptide, peptide, or polypeptide tag fused to the protein of interest.
The detectable signal may also be caused by the binding of the polypeptide to the protein of interest fused to a His-tag.
Any in-vitro transcription and translation may be used, for example extract-based systems derived from rabbit reticulocyte lysate, human lysate, Chinese Hamster Ovary lysate, a wheat germ, HEK293 lysate, E. coli lysate, yeast lysate.
Alternatively the in-vitro transcription and translation may be assembled from purified components, for example a system of purified recombinant elements (PURE).
The in-vitro transcription and translation may be coupled or uncoupled.
The peptide tag may be one component of a fluorescent protein and the further polypeptide a complementary portion of the fluorescent protein. The fluorescent protein could include sfGFP, GFP, eGFP, deGFP, frGFP, eYFP, eBFP, eCFP, Citrine, Venus, Cerulean, Dronpa, DsRED, mKate, mCherry, mRFP, FAST, SmURFP, miRFP670nano. For example the peptide tag may be GFPn and the further polypeptide GFPi-io. The peptide tag may be one component of sfCherry. The peptide tag may be sfCherryn and the further polypeptide sfCherryi-io. The peptide tag may be CFASTn or CFASTio and the further polypeptide NFAST in the presence of a hydroxybenzylidene rhodanine analog.
The peptide tag may be ccGFPn and the further polypeptide ccGFPi-io.
For example, the GFPi-io polypeptide amino acid sequence could be derived from sfGFP:
SEQ. ID NO: 1
MSKGEELFTGVVPILVELDGDVNGHKFSVRGEGEGDATNGKLTLKFICTTGKLPVPWPTLVTTLTYGVQCFS RYPDHMKRHDFFKSAMPEGYVQERTISFKDDGTYKTRAEVKFEGDTLVNRIELKGIDFKEDGNILGHKLEYN FNSHNVYITADKQKNGIKANFKIRHNVEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQSVLSKDPNEK
Alternatively, the GFPi-io polypeptide amino acid sequence could be further mutated from the sequence above to become brighter more quickly upon complementation. The sequence may have a greater than 90 % homology to any sequence mentioned herein. The sequence may have a greater than 95 % homology to any sequence mentioned herein.
SEQ. ID NO: 2
MSKGEELFTGVVPILVELDGDVNGHKFSVRGEGEGDATIGKLTLKFICTTGKLPVPWPTLVTTLTYGVQCFSR
YPDHMKRHDFFKSAMPEGYVQERTISFKDDGKYKTRAVVKFEGDTLVNRIELKGTDFKEDGNILGHKLEYNF NSHNVYITADKQKNGIKANFTVRHNVEDGSVQLADHYQQNTPIGDGPVLLPDNHYLSTQTVLSKDPNEK
The GFPi-io polypeptide amino acid sequence could also be derived from ccGFP, having a greater than 90 or 95% homology to:
SEQ. ID NO: 3 (CCGFPI-H)
MSLSKQVVKEDMKMTYHMDGCVNGHYFTIEGEGTGKPFKGQKTLKLRVTEGGPLPFAFDILSATFTYGNR
CFCDYPEDMPDYFKQSLPEGYSWERTMMYEDGACGTASAHISLDKNGFVHNSTFHGVNFPANGPVMKK
KGVNWEPSSEKITACDGILKGDVTMFLVLEGGHRLKCLFQTTYKADKVVKM PPNHIIEHRLVRSEDGDAVQI
QEHAVAKYFTV
SEQ. ID NO: 4 (ccGFPi-io)
MSLSKQVVKEDMKMTYHMDGCVNGHYFTIEGEGTGKPFKGQKTLKLRVTEGGPLPFAFDILSATFTYGNR
CFCDYPEDMPDYFKQSLPEGYSWERTMMYEDGACGTASAHISLDKNGFVHNSTFHGVNFPANGPVMKK
KGVNWEPSSEKITACDGILKGDVTMFLVLEGGHRLKCLFQTTYKADKVVKM PPNHIIEHRLVRSED
SEQ ID NO 5 (ccGFPi-io)
MSM EKQVLKENMKTTYHM DGSVDGHYFEIEGEGTGNPFKGEQELKLRVTKGGPLPFAFDILSPTFTYGNR
VFTDYPEDMPDYFKQSLPEGYSWERTM MYEDGATATASARISLDKNGFVHKSTFHGENFPANGPVMKKK
GVDWEPSSETITPEDGILKGDVEMFLVLEGGQRLKALFQTTYKANKVVKM PPRHKIEHRLVRS
Nucleic acid sequence to express seq ID No 5 ccGFPi-io; SEQ ID NO: 6
5'atgagcatggaaaaacaggtgctgaaagaaaacatgaaaaccacctatcacatggatggtagcgttgatggtcactattttgaaatt gaaggtgaaggcaccggcaatccgtttaaaggtgaacaagaactgaaactgcgtgttaccaaaggtggtccgctgccgtttgcatttga tattctgagcccgacctttacctatggtaatcgtgtttttaccgactatccggaagatatgccggattatttcaaacagagcctgccggaa ggttatagctgggaacgtaccatgatgtatgaagatggtgcaaccgcaaccgccagcgcacgtattagcctggataaaaatggttttgt gcataagagcacctttcacggtgaaaactttccggcaaatggtccggttatgaaaaagaaaggtgttgattgggaaccgagcagcgaa accattacaccggaagatggtattctgaaaggtgatgttgaaatgtttctggttctggaaggtggtcagcgtctgaaagccctgtttcaga ccacctataaagccaataaagtggttaaaatgcctccgcgtcataaaattgaacatcgtctggttcgtagc
SEQ ID NO 7
MSMSKQVLKENM KTTYHMDGSVNGHYFTIEGEGTGNPFKGQQSLKLRVTKGGPLPFAFDILSPTFTYGNR
VFTDYPEDMPDYFKQSLPEGYSWERTM MYEDGATATASARISLDKNGFVHKSTFHGENFPANGPVMKKK
GVNWEPSSETITPSDGILKGDVTMFLVLEGGQRLKALFQTTYKANKVVKM PPRHKIEHRLVRS
Nucleic acid sequence to express seq ID No 7 ccGFPi-io; SEQ ID NO: 8
5'atgagcatgagcaaacaggtgctgaaagaaaatatgaaaaccacctatcacatggatggtagcgttaatggtcactattttaccatt gaaggtgaaggcaccggtaatccgtttaaaggtcagcagagcctgaaactgcgtgttaccaaaggtggtccgctgccgtttgcatttga tattctgagcccgacctttacctatggtaatcgtgtttttaccgactatccggaagatatgccggattatttcaaacagagcctgccggaa ggttatagctgggaacgtaccatgatgtatgaagatggtgcaaccgcaaccgccagcgcacgtattagcctggataaaaatggttttgt gcataagagcacctttcacggtgaaaactttccggcaaatggtccggttatgaaaaagaaaggtgttaattgggaaccgagcagcgaa accattacaccgagtgatggtattctgaaaggtgatgttaccatgtttctggttctggaaggtggtcagcgtctgaaagccctgtttcaga ccacctataaagccaataaagtggttaaaatgcctccgcgtcataaaattgaacatcgtctggttcgtagc
The complementary GFPn peptide amino acid sequence could be the following:
1. KRDHMVLLEFVTAAGITGT (SEQ. ID NO: 9)
2. KRDHMVLHEFVTAAGITGT (SEQ ID NO: 10)
3. KRDHMVLHESVNAAGIT (SEQ ID NO: 11)
4. RDHMVLHEYVNAAGIT (SEQ ID NO: 12)
5. GDAVQIQEHAVAKYFTV (SEQ ID NO: 13)
6. GDTVQLQEHAVAKYFTV (SEQ ID NO: 14)
7. GETIQLQEHAVAKYFTE (SEQ ID NO: 15) or a truncated version thereof. Truncations may involve a shortening of up to 5 amino acids from the N terminus, the C terminus or a combination thereof.
GFPn or GFPi-io can be fused to the protein of interest through an amino acid linker. In one embodiment, the oligopeptide, peptide, or polypeptide linker can be 0 - 50 amino acids.
Also disclosed are nucleic acid sequences for expressing particular tags. Nucleic acid sequences include
SEQ ID NO: 16
5'GGTGATACCGTTCAGCTGCAAGAACATGCAGTTGCAAAATACTTTACCGTG
SEQ ID NO: 17
5'GGTGAAACCATCCAGTTACAAGAACACGCCGTGGCCAAATATTTCACCGAA or a truncated version thereof.
These sequences may be repeated one or more times to produce a protein having multiple GFPn domains.
For example, the sfCherryi-io polypeptide amino acid sequence could be:
SEQ ID NO: 18
MEEDNMAIIKEFMRFKVHM EGSVNGHEFEIEGEGEGHPYEGTQTAKLKVTKGGPLPFAWDILSPQFMYGS
KAYVKHPADIPDYLKLSFPEGFTWERVM NFEDGGVVTVTQDSSLQDGEFIYKVKLLGTNFPSDGPVMQKKT
MGWEASTERMYPEDGALKGEINQRLKLKDGGHYDAEVKTTYKAKKPVQLPGAYNVDIKLDITSHNED
The complementary sfCherryll peptide amino acid sequence could be:
SEQ. ID NO: 19
YTIVEQYERAEGRHSTGG sfCherryll or sfCherryi-io can be fused to the protein of interest through an amino acid linker.
In one embodiment, the oligopeptide, peptide, or polypeptide linker can be 0 - 50 amino acids.
For example, the NFAST polypeptide amino acid sequence could be:
SEQ IS NO: 20
MEHVAFGSEDIENTLAKMDDGQLDGLAFGAIQLDGDGNILQYNAAEGDITGRDPKQVIGKNFFKDVAPGT
DSPEFYGKFKEGVASGNLNTMFEWM IPTSRGPTKVKVHM KKALS
The complementary CFAST11 peptide amino acid sequence could be:
SEQ ID NO: 21
GDSYWVFVKRV
Or the complementary CFAST10 peptide amino acid sequence could be:
SEQ ID NO: 22
GDSYWVFVKR
NFAST, CFAST11, and/or CFAST10 can be fused to the protein of interest through an amino acid linker. In one embodiment, the oligopeptide, peptide, or polypeptide linker can be 0 - 50 amino acids.
The peptide tag may also be one component of a protein that forms a detectable substrate, such as a luminescent or colorigenic substrate. The protein could include beta-galactosidase, betalactamase, or luciferase.
The protein may be fused to multiple tags. For example the protein may be fused to multiple GFPn peptide tags and the synthesis occurs in the presence of multiple GFPi-io polypeptides. The GFP may be ccGFP or sfGFP. For example the protein may be fused to multiple ccGFPn peptide tags and the synthesis occurs in the presence of multiple ccGFPi-io polypeptides. The protein of interest may be fused to one or more sfCherryn peptide tags and one or more GFPn peptide tags and the synthesis occurs in the presence of one or more GFPi-io polypeptides and one or more sfCherryno polypeptides.
Any protein of interest may be synthesised. The protein may be an enzyme, for example a terminal deoxynucleotidyl transferase (TdT) enzyme or a truncated version thereof or the homologous amino acid sequence of a terminal deoxynucleotidyl transferase (TdT) enzyme in other species or the homologous amino acid sequence of Polp, Poip, PoIX, and Pol0 of any species or the homologous amino acid sequence of X family polymerases of any species.
Protein sequences disclosed herein may be attached to further elements to improve solubility. The variant may be attached to one or more solubility enhancing sequences. The solubility enhancing sequence may be a peptide sequence or a naturally occurring sequence. The solubility enhancing sequence may be selected from for example maltose binding protein (MBP), Small Ubiquitin-like Modifier (SUMO), Glutathione S-transferase (GST) or thioredoxin (TRX). The tags may be attached to either the C or N terminus. Any example of a solubility enhancer may be used. A list of possible proteins is shown below. Any sequence selected from the list below may be chosen:
The purification tag can be a region of amino acid/peptide sequence. The affinity binding site can be a region of amino acid/peptide sequences specific to a particular antibody. The tag can be attached to the N or C terminus. For example the purification tag can be selected from the list of exemplary peptide affinity binding sites below:
Alfa-tag (SRLEEELRRRLTE) (SEQ ID NO: 23)
Avi-tag (GLNDIFEAQKIEWHE) (SEQ. ID NO: 24) C-tag (EPEA) (SEQ ID NO: 25)
Calmodulin-tag (KRRWKKNFIAVSAANRFKKISSSGAL) (SEQ ID NO: 26)
Dogtag (DIPATYEFTDGKHYITNEPIPPK) (SEQ ID NO: 27)
E-tag (GAPVPYPDPLEPR) (SEQ ID NO: 28)
FLAG (DYKDDDDK) (SEQ ID NO: 29) G4T (EELLSKNYHLENEVARLKK) (SEQ ID NO: 30)
HA (YPYDVPDYA) (SEQ ID NO: 31)
His (HHHHHH) (SEQ ID NO: 32)
Isopeptag (TDKDMTITFTNKKDAE) (SEQ ID NO: 33) lanthanide binding tag (LBT) (FIDTNNDGWIEGDELLLEEG) (SEQ ID NO: 34)
Myc (EQKLISEEDL) (SEQ ID NO: 35)
NE-Tag (TKENPRSNQEESYDDNES) (SEQ. ID NO: 36)
Poly Glutamate-tag (EEEEEEE) (SEQ ID NO: 37)
Poly Arginine-tag (RRRRRRR) (SEQ ID NO: 38)
RholD4-tag (TETSQVAPA) (SEQ ID NO: 39)
SBP-tag (MDEKTTGWRGGHVVEGLAGELEQLRARLEHHPQGQREP) (SEQ ID NO: 40)
Sdytag (DPIVMIDNDKPIT) (SEQ ID NO: 41)
SH3 (STVPVAPPRRRRG) (SEQ ID NO: 42)
Snooptag (KLGDIEFIKVNK) (SEQ ID NO: 43)
Softag 1 (SLAELLNAGLGGS) (SEQ ID NO: 44)
Softag 3 (TQDPSRVG) (SEQ ID NO: 45)
Spot-tag (PDRVRAVSHWSS) (SEQ ID NO: 46)
Spytag (AHIVMVDAYKPTK) (SEQ ID NO: 47)
S-tag (KETAAAKFERQHMDS) (SEQ ID NO: 48)
Strep-tag (AWAHPQPGG) (AWRHPQFGG) (SEQ ID NO: 49)
Strep-tag II (WSHPQFEK) (SEQ ID NO: 50)
T7tag (MASMTGGQQMG) (SEQ ID NO: 51)
TC-tag (EVHTNQDPLD) (SEQ ID NO: 52)
Ty-tag (CCPGCC) (SEQ ID NO: 53)
VSV-tag (YTDIEMNRLGK) (SEQ ID NO: 54)
Xpress-tag (DLYDDDDK) (SEQ ID NO: 55)
The binding moiety tag can be a sub-component of a fluorescent protein. Thus the fully assembled protein becomes fluorescent. For example if the detector protein contains GFPi-io and the expressed protein tag contains a GFPn peptide, complementation forms fluorescent GFP, allowing simultaneous monitoring and stability evaluation. The expressed material can be monitored over time or conditions by monitoring changes in the level of material generating a fluorescent signal.
Devices
The manipulation of droplets by the application of electrical potential can be achieved on electrodes covered with an insulator or a dielectric or a series of insulators or dielectrics. Droplet manipulation as a result of an applied electrical potential is known as electrowetting. Electrokinesis occurs as result of a non-uniform electric field that influences the hydrostatic equilibrium of a dielectric liquid (dielectrophoresis or DEP) or a change in the contact angle of
the liquid on solid surface (electrowetting-on-dielectric or EWoD). DEP can also be used to create forces on polarizable particles to induce their movement. The electrical signal can be transmitted to a discrete electrode, a transistor, an array of transistors, or a sheet of semiconductor film whose electrical properties can be modulated by an optical signal.
EWoD phenomena occur when droplets are actuated between two parallel electrodes covered with a hydrophobic insulator or dielectric. The electric field at the electrode-electrolyte interface induces a change in the surface tension, which results in droplet motion as a result of a change in droplet contact angle. The electrowetting effect can be quantitatively treated using Young- Lippmann equation: cos0 - cos0o= (l/2yLG) c.V2 where 0o is the contact angle when the electric field across the interfacial layer is zero, yLG is the liquid-gas tension, c is the specific capacitance (given as sr. so/t, where sr is dielectric constant of the insulator/dielectric, so is permittivity of vacuum, t is thickness) and V is the applied voltage or electrical potential. The change in contact angle (inducing droplet movement) is thus a function of surface tension, electrical potential, dielectric thickness, and dielectric constant.
When a droplet is actuated by EWoD, there are two opposing sets of forces that act upon it: an electrowetting force induced by electric field and resistant forces that include the drag forces resulting from the interaction of the droplet with filler medium and the contact line friction (ref). The minimum voltage applied to balance the electrowetting force with the sum of all drag forces (threshold voltage) is variably determined by the thickness-to-dielectric contact ratio of the insulator/dielectric, (t/e)1/2. Thus, to reduce actuation voltage, it is required to reduce (t/e)1/2 (i.e., increase dielectric constant or decrease insulator/dielectric thickness). To achieve low voltage actuation, thin insulator/dielectric layers must be used. However, the deposition of high quality thin insulator/dielectric layers is a technical challenge, and these thin layers are easily damaged before the desired electrowetting contact angle is large enough to drive the droplet is achieved. Most academic studies thus report the use of much higher voltages >100 V on easily fabricated, thick dielectric films (>3 pm) to effect electrowetting.
High voltage EWoD-based devices with thick dielectric films, however, have limited industrial applicability largely due to their limited droplet multiplexing capability. The use of low voltage
devices including thin-film transistors (TFT) and optically-activated amorphous silicon layers (a- Si) have paved the way for the industrial adoption of EWoD-based devices due to their greater flexibility in addressing electrical signals in a highly multiplex fashion. The driving voltage for TFTs or optically-activated a-Si are low (typically <15 V). The bottleneck for fabrication and thus adoption of low voltage devices has been the technical challenge of depositing high quality, thin film insulators/dielectrics. Hence there has been a particular need for improving the fabrication and composition of thin film insulator/dielectric devices.
Typically, the electrodes (or the array elements) used for EWoD are covered with (i) a hydrophilic insulator/dielectric and a hydrophobic coating or (ii) a hydrophobic insulator/dielectric. Commonly used hydrophobic coatings comprise of fluoropolymers such as Teflon AF 1600 or CYTOP. The thickness of this material as a hydrophobic coating on the dielectric is typically <100 nm and can have defects in the form of pinholes or a porous structure; hence, it is particularly important that the insulator/dielectric is pinhole free to avoid electrical shorting. Teflon has also been used as an insulator/dielectric, but it has higher voltage requirements due to its low dielectric constant and the thickness required to make it pinhole free. Other hydrophobic insulator/dielectric materials can include polymer-based dielectrics such as those based on siloxane, epoxy (e.g. SU-8), or parylene (e.g., parylene N, parylene C, parylene D, or parylene HT). Due to minimal contact angle hysteresis and a higher contact angle with aqueous solutions, Teflon is still used as a hydrophobic topcoat on these insulator/dielectric polymers. However, there are difficulties in reliably producing <1 micron pinhole-free coatings of parylene or SU-8; thus, the thickness of these materials is typically kept at 2-5 microns at the cost of increased voltage requirements for electrowetting. It has also been reported that traditional EWoD devices with parylene C are easily broken and unstable for repeated droplet manipulation with cell culture medium. Multi-layer insulator devices deposited with metal-oxide and parylene C films have been used to produce a more robust insulator/dielectric and enable operations with lower applied voltages. Inorganic materials, such metal oxides and semiconductor oxides, commonly used in the CMOS industry as "gate dielectrics", have been used as insulator/dielectric for EWoD devices. They offer the advantage of utilizing standard cleanroom processes for thin film depositions (<100 nm). These materials are inherently hydrophilic, requiring an additional hydrophobic coating, and can be prone to pinhole formation as a result of thin film layer deposition process. Together with the need for lower voltage operations of EWoD, recent developmental work has focused on (1) using materials with improved dielectric properties (e.g., using high-dielectric constant insulators/dielectrics), (2) optimizing the fabrication process to make the insulator/dielectric pinhole free to avoid dielectric breakdown.
Operation of EWoD devices suffers from contact angle saturation and hysteresis, which is believed to be brought about by either one or combination of these phenomena: (1) entrapment of charges in the hydrophobic film or insulator/dielectric interface, (2) adsorption of ions, (3) thermodynamic contact angle instabilities, (4) dielectric breakdown of dielectric layer, (5) the electrode-electrode-insulator interface capacitance (arising from the double layer effect), and (6) fouling of the surface (such as by biomacromolecules). One of the adverse effects of this hysteresis is reduced operational lifetime of the EWoD-based device.
Contact angle hysteresis is believed to be a result of charge accumulation at the interface or within the hydrophobic insulator after several operations. The required actuation voltage increases due to this charging phenomenon resulting in eventual catastrophic dielectric breakdown. The most probable explanation is that pinholes at the insulator/dielectric may allow the liquid to come into contact with the electrode causing electrolysis. Electrolysis is further facilitated by pinhole-prone or porous hydrophobic insulators.
Most of the studies to understand contact angle hysteresis on EWoD have been conducted on short time scales and with low conductivity solutions. Long duration actuations (e.g., >1 hour) and high conductivity solutions (e.g., 1 M NaCI) could produce several effects other than electrolysis. The ions in solution can permeate through the hydrophobic coat (under the applied electric field) and interact with the underlying insulator/dielectric. Ion permeation can result in (1) change in dielectric constant due to charge entrapment (which is different from interfacial charging) and (2) change in surface potential of a pH sensitive metal oxide. Both can result in reduction of electrowetting forces to manipulate aqueous droplets, leading to contact angle hysteresis. The inventors have previously found that the damage from high conductivity solutions reduces or disables electrowetting on electrodes by inhibiting the modulation of contact angle when an electric field is applied.
An electrokinetic device includes a first substrate having a matrix of electrodes, wherein each of the matrix electrodes is coupled to a thin film transistor, and wherein the matrix electrodes are overcoated with a functional coating comprising: a dielectric layer in contact with the matrix electrodes, a conformal layer in contact with the dielectric layer, and a hydrophobic layer in contact with the conformal layer; a second substrate comprising a top electrode; a spacer disposed between the first substrate and the second substrate and defining an electrokinetic workspace; and a voltage source operatively coupled to the matrix electrodes.
The dielectric layer may comprise silicon dioxide, silicon oxynitride, silicon nitride, hafnium oxide, yttrium oxide, lanthanum oxide, titanium dioxide, aluminium oxide, tantalum oxide, hafnium silicate, zirconium oxide, zirconium silicate, barium titanate, lead zirconate titanate, strontium titanate, or barium strontium titanate. The dielectric layer may be between 10 nm and 100 pm thick. Combinations of more than one material may be used, and the dielectric layer may comprise more than one sublayer that may be of different materials.
The conformal layer may comprise a parylene, a siloxane, or an epoxy. It may be a thin protective parylene coating in between the insulating dielectric and the hydrophobic coating. Typically, parylene is used as a dielectric layer on simple devices. In this invention, the rationale for deposition of parylene is not to improve insulation/dielectric properties such as reduction in pinholes, but rather to act as a conformal layer between the dielectric and hydrophobic layers. The inventors find that parylene, as opposed to other similar insulating coatings of the same thickness such as PDMS (polydimethylsiloxane), prevent contact angle hysteresis caused by high conductivity solutions or solutions deviating from neutral pH for extended hours. The conformal layer may be between 10 nm and 100 pm thick.
The hydrophobic layer may comprise a fluoropolymer coating, fluorinated silane coating, manganese oxide polystyrene nanocomposite, zinc oxide polystyrene nanocomposite, precipitated calcium carbonate, carbon nanotube structure, silica nanocoating, or slippery liquid-infused porous coating.
The elements may comprise one or more of a plurality of array elements, each element containing an element circuit; discrete electrodes; a thin film semiconductor in which the electrical properties can be modulated by incident light; and a thin film photoconductor whose properties can be modulated by incident light.
The functional coating may include a dielectric layer comprising silicon nitride, a conformal layer comprising parylene, and a hydrophobic layer comprising an amorphous fluoropolymer. This has been found to be a particularly advantageous combination.
The electrokinetic device may include a controller to regulate a voltage provided to the individual matrix electrodes. The electrokinetic device may include a plurality of scan lines and a plurality of gate lines, wherein each of the thin film transistors is coupled to a scan line and a
gate line, and the plurality of gate lines are operatively connected to the controller. This allows all the individual elements to be individually controlled.
The second substrate may also comprise a second hydrophobic layer disposed on the second electrode. The first and second substrates may be disposed so that the hydrophobic layer and the second hydrophobic layer face each other, thereby defining the electrokinetic workspace between the hydrophobic layers.
The method is particularly suitable for aqueous droplets with a volume of 1 pL or smaller.
The EWoD-based devices shown and described below are active matrix thin film transistor devices containing a thin film dielectric coating with a Teflon hydrophobic top coat. These devices are based on devices described in the E Ink Corp patent filing on "Digital microfluidic devices including dual substrate with thin-film transistors and capacitive sensing", US patent application no 2019/0111433, incorporated herein by reference.
Described herein are electrokinetic devices, including: a first substrate having a matrix of electrodes, wherein each of the matrix electrodes is coupled to a thin film transistor, and wherein the matrix electrodes are overcoated with a functional coating comprising: a dielectric layer in contact with the matrix electrodes, a conformal layer in contact with the dielectric layer, and a hydrophobic layer in contact with the conformal layer; a second substrate comprising a top electrode; a spacer disposed between the first substrate and the second substrate and defining an electrokinetic workspace; and a voltage source operatively coupled to the matrix electrodes;
Described herein is an electrokinetic device, including: a first substrate having a matrix of electrodes, wherein each of the matrix electrodes is coupled to a thin film transistor, and wherein the matrix electrodes are overcoated with a functional coating comprising: one or more dielectric layer(s) comprising silicon nitride, hafnium oxide or aluminum oxide in contact with the matrix electrodes, a conformal layer comprising parylene in contact with the dielectric layer, and a hydrophobic layer in contact with the conformal layer;
a second substrate comprising a top electrode; a spacer disposed between the first substrate and the second substrate and defining an electrokinetic workspace; and a voltage source operatively coupled to the matrix electrodes;
The electrokinetic devices as described may be used with other elements, such as for example devices for heating and cooling the device or reagent cartridges for the introduction of reagents as needed.
The device can be an active-matrix thin film transistor (AM-TFT) based device.
Also disclosed is an active-matrix thin film transistor (AM-TFT) device having a substrate bearing a plurality of electrodes, the device comprising multiple fluidic inlet ports on at least two sides of the device, wherein the inlet ports on each side of the device are evenly spaced and wherein the device is connected to a syringe pump.
The device may comprise two substrates, wherein at least one substrate has a plurality of electrodes, and the two substrates define parallel plates that are separated by a spacer to define a volume.
The fluidic entry may come via holes in the upper plate or through the spacer. The entry holes may be in the top substrate. The plurality of electrodes may be on a bottom substrate. The top substrate be of glass or polymer and may have a thickness ranging from 0.5 mm to 20 mm.
The spacer may comprise an adhesive with beads of a defined size distribution. The spacer may comprise a polymer material of a defined thickness. The spacer may comprise glass, in which case the layers can be fused together. The spacer gap and therefore height of fluid in the device may be between 50 microns and 250 microns. The spacer gap and therefore height of fluid in the device may be between 100 microns and 150 microns.
The filler liquid may be moved via an automated manner, or may be moved under gravity. A hydrostatic head of pressure can be used to move the liquid within the device. The wells are at least partially filled with filler fluid before the aqueous reagents are loaded. The filler fluid may be less dense than the aqueous phase such that the aqueous phase sinks in the wells.
Alternatively the aqueous phase may sit above the filler fluid, in which case all the filler fluid must be withdrawn from the wells in order to enable entry of the aqueous fluid.
The device may be connected to a pump, for example a syringe pump, a peristaltic pump, a disc pump, a diaphragm pump, or a pneumatic pump. The pump enables filling of the device with filler liquid in an automated manner. Once filled, the pump enables partial withdrawal of the filler fluid to create a negative pressure in the device which draws in reagents from the wells. Thus the filling and withdrawal of fluid may be performed in an automated manner to allow largely 'hands-free' loading of the aqueous reagents. An automated filler liquid filling and withdrawal method may be integrated into an instrument that provides other functions relating to the digital microfluidic device, including heating, cooling, optical, sensing, mechanical, and magnetic functions.
The wells of the device may be at 90 degrees to each other. The inlets may be at 180 degrees to each other. The inlets may be on 4 sides of the device. Each side may have at least 4, 8 or 12 ports. Each side may have 8 ports. The device may have 4 sets of 8 ports. The number of ports may vary on different sides of the device, for example one side may have 8 ports and one side 4 ports. The device may have 8 ports on 3 sides and 16 ports on a fourth side. The ports may be offset to give multiple rows of linear ports on one side, for example a first and second row where the second row is behind by offset from the first row such that the source liquid can flow between the ports of the first row. The rows may be a zig-zag fashion.
The pitch between inlet ports may be 9 mm. The pitch between inlet ports may be 4.5 mm. The inlet ports have a pitch of 4.5 mm or a multiple of thereof. This would cover 24 well, 48 well, 96 well, 384 well ports. The pitch of the ports may be the same on each side of the device, or may be different sized. In this context the pitch refers to the distance between the centre of each inlet.
The volume of aqueous reagents loaded per inlet port may be between 1 microlitre and 50 microlitres. The volume may be between 1 microlitre and 20 microlitres.
The aqueous liquid may be introduced to the wells by a pipette, a multichannel pipette, a syringe, a blister pack, an acoustic dispenser, or a robotic liquid handler. The aqueous liquids may be loaded simultaneously from multiple wells, which may be on the same side or multiple sides of the devices. Each well is a separate liquid, and can be the same or different to the
contents of the aqueous volume in other wells. The volume of aqueous liquid loaded in each port can be the same or can be different.
The automated filling and/or withdrawing of filler fluid may be controlled by software. The device may be part of a larger instrument system that provides environmental control such as temperature control or light control and may have analytical capabilities such as optical systems for fluorescence or luminescence assay detection.
The location of the aqueous layer is controlled by the actuation of electrodes to form reservoirs in defined areas. A plurality of electrodes is actuated to control the location of the aqueous liquid once it has been drawn onto the substrate bearing a plurality of electrodes. Multiple reservoirs may be formed on the device.
Experimental & Results
Figure 5 shows the experimental result from 24 different proteins expressed in a reconstituted cell-free protein synthesis system in droplets on an electrowetting on dielectric (EWoD) device. Each construct contains a GFPn tag. In the rows marked Screen, the GFPi-io detector species is present from the start of expression. The rows marked Endpoint shows the fluorescence signal from 10 hours expression in the absence of GFPi-io detector species followed by 5 hours complementation with the GFPi-io detector species. This experiment showed significant differences between expression/complementation in Screen BioInk compared to Endpoint detection. The detected protein clusters formed after expression only with endpoint detection mean that the protein aggregated after expression, thereby lowering the soluble yield. It is evident from the image the presence of speckles for several constructs, indicating the likely presence of protein aggregates. The level of aggregation enables identification of conditions which are worthy of further testing, and identification of conditions having high level of aggregated protein from which further purification is unlikely to give material.
Example using structure prediction to guide a higher level of expression
AlphaFold structures were generated for 4 wild type terminal transferase (TdT) sequences. Each TdT sequence has a length region of unstructured sequence known as the BRCT domain. The "BRCA-1 C-terminal (BRCT) domain" refers to the C-terminal domain of a breast cancer susceptibility protein. This domain is found predominantly in proteins involved in cell cycle checkpoint functions responsive to DNA damage, for example as found in the breast cancer
DNA-repair protein BRCA1. The domain is an approximately 100 amino acid tandem repeat, which appears to act as a phospho-protein binding domain. For example, the BRCT domain is present in all TdT sequences in the region from amino acid residues 1-130. Each of the TdT structures shown in the left column of Figure 6 has a long unstructured region at the C terminus corresponding to the BRCT domain. Expression of the wild type proteins (Figure 7) gives a low level of soluble expression.
Removal of the BRCT domain gives the structures shown in the right column of Figure 6. The long unstructured region has been removed. The structure of the central active region is not altered by the deletion. Figure 7 shows expression levels of the BRCT deleted constructs. Removal of the unstructured regions flanking the BRCT domain and the BRCT domain itself increased expression and yields of the TdT ortholog candidate TdT proteins.
Example using structure prediction to shorten sequences for increase in protein expression.
Aim: generate length variants of 9°N DNA polymerase based on domain prediction, structure data and AlphaFold predictions for expression screening.
UniProt ID: Q.56366
PDBJD: https://www.rcsb.org/structure/5omq
Structure predictions were generated for a selection of truncated fragments, the full length WT and three truncated sequences are shown in Figure 8.
>9oN_WT (1-775) (SEQ ID NO: 56)
MILDTDYITENGKPVIRVFKKENGEFKIEYDRTFEPYFYALLKDDSAIEDVKKVTAKRHGTVVKVKRAEKVQKK FLGRPIEVWKLYFNHPQDVPAIRDRIRAHPAVVDIYEYDIPFAKRYLIDKGLIPMEGDEELTMLAFAIATLYHE GEEFGTGPILM ISYADGSEARVITWKKIDLPYVDVVSTEKEMIKRFLRVVREKDPDVLITYNGDNFDFAYLKKR CEELGIKFTLGRDGSEPKIQ.RMGDRFAVEVKGRIHFDLYPVIRRTINLPTYTLEAVYEAVFGKPKEKVYAEEIAQ. AWESGEGLERVARYSMEDAKVTYELGREFFPM EAQ.LSRLIGQ.SLWDVSRSSTGNLVEWFLLRKAYKRNELA PNKPDERELARRRGGYAGGYVKEPERGLWDNIVYLDFRSLYPSIIITHNVSPDTLNREGCKEYDVAPEVGHKF CKDFPGFIPSLLGDLLEERQ.KIKRKMKATVDPLEKKLLDYRQ.RAIKILANSFYGYYGYAKARWYCKECAESVTA WGREYIEMVIRELEEKFGFKVLYADTDGLHATIPGADAETVKKKAKEFLKYINPKLPGLLELEYEGFYVRGFFV TKKKYAVIDEEGKITTRGLEIVRRDWSEIAKETQARVLEAILKHGDVEEAVRIVKEVTEKLSKYEVPPEKLVIHE Q.ITRDLRDYKATGPHVAVAKRLAARGVKIRPGTVISYIVLKGSGRIGDRAIPADEFDPTKHRYDAEYYIENQ.VL PAVERILKAFGYRKEDLRYQ.KTKQ.VGLGAWLKVKGKK
>9oN_341-758 (2XGSA) (SEQ ID NO: 57)
GSAGSALWDVSRSSTGNLVEWFLLRKAYKRNELAPNKPDERELARRRGGYAGGYVKEPERGLWDNIVYLD FRSLYPSIIITHNVSPDTLNREGCKEYDVAPEVGHKFCKDFPGFIPSLLGDLLEERQKIKRKM KATVDPLEKKLL DYRQRAIKILANSFYGYYGYAKARWYCKECAESVTAWGREYIEMVIRELEEKFGFKVLYADTDGLHATIPGAD AETVKKKAKEFLKYINPKLPGLLELEYEGFYVRGFFVTKKKYAVIDEEGKITTRGLEIVRRDWSEIAKETQARVLE AILKHGDVEEAVRIVKEVTEKLSKYEVPPEKLVIHEQITRDLRDYKATGPHVAVAKRLAARGVKIRPGTVISYIV LKGSGRIGDRAIPADEFDPTKHRYDAEYYIENQVLPAVERILKAFGYRKEDLRYQ
>9oN_347-End (2XGSA) (SEQ. ID NO: 58)
GSAGSASSTGNLVEWFLLRKAYKRNELAPNKPDERELARRRGGYAGGYVKEPERGLWDNIVYLDFRSLYPSII ITHNVSPDTLNREGCKEYDVAPEVGHKFCKDFPGFIPSLLGDLLEERQKIKRKMKATVDPLEKKLLDYRQRAIK ILANSFYGYYGYAKARWYCKECAESVTAWGREYIEMVIRELEEKFGFKVLYADTDGLHATIPGADAETVKKK AKEFLKYINPKLPGLLELEYEGFYVRGFFVTKKKYAVIDEEGKITTRGLEIVRRDWSEIAKETQARVLEAILKHGD VEEAVRIVKEVTEKLSKYEVPPEKLVIHEQITRDLRDYKATGPHVAVAKRLAARGVKIRPGTVISYIVLKGSGRI GDRAIPADEFDPTKHRYDAEYYIENQVLPAVERILKAFGYRKEDLRYQKTKQVGLGAWLKVKGKK
>9oN_305-end+Flanks (SEQ ID NO: 59)
MSKEKRLEVLFQGPGSAGSALERVARYSMEDAKVTYELGREFFPMEAQLSRLIGQSLWDVSRSSTGNLVEW FLLRKAYKRNELAPNKPDERELARRRGGYAGGYVKEPERGLWDNIVYLDFRSLYPSIIITHNVSPDTLNREGCK EYDVAPEVGHKFCKDFPGFIPSLLGDLLEERQKIKRKMKATVDPLEKKLLDYRQRAIKILANSFYGYYGYAKAR WYCKECAESVTAWGREYIEMVIRELEEKFGFKVLYADTDGLHATIPGADAETVKKKAKEFLKYINPKLPGLLEL EYEGFYVRGFFVTKKKYAVIDEEGKITTRGLEIVRRDWSEIAKETQARVLEAILKHGDVEEAVRIVKEVTEKLSK YEVPPEKLVIHEQITRDLRDYKATGPHVAVAKRLAARGVKIRPGTVISYIVLKGSGRIGDRAIPADEFDPTKHR YDAEYYIENQVLPAVERILKAFGYRKEDLRYQKTKQVGLGAWLKVKGKKENLYFQSGGGGSGGGGSGGGG SGETIQLQEHAVAKYFTEEAAAKEAAAKEAAAKWSHPQFEK
Suitable nucleic acid templates were prepared to allow expression using an E. coli derived expression system. The full length nucleic acid was transformed into a vector suitable for growth in NEB 5-alpha Competent E. coli cells, which were grown overnight at 37 °C. The full length nucleic acid material was amplified using colony PCR to obtain the full length template. Varying length nucleic acid templates were prepared by further PCR amplifications using suitable primers. The PCR products corresponded to the expected MW of the templates and showed as single gel bands. The samples were purified using the GeneJet PCR purification kit.
Further nucleic acid amplification reactions were performed in order to add the Nuclera adaptor sequences required for protein expression. Templates for expression were prepared and purified using Nuclera's commercially available eGene™ preparation kit and instructions therein. Adapter sequences add an optional solubility tag or no solubility tag depending on the sequence of the adapter.
Expression was performed using Nuclera's eProtein Discovery™ instrument. Results of the Expression are shown in Figure 9. 8 length variants with either no solubility tag or a solubility tag P17 or FH8 were screened in 8 different expression mixes to screen 192 (8x3x8) conditions in parallel. Full length protein expression of the full 1-775 sequence was the lowest for all constructs tested. The C-terminal truncation 1-758 showed slightly better expression and had a higher ratio of purifiable to total protein. Removal of the putatively unstructured C-terminal region (759-End) appears to have been beneficial towards expression rate for all length variants tested. Overall, the results are in agreement with the AlphaFold structure predictions that predict that removal of the exonuclease domain in the upstream proximity to the alpha helix beginning with T349 (full length protein reference sequence) does not impair correct folding of the DNA pol. family B domain.
The 192 datapoints are shown in table 1 and plotted as a heatmap in Figure 9. The data shows a clear correlation between molecular weight and expression, as expected. Across all 192 compositions tested, expression of full length transcripts are typically less than 3 pM, whereas the truncations are greater than 4 pM. The best truncation conditions using manganese give expression greater than 5 pM (5.21 pM) for the truncation with FH8 tags, compared to the full length sequence having FH8 tags expressing at 1.4 pM (1.44 pM).
Table 1:
Interestingly, S347 N-terminal truncations appear to benefit most from the addition of a soluble tag. Both P17 and FH8 increase the yield by about 25%. L341 N-Terminal truncations, in contrast, do not appear to benefit from a N-Terminal tag. The latter have an additional small alpha helix between the polymerase family domain and the 2xGSA linker, that was still partially intact in the AlphaFold structure prediction.
In summary, all length variants designed were expressed in a purifiable form and with higher expression rates as the full length parent sequence. The results agree with the AlphaFold predictions, showing how helpful structure predictions can be in the design of length variants for eProtein Discovery screening assays.
Claims
1. A method for analysing protein sequences, the method comprising: receiving input data comprising a plurality of input sequences, each of the input sequences coding for a variant of a target protein; for each of the input sequences in the plurality of input sequences: predicting a folded protein structure of the coded protein of the input sequence; and determining one or more structural properties of the predicted folded protein structure; identifying one or more preferred input protein sequences based on the determined structural properties; taking nucleic acid input sequences for expressing the input protein sequences; expressing the coded variant of each of the identified one or more preferred input protein sequences; characterizing the protein variant sequences based on the expressed one or more coded variants; and determining one or more optimal input sequences based on the expressed one or more coded variants.
2. The method of claim 1 wherein the variants are selected from length variants, isoforms, polymorphisms, orthologs, homologs or the structure of binding sites.
3. The method of claim 1 or claim 2 wherein the expression is performed in a cell-free expression system.
4. The method of any one of the preceding claims, wherein determining a structural property of the predicted folded protein structure comprises calculating a root mean square deviation (RMSD) value.
5. The method of any one of the preceding claims, wherein predicting a folded protein structure of the coded protein of the input sequence is performed using one or more machine learning models.
6. The method of any one of the preceding claims, wherein expressing the coded variant DNA sequences is performed in a digital microfluidic device.
7. The method of any one of the preceding claims, wherein determining one or more optimal input sequences based on the expressed one or more coded variants comprises determining one or more assay metrics by measuring expression yield, solubility, stability or performing one or more protein assays on each expressed coded variant.
8. The method of any one of the preceding claims, wherein receiving the input data comprises receiving a sequence coding for the target protein, and determining one or more input sequences that each code for variants of the target protein.
9. The method of any one of the preceding claims, wherein identifying one or more preferred input sequences comprises identifying one or more input sequences for which the predicted folded protein structure is most structurally similar to the target protein.
10. The method of any one of the preceding claims, wherein identifying one or more preferred input sequences comprises identifying one or more input sequences for which the predicted folded protein structure is structurally similar to the target protein, but the length of the target protein is reduced through truncations.
11. The method of any one of the preceding claims comprising receiving input data comprising a plurality of input sequences, each of the input sequences coding for an expressed protein sequence having a target protein and one or more binding or detection tags; for each of the DNA sequences in the plurality of DNA sequences: predicting a folded protein structure of the expressed protein of the DNA sequence; and determining whether the binding or detection tags are accessible for binding; identifying one or more preferred input sequences based on the accessibility of binding or detection tags; expressing the protein sequence of each of the identified one or more preferred input sequences in a cell-free system; and
using the tags for binding to a moiety which enables detection of expression or purification of the protein.
12. The method according to claim 11, wherein the proteins of interest have binding tag and a detection tag and the predicted structures show both tags are accessible in the predicted folded protein structure.
13. The method according to claim 11 or claim 12, wherein one of the tags is a purification tag.
14. The method according to any one of claims 11 to 13, wherein one of the tags is a tag which binds to a detector species which is a sub-component of a fluorescent protein.
15. The method of claim 14, wherein the tag is selected from GFPn, sfGFPn, or ccGFPn.
16. A method according to any one preceding claim for identifying protein variants, the method comprising: designing multiple input sequences that code for variants of a target protein wherein the target proteins have flank sequences allowing for protein expression, the flank sequences including binding or detection tags; screening the input sequences through software which predicts folded protein structure to determine structural similarity between the coded protein and the target protein; identifying the coded proteins which are structurally not affected by the presence of binding or detection tags; taking the identified input sequences and expressing them in a cell-free system to yield protein variants having an available binding or detection tag; and using the tags for binding to a moiety which enables detection of expression or purification of the protein.
17. A system for analysing protein variants, the system comprising: an input module configured to receive input data comprising a plurality of input sequences, each of the input sequences coding for a variant of a target protein; a processing module configured to receive the input data from the input module, and provide the input data to a sequence analysis module; a sequence analysis module configured to:
for each of the input sequences in the plurality of input sequences: predict a folded protein structure of the coded protein of the input sequence; and determine a structural similarity between the predicted folded protein structure and the target protein; the processing module being further configured to receive the determined structural similarities from the sequence analysis module, and identify one or more preferred input sequences based on the determined structural similarities; a protein expression module configured to express and characterise the coded variant of each of the identified one or more preferred input sequences; the processing module being further configured to determine one or more optimal input sequences based on the expressed one or more coded variants.
18. The system of claim 17, wherein the sequence analysis module is configured to determine a structural similarity between the predicted folded protein structure and the target protein by calculating a root mean square deviation (RMSD) value.
19. The system of claim 17 or claim 18, wherein the sequence analysis module is configure to predict a folded protein structure of the coded protein of the DNA sequence using one or more machine learning models.
20. The system of any one of claims 16 to 19, wherein the protein expression module is configured to express the coded variant of each of the identified one or more preferred DNA sequences in a cell-free expression system on a digital microfluidic device.
21. The system of any one of claims 17 to 20, wherein the protein expression module is configured to determine one or more optimal input sequences based on the expressed one or more coded variants by determining one or more assay metrics by performing one or more measurements on each expressed coded variant.
22. The system of any one of claims 17 to 21, wherein receiving the input data comprises receiving an input sequence coding for the target protein, and determining one or more DNA sequences that each code for a variant of the target protein.
23. The system of any one of claims 17 to 22, wherein the processing module is configured to identify one or more preferred input sequences by identifying one or more input sequences for which the predicted folded protein structure is most structurally similar to the target protein.
24. The system of any one of claims 17 to 23, wherein the a protein expression module measures the yield of expression by binding an accessible tag to a detector species which is a sub-component of a fluorescent protein.
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| GBGB2303808.6A GB202303808D0 (en) | 2023-03-15 | 2023-03-15 | System and method for protein sequence screening |
| PCT/GB2024/050726 WO2024189383A1 (en) | 2023-03-15 | 2024-03-15 | System and method for protein sequence screening |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| EP4681203A1 true EP4681203A1 (en) | 2026-01-21 |
Family
ID=86052795
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| EP24714556.8A Pending EP4681203A1 (en) | 2023-03-15 | 2024-03-15 | System and method for protein sequence screening |
Country Status (3)
| Country | Link |
|---|---|
| EP (1) | EP4681203A1 (en) |
| GB (1) | GB202303808D0 (en) |
| WO (1) | WO2024189383A1 (en) |
Family Cites Families (6)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20040010376A1 (en) * | 2001-04-17 | 2004-01-15 | Peizhi Luo | Generation and selection of protein library in silico |
| US20150099271A1 (en) * | 2013-10-04 | 2015-04-09 | Los Alamos National Security, Llc | Fluorescent proteins, split fluorescent proteins, and their uses |
| CA3075408C (en) | 2017-10-18 | 2022-06-28 | E Ink Corporation | Digital microfluidic devices including dual substrates with thin-film transistors and capacitive sensing |
| US20210043272A1 (en) | 2018-02-26 | 2021-02-11 | Just Biotherapeutics, Inc. | Determining protein structure and properties based on sequence |
| JP7128346B2 (en) | 2018-09-21 | 2022-08-30 | ディープマインド テクノロジーズ リミテッド | Determining a protein distance map by combining distance map crops |
| GB202013063D0 (en) | 2020-08-21 | 2020-10-07 | Nuclera Nucleics Ltd | Real-time monitoring of in vitro protein synthesis |
-
2023
- 2023-03-15 GB GBGB2303808.6A patent/GB202303808D0/en not_active Ceased
-
2024
- 2024-03-15 WO PCT/GB2024/050726 patent/WO2024189383A1/en not_active Ceased
- 2024-03-15 EP EP24714556.8A patent/EP4681203A1/en active Pending
Also Published As
| Publication number | Publication date |
|---|---|
| GB202303808D0 (en) | 2023-04-26 |
| WO2024189383A1 (en) | 2024-09-19 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Brown et al. | Characterizing protein‐protein interactions by sedimentation velocity analytical ultracentrifugation | |
| EP3727692B1 (en) | Droplet interfaces in electro-wetting devices | |
| Si et al. | A nanoparticle-DNA assembled nanorobot powered by charge-tunable quad-nanopore system | |
| US20240359181A1 (en) | Methods and compositions for improved biomolecule assays on digital microfluidic devices | |
| US20250146041A1 (en) | Methods for cell free protein synthesis and post translational modification of the expressed proteins | |
| EP4681203A1 (en) | System and method for protein sequence screening | |
| TWI797601B (en) | Digital microfluidic device and method of driving a digital microfluidic system | |
| US20240352451A1 (en) | A method of loading devices using electrowetting | |
| US20250214081A1 (en) | Controlled reservoir filling | |
| US20260072018A1 (en) | Protein binding assays | |
| US20250196132A1 (en) | Loading and formation of multiple reservoirs | |
| US20260092925A1 (en) | Improved fluorescent proteins | |
| Bauer et al. | E. coli RNA polymerase pauses during initial transcription | |
| US12558690B2 (en) | Method of electrowetting | |
| WO2023161640A1 (en) | Monitoring of in vitro protein synthesis | |
| EP4689663A1 (en) | Protein expression systems | |
| WO2024236295A1 (en) | A method to homogenize the concentration of beads across multiple liquid volumes | |
| EP4649316A1 (en) | Protein aggregation assays | |
| US20250360503A1 (en) | Controlled reservoir filling | |
| US20250171822A1 (en) | Monitoring of in vitro protein synthesis | |
| GB2629179A (en) | System and method for detecting and reporting errors in a digital microfluidic experiment | |
| WO2025017327A2 (en) | Protein expression reagents | |
| WO2025037122A1 (en) | Protein expression reagents | |
| HK40089114A (en) | Monitoring of in vitro protein synthesis | |
| Anazawa et al. | Novel concept of Escherichia coli minimum genome cell factory |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: UNKNOWN |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: THE INTERNATIONAL PUBLICATION HAS BEEN MADE |
|
| PUAI | Public reference made under article 153(3) epc to a published international application that has entered the european phase |
Free format text: ORIGINAL CODE: 0009012 |
|
| STAA | Information on the status of an ep patent application or granted ep patent |
Free format text: STATUS: REQUEST FOR EXAMINATION WAS MADE |
|
| 17P | Request for examination filed |
Effective date: 20251006 |
|
| AK | Designated contracting states |
Kind code of ref document: A1 Designated state(s): AL AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HR HU IE IS IT LI LT LU LV MC ME MK MT NL NO PL PT RO RS SE SI SK SM TR |