WO2024252392A1 - Computational models for effecting copy number of plasmids - Google Patents
Computational models for effecting copy number of plasmids Download PDFInfo
- Publication number
- WO2024252392A1 WO2024252392A1 PCT/IL2024/050552 IL2024050552W WO2024252392A1 WO 2024252392 A1 WO2024252392 A1 WO 2024252392A1 IL 2024050552 W IL2024050552 W IL 2024050552W WO 2024252392 A1 WO2024252392 A1 WO 2024252392A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- rna
- plasmid
- copy number
- promoter
- complexation
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N15/00—Mutation or genetic engineering; DNA or RNA concerning genetic engineering, vectors, e.g. plasmids, or their isolation, preparation or purification; Use of hosts therefor
- C12N15/09—Recombinant DNA-technology
- C12N15/63—Introduction of foreign genetic material using vectors; Vectors; Use of hosts therefor; Regulation of expression
- C12N15/79—Vectors or expression systems specially adapted for eukaryotic hosts
- C12N15/85—Vectors or expression systems specially adapted for eukaryotic hosts for animal cells
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2800/00—Nucleic acids vectors
- C12N2800/10—Plasmid DNA
- C12N2800/101—Plasmid DNA for bacteria
-
- C—CHEMISTRY; METALLURGY
- C12—BIOCHEMISTRY; BEER; SPIRITS; WINE; VINEGAR; MICROBIOLOGY; ENZYMOLOGY; MUTATION OR GENETIC ENGINEERING
- C12N—MICROORGANISMS OR ENZYMES; COMPOSITIONS THEREOF; PROPAGATING, PRESERVING, OR MAINTAINING MICROORGANISMS; MUTATION OR GENETIC ENGINEERING; CULTURE MEDIA
- C12N2830/00—Vector systems having a special element relevant for transcription
Definitions
- the present invention is in the field of synthetic biology, and biotechnology, and relates to methods of controlling gene expression by tuning plasmid copy number.
- Plasmid copy number refers to the number of plasmid copies in a single cell. Plasmid copy number is a crucial factor in synthetic biology, where plasmids are commonly used to transfer and express genes of interest in various organisms. Plasmid copy number can impact the level of protein expression or the level of expression of RNA molecules and help regulate gene expression. Therefore, controlling the copy number can be crucial in designing and optimizing synthetic biology systems.
- a method for maximizing plasmid copy number in a host cell comprising: (a) modeling origin of replication (orz) associated interactions the affect plasmid copy number in the host cell; (b) determining free energy, affinity, abundance, or any combination thereof, of the ori associated interactions; and (c) based on the determined free energy, affinity, abundance, or any combination thereof of step (b) introducing to the plasmid any one of: i. at least one mutation increasing the affinity, reducing the free energy, increasing the abundance, or any combination thereof, of ori associated interactions supporting high plasmid copy number of the plasmid; ii. at least one mutation decreasing the affinity, increasing the free energy, reducing the abundance, or any combination thereof, of ori associated interactions suppressing high plasmid copy number of the plasmid; or iii. both (i) and (ii).
- a method for optimizing a copy number of a plasmid to host cell comprising: (a) calculating an activity index for a host cell comprising a pre-determined copy number of the plasmid, wherein an activity index above or below a cut-off range is indicative of the pre-determined copy number of the plasmid being suboptimal for expression of a polynucleotide being integrated into the plasmid in the host cell; and (b) introducing to the plasmid at a pre-determined copy number determined as being suboptimal for expression of the polynucleotide in the host cell any one of: i.
- At least one mutation increasing the affinity, reducing the free energy, increasing the abundance, or any combination thereof, of ori associated interactions supporting high plasmid copy number of the plasmid; ii. at least one mutation decreasing the affinity, increasing the free energy, reducing the abundance, or any combination thereof, of ori associated interactions suppressing high plasmid copy number of the plasmid; or iii.
- RNA I - RNA II complexation RNA II - DNA complexation
- RNA I - RNA II - rop complexation RNA I - tRNA complexation
- RNA II - tRNA complexation RNA II - RNase H complexation
- RNA II - sigma factor complexation RNA II - sigma factor - RNA polymerase, and any combination thereof.
- a method for obtaining a predetermined copy number of a plasmid in a host cell for a user in need thereof comprising introducing at least one mutation in the ori of the plasmid, a constituent of the ori, or both, thereby tuning the copy number of the plasmid to a pre-determined copy number in the host cell for the user.
- the ori associated interactions are selected from the group consisting of: RNA I - RNA II complexation, RNA II - DNA complexation, RNA I - RNA II - repressor of primer (rop) complexation, RNA I - tRNA complexation, RNA II - tRNA complexation, RNA II - RNase H complexation, RNA II - sigma factor complexation, RNA II - sigma factor - RNA polymerase, and any combination thereof.
- the at least one mutation forms an RpoD sigma factor recognition site in RNA I of the ori.
- the plasmid comprises a polynucleotide integrated therein.
- the optimizing comprises increasing or decreasing the copy number of the plasmid such that it is optimal for expression of the polynucleotide in the host cell.
- the tRNA is an uncharged tRNA.
- the plasmid is a ColEl-like plasmid.
- the at least one mutation comprises: substitution, deletion, insertion, or any combination thereof.
- the ori comprises: RNA I, RNA II, rop, an uncharged tRNA, a promoter thereof, or any combination thereof.
- the at least one mutation is introduced in RNA I or a promoter thereof, and wherein any one of: the mutated RNA I is a deleterious RNA I, the RNA I is characterized by reduced affinity to RNA II compared to a control, the mutated promoter is characterized by increased or decreased transcription rate compared to a control, and any combination thereof.
- the mutation is in RNA II or a promoter thereof, and wherein any one of: the mutated RNA II is characterized by reduced affinity to RNA I compared to a control, the mutated RNA II is characterized by increased affinity to DNA compared to a control, the mutated promoter is characterized by increased or decreased transcription rate compared to a control, the mutated RNA II is mutated at a cleavage site of RNase H, and any combination thereof.
- the mutation is in rop, and wherein the mutated rop is characterized by increased or reduced affinity to a complex comprising RNA I and RNA II (kissing complex) compared to a control.
- the mutated rop stabilizes or destabilizes said kissing complex.
- the method further comprises a step proceeding the introducing comprising contacting the cell with the plasmid comprising the at least one mutation.
- the obtaining comprises increasing or decreasing the copy number of the plasmid such that it is suitable or optimal for expression of a polynucleotide being integrated into the plasmid in the cell.
- the polynucleotide encodes a polypeptide of interest.
- Fig. 1 includes a plot and a horizontal bar graph showing Lasso For RNAp: R- squared equals 0.31 with the optimal alpha of 0.127.
- Fig. 2 includes a plot and a horizontal bar graph showing Lasso for RNAi: R-squared equals 0.13 with the optimal alpha of 1.162.
- Fig. 3 includes a plot and a horizontal bar graph showing Lasso for the shared model: R-squared equals 0.196 with the optimal alpha of 0.74.
- Fig. 4 includes a horizontal bar graph and a plot showing XGBoost for RNAp: R- squared equals 0.908.
- Fig. 5 includes a plot and a horizontal bar graph showing XGBoost for RNAi: R- squared equals 0.489.
- Fig. 6 includes a plot and a horizontal bar graph showing XGBoost for RNAi: R- squared equals 0.839.
- Fig. 7 includes a plot showing linear regression of feature selection using forward sequential feature selection of for priming RNA.
- Fig. 8 includes a plot showing linear regression of feature selection using forward sequential feature selection of for inhibitor RNA.
- Fig. 9 includes graphs showing XGBoost complexity analysis.
- the XGBoost model was trained with different numbers of trees with a range of 10-300 trees and default maximum depth of 6, and with different values of maximum depth with a range of 2-20 and a default number of trees of 100.
- the R2 and Spearman correlation between the true values of the test set and the predicted values of each model were calculated and displayed.
- Fig. 10 includes bar graphs showing features comparison between the models.
- the models were trained 1,000 times and the 40 features of each model with the highest features importance score on average were intercrossed.
- the y axis represents the feature importance of the two models: F score for XGBoost and the coefficient for Lasso. The common features are displayed in the graphs.
- Fig. 11 includes dot plots showing correlation of all features correlation by importance (XGboost) and average (Lasso) in inhibitor RNA (RNAi) and priming RNA (RNAp).
- Fig. 12 includes dot plots showing correlation of 40 highest features correlation by importance (XGboost) and average (Lasso) in inhibitor RNA (RNAi) and priming RNA (RNAp).
- Fig. 13 includes dot plots showing correlation of all features correlation by importance (XGboost) and average (Lasso) in inhibitor RNA (RNAi) and priming RNA (RNAp).
- Fig. 14 includes dot plots showing correlation of 100 highest features correlation by importance (XGboost) and average (Lasso) in inhibitor RNA (RNAi) and priming RNA (RNAp).
- Fig. 15 includes correlation between top MIC features and copy number.
- Fig. 16 includes correlation between top Pearson features and copy number.
- Fig. 17 includes correlation between top Spearman features and copy number.
- Fig. 18 includes plots showing promoter correlations. A correlation of MIC being 0.3 and 0.23 with a significant p-value is observed for RNAi, and for RNAp, respectively.
- Fig. 19 includes plots showing sigma factor of RNA polymerase correlations.
- the rpoD17 and rpoD18 motifs are mutations found in the gene encoding the sigma factor of RNA polymerase in the bacterium Escherichia coli. These mutations affect the ability of the sigma factor to bind to DNA and initiate transcription, leading to changes in gene expression patterns in E. coli.
- Fig. 20 includes graphs showing AA count and TT count. This feature represents the number of times ‘AA’ or ‘TT’ being presented in the promoter sequence.
- Fig. 21 includes PSSM score calculated with the 20% high copy number only, for RNAp.
- the presented sequence is set forth in SEQ ID NO: 10.
- Fig. 22 includes PSSM score calculated with the 20% high copy number only, for RNAi.
- Fig. 23 includes vertical bar graphs showing the number of times each feature was used in the models out of 5,000 iterations is represented by count (the height of the bin), the feature importance is represented by the features orders on the x axis (from left to right) and the average R2 score of the models in the iterations the feature was used is represented by the color bar.
- Fig. 24 includes a table listing the statistical information of the R2 score of Fig. 23.
- Figs. 25A-25F include graphs, a table, and horizontal bar graphs.
- 25C A table of the results of 25A and 25B.
- RNA feature importance for CatBoost model the higher the feature importance score is the higher feature contribution to the predicted copy number.
- (25E) Priming RNA feature importance for XGBoost model, the higher the feature importance score is the higher feature contribution to the predicted copy number.
- Figs. 26A-26D include plots and horizontal bar graphs.
- 26A Features distribution corresponding to their copy number for the common features of XGBoost and CatBoot in the voting model for RNAp.
- 26B Features distribution corresponding to their copy number for the 2 most important features CatBoot in RNAi model.
- 26C The absolute mean SHAP value for each feature chosen during the feature selection process in the model based on the RNAp model.
- 26D The absolute mean SHAP value for each feature chosen during the feature selection process in the model based on the RNAi model.
- Figs. 27A-27B include graphs.
- Figs. 28A-28D include illustrations, a graph, and a sequence.
- 28A A modified origin of replication is synthesized from tool-generated data, followed by PCR amplification of pUC19 to incorporate new flanking regions around the ORI. Subsequently, a restriction enzyme is utilized to target a specific cut site in the original pUC19 sequence. The modified ORI is integrated into the pUC19 backbone via homologous joining, resulting in successful transformation of the modified pUC19 vectors Created with BioRender.
- Figs. 29A-29D include illustrations, a flow chart, and sequences.
- 29A The interaction strengths between RNA polymerase and the subunit of RNA polymerase c70 factor and promoter DNA.
- 29B RNA polymerase binding energy matrix. Colors indicate the contribution of each base-pair to the total binding energy.
- 29C Model training and evaluation workflow.
- 29D A few examples of RNAp variants and their corresponding copy number. All the variants are mutated in positions -33 to -30, -8 to -11, and +1. Sequence in 29B is set forth in SEQ ID NO: 12. The sequences of the variants of 29D are set forth in SEQ ID Nos: 13-18. [056] Figs.
- 30A-30B include vertical bar graphs showing the models performance in the model selection section.
- (30A) RNAp Models exploration showing Pearson correlation and Spearman correlation results on train vs validation. Chosen models are surrounded by a rectangle. The final model is a voting model that utilizes these models (see more in methods).
- (30B) RNAi Models exploration showing Pearson correlation and Spearman correlation results on train vs validation. Chosen models for the final model are surrounded by a rectangle.
- the method comprises introducing at least one mutation in the origin of replication (orz) of the plasmid, a constituent of the ori, or both, thereby tuning a copy number of a plasmid to a pre-determined copy number in a cell.
- At least one mutation comprises: substitution, deletion, insertion, or any combination thereof.
- a constituent of an ori comprises: RNA I, RNA II, repressor of primer (rop), an uncharged transfer RNA (tRNA), a promoter thereof, or any combination thereof.
- tRNA is or comprises an uncharged tRNA.
- the at least one mutation is introduced in RNA I or a promoter thereof.
- the mutated RNA I is a deleterious RNA I.
- the RNA I is characterized by reduced affinity to RNA II compared to a control.
- the mutated promoter is characterized by increased or decreased transcription rate compared to a control.
- the mutated RNA I is a deleterious RNA I, the mutated RNA I is characterized by reduced affinity to RNA II compared to a control, the mutated promoter is characterized by increased or decreased transcription rate compared to a control, or any combination thereof.
- the mutation is in RNA II or a promoter thereof.
- the mutated RNA II is characterized by reduced affinity to RNA I compared to a control. In some embodiments, the mutated RNA II is characterized by increased affinity to DNA compared to a control. In some embodiments, the mutated promoter is characterized by increased or decreased transcription rate compared to a control. In some embodiments, the mutated RNA II is mutated at a cleavage site of RNase H.
- the mutated RNA II is characterized by reduced affinity to RNA I compared to a control, the mutated RNA II is characterized by increased affinity to DNA compared to a control, the mutated promoter is characterized by increased or decreased transcription rate compared to a control, the mutated RNA II is mutated at a cleavage site of RNase H, or any combination thereof.
- the mutation is in rop.
- a mutated rop is characterized by increased or reduced affinity to a complex comprising RNA I and RNA II (kissing complex) compared to a control.
- a mutated rop stabilizes or destabilizes a kissing complex.
- the method further comprises a step proceeding the introducing, comprising contacting a cell with a plasmid comprising the at least one mutation.
- RNA II is an RNA primer or pre -primer (RNAp).
- RNA I is an inhibitory or interfering RNA (RNAi).
- RNA I and RNA II are at least partially complementary (to one another). In some embodiments, RNA I is at least partially complementary to RNA II. In some embodiments, complementary comprises reverse and complementary. In some embodiments, RNA I is complementary to RNA II in a level that effectively or sufficiently reduces or inhibits hybridization of RNA II to the plasmid. In some embodiments, RNA I is characterized by a complementarity level to RNA II that effectively or sufficiently reduces or inhibits hybridization of RNA II to the plasmid.
- the plasmid is or comprises a ColEl-like plasmid.
- Types and sources of ColEl-like plasmids are common and would be apparent to one of ordinary skill in the art, such as reviewed by Chaillou et al., Nucleic Acids Res. 2022; 50(16): 9568-9579.
- the plasmid is or comprises pUC19.
- a ColEl-like plasmid is or comprises pUC19.
- At least partially complementary comprises 50-100%, 60- 100%, 70-100%, 80-100%, or 90-100% complementary. Each possibility represents a separate embodiment of the invention. In some embodiments, complementarity is over or in an equal-length portion.
- the obtaining, tuning, or both comprises increasing or decreasing the copy number of the plasmid such that it is suitable or optimal for expression of the polynucleotide in the host cell. In some embodiments, the obtaining, tuning, or both is to a particular or definitive copy number of the plasmid.
- the polynucleotide encodes a polypeptide of interest.
- the polypeptide comprises a recombinant polypeptide.
- the cell or the host cell is further cultured under conditions sufficient for expression of the polypeptide of interest. In some embodiments, the cell or the host cell is further cultured under conditions such that the polypeptide of interest sufficiently, effectively, or both, is expressed, secreted, or both. In some embodiments, the cell or the host cell is further cultured under conditions such that the polypeptide of interest is expressed, secreted, or both.
- the methods further comprises culturing the cell comprising the plasmid being modified, optimized, tuned, maximized, or any combination thereof for copy number, according to the method of the invention.
- the method further comprises isolating, purifying, retrieving, or any combination thereof, the polypeptide of interest from the cultured cell or cultured host cell, or a medium in which the cell or host cell being cultured.
- the expression “of interest” is to a user or a human user.
- the polypeptide of interest comprises any polypeptide, the expression of a recombinant version thereof is of importance to the user. In some embodiments, importance is commercial importance.
- the polypeptide is a product or an ingredient thereof. In some embodiments, the polypeptide is a therapeutic agent or an active ingredient thereof.
- the terms “peptide”, “polypeptide” and “protein” are used interchangeably to refer to a polymer of amino acid residues.
- the terms “peptide”, “polypeptide” and “protein” as used herein encompass native peptides, peptidomimetics (typically including non-peptide bonds or other synthetic modifications) and the peptide analogues peptoids and semipeptoids or any combination thereof.
- the peptides polypeptides and proteins described have modifications rendering them more stable while in the body or more capable of penetrating into cells.
- the terms “peptide”, “polypeptide” and “protein” apply to naturally occurring amino acid polymers.
- the terms “peptide”, “polypeptide” and “protein” apply to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid.
- the method comprises: (a) modeling ori associated interactions the affect plasmid copy number in the host cell; (b) determining free energy, affinity, abundance, or any combination thereof, of the ori associated interactions; and (c) based on the determined free energy, affinity, abundance, or any combination thereof of step (b) introducing to the plasmid at least one mutation.
- RNA II Ori associated interaction of RNA II or its promoter RNAp
- the modeling comprises determining: (i) free energy of promoter binding; (ii) abundance of complexes comprising the promoter and sigma factor, RNA polymerase, or both; or (iii) any combination of (i) and (ii).
- the promoter comprises a promoter of RNA II.
- the complex comprises a promoter of RNA II and sigma factor, a promoter of RNA II and RNA polymerase, or a promoter of RNA II, sigma factor, and RNA polymerase.
- the abundance reflects and the strength of sigma factor recognition sites. In some embodiments, the strength of sigma factor recognition sites is reflected by the abundance, e.g., of the complexes disclosed herein.
- the promoter, e.g., RNA II comprises one or more sigma factor recognition sites. In some embodiments, the promoter, e.g., RNA II comprises at least one sigma factor recognition sites. In some embodiments, the promoter, e.g., RNA II comprises a plurality of sigma factor recognition sites.
- the method comprises determining an ori sequence of a plasmid.
- the method further comprises extracting, from the determined ori sequence, a set of features representing transcription related properties.
- the set of features comprises at least one of: motif score; P-values; PSSM scores (reflecting frequency of nucleotides in high copy number sequences); promoter strength; nucleotide counts; interaction transcription rate measures, or any combination thereof.
- the method comprises determining at least one feature related to the free energy of the promoter.
- the set of features comprises at least one feature related to the free energy of the promoter.
- the free energy is of an unbound promoter.
- the free energy is of a bound promoter.
- the free energy is the delta free energy of a bound promoter compared to free energy of an unbound promoter, or vice versa.
- a bound promoter is bound to a transcription factor or a plurality thereof, a DNA binding protein, RNA polymerase, or any combination thereof.
- a transcription factor or a plurality thereof, a DNA binding protein, or both comprises a sigma factor.
- free energy is defined or reflected as delta G (dG or AG).
- free energy is dG is total dG (e.g., dG_total).
- total dG comprises or refers to the overall change in free energy between an unbound promoter, such as of RNA II, and a promoter, such as of RNA II bound to a sigma factor, RNA polymerase, or both.
- total dG comprises or refers to the overall change in free energy between an unbound RNA II promoter and an RNA II promoter being bound to a sigma factor, RNA polymerase, or both.
- the method further comprises determining the abundance of complexes comprising the promoter and: sigma factor, RNA polymerase, or both, comprises determining the presence of the rpoD16 motif (a sigma factor recognition site) in the promoter.
- the method further comprises determining rpoD16_score as a numerical value representing the strength or likelihood of the presence of the rpoD16 motif in the promoter.
- the set of features comprises a feature representing the abundance or presence of rpoD16 motif in the promoter.
- a , or, more specifically, a feature representing the abundance or presence of rpoD16 motif in the promoter is or comprises rpoD16_score.
- the method further comprises calculating a vector representation of the determined ori sequence, based on the determined set of features, thereby performing an embedding of the ori sequence into a respective feature space.
- the method further comprises inferring a respectively pretrained machine learning (ML)-based model on the calculated vector representation to predict the plasmid copy number.
- ML machine learning
- the pretrained ML -based model comprises one of gradient boosting algorithms known in the art.
- an ML-based model is based on Categorical Boosting (CatBoost) algorithm.
- CatBoost Categorical Boosting
- XGBoost Extreme Gradient Boosting
- ML-based model is pretrained based on samples representing mutations in the RNAp (alternatively used herein for RNA II) promoter region.
- the method further comprises generating at least one mutation (or analyzing a suggested mutation) to be introduced to the plasmid, based on the predicted plasmid copy number.
- At least one mutation reducing dG to about -3.7 to -2.5 increases copy number of the plasmid. In some embodiments, the at least one mutation increasing copy number of the plasmid reduces dG to about -3.7 to -2.5. In some embodiments, at least one mutation reducing dG to about -3.8 or less reduces copy number of the plasmid. In some embodiments, the at least one mutation reducing copy number of the plasmid reduces dG to about -3.8 or less. In some embodiments, at least one mutation increasing dG to about -2.4 or more reduces copy number of the plasmid. In some embodiments, the at least one mutation reducing copy number of the plasmid increases dG to about -2.4 or more.
- dG of about -3.7 to -2.5 is or comprises about -3.4 to -2.7.
- RNAi Ori associated interaction of RN A I
- the method comprises determining the probability of adjacent RNA nucleotides of RNA I to hybridize. In some embodiments, the determining comprises determining the probability of secondary structure formation(s) of RNA I. In some embodiments, the determining is by using RNAFold or an equivalent thereof. In some embodiments, the determining comprises determining the probability of the nucleotide in index 375 to pair with the nucleotide in index 385 of the first 400 bases of RNA I (e.g., seq_end_400_375_385).
- the set of features further comprises the feature representing the determined probability of secondary structure(s) formation (RNAFold).
- the determined probability of secondary structure(s) formation is further analyzed using Catboost. In some embodiments, the method further comprises analyzing the determined probability of secondary structure(s) formation using Catboost.
- the method comprises determining the presence of the de- novo motif in a promoter of RNA I.
- the de-novo motif comprises the nucleic acid sequence: TGCGCGTTTAAT (SEQ ID NO: 7) or having 80-100% sequence identity thereto.
- the de-novo motif is present with higher rates in promoters of high copy number plasmids, compared to control promoters of low copy number plasmids.
- the method comprises determining the presence of the nucleic acid sequence: TGCGCGTTTAAT (SEQ ID NO: 7) in a promoter of RNA I.
- the method further comprises analyzing the determined presence of the de-novo motif in a promoter of RNA I.
- the set of features further comprises a feature representing a presence of the nucleic acid sequence: TGCGCGTTTAAT (SEQ ID NO: 7).
- the ML-based model is pretrained based on samples representing mutations in the RNAi (as used herein interchangeably with RNA I) promoter region.
- the abovementioned variants of ML-based model implementations may be used in combination, e.g., be arranged as an ensemble of ML-based models, for further improvement of the plasmid copy number prediction.
- the at least one mutation comprises any one of: (i) at least one mutation increasing the affinity, reducing the free energy, increasing the abundance, or any combination thereof, of ori associated interactions supporting high plasmid copy number of the plasmid; (ii) at least one mutation decreasing the affinity, increasing the free energy, reducing the abundance, or any combination thereof, of ori associated interactions suppressing high plasmid copy number of the plasmid; or (iii) both (i) and (ii).
- the ori associated interactions are selected from: RNA I - RNA II complexation, RNA II - DNA complexation, RNA I - RNA II - rop complexation, RNA I - tRNA complexation, RNA II - tRNA complexation, RNA II - RNase H complexation, RNA II - sigma factor complexation, RNA II - sigma factor - RNA polymerase, or any combination thereof.
- the at least one mutation forms or introduces an RpoD sigma factor recognition site in the RNA II promoter.
- the at least one mutation increases the abundance or presence of an RpoD sigma factor recognition site(s) in the RNA II promoter. In some embodiments, the at least one mutation forms or introduces the nucleic acid sequence TGCGCGTTTAAT (SEQ ID NO: 7) in a promoter of RNA I. In some embodiments, the at least one mutation increases the abundance or presence of the nucleic acid sequence TGCGCGTTTAAT (SEQ ID NO: 7) in a promoter of RNA I. In some embodiments, the at least one mutation forms or introduces at least one secondary structure in RNA I or a promoter thereof. In some embodiments, the at least one mutation increases the abundance or presence of at least one secondary structure in RNA I or a promoter thereof. In some embodiments, at the least one mutation reduces the total dG of a complex comprising the RNA II promoter and: a sigma factor, RNA polymerase, or both.
- increase the abundance or presence of a feature as disclosed herein, e.g., dG, rpol6D, secondary structure, SEQ ID NO: 7, etc. comprises increase the probability to determine the presence of the feature.
- increase the abundance or presence of a feature as disclosed herein, e.g., dG, rpol6D, secondary structure, SEQ ID NO: 7, etc. means that the probability to determine the presence of the feature is increased.
- reduces or reduce or increase or increases is compared to a control promoter being: non-modified, non-mutated, wildtype, naive, or any combination thereof.
- the method comprises: (a) calculating an activity index for a host cell comprising a pre-determined copy number of the plasmid, wherein an activity index above or below a cut-off range is indicative of the pre-determined copy number of the plasmid being suboptimal for expression of a polynucleotide being integrated into the plasmid in the host cell; and (b) introducing to the plasmid at a pre-determined copy number determined as being suboptimal for expression of the polynucleotide in the host cell any one of: (i) at least one mutation increasing the affinity of ori associated interactions supporting high plasmid copy number of the plasmid; (ii) at least one mutation decreasing the affinity of ori associated interactions suppressing high plasmid copy number of the plasmid; or (iii) both (i) and (ii).
- the optimizing comprises increasing or decreasing the copy number of the plasmid such that it is optimal for expression of the polynucleotide in the host cell. In some embodiments, the optimizing is to a particular or definitive copy number of the plasmid.
- an activity index equivalent to a cut-off range is indicative of the pre-determined copy number of the plasmid being optimal for expression of a polynucleotide being integrated into the plasmid in the host cell.
- the method is an in vitro method. In some embodiments, the method is a computerized method.
- expression refers to the biosynthesis of a gene product, including the transcription and/or translation of said gene product.
- expression of a nucleic acid molecule may refer to transcription of the nucleic acid fragment (e.g., transcription resulting in mRNA or other functional RNA) and/or translation of RNA into a precursor or mature protein (polypeptide).
- a gene within a cell is well known to one skilled in the art. It can be carried out by, among many methods, transfection, viral infection, or direct alteration of the cell’s genome.
- the gene is in an expression vector such as plasmid or viral vector.
- an expression vector containing pl6-Ink4a is the mammalian expression vector pCMV pl6 INK4A available from Addgene.
- a vector nucleic acid sequence generally contains at least an origin of replication for propagation in a cell and optionally additional elements, such as a heterologous polynucleotide sequence, expression control element (e.g., a promoter, enhancer), selectable marker (e.g., antibiotic resistance), poly-Adenine sequence.
- expression control element e.g., a promoter, enhancer
- selectable marker e.g., antibiotic resistance
- the vector may be a DNA plasmid delivered via non-viral methods or via viral methods.
- the viral vector may be a retroviral vector, a herpesviral vector, an adenoviral vector, an adeno-associated viral vector or a poxviral vector.
- the promoters may be active in mammalian cells.
- the promoters may be a viral promoter.
- the gene is operably linked to a promoter.
- operably linked is intended to mean that the nucleotide sequence of interest is linked to the regulatory element or elements in a manner that allows for expression of the nucleotide sequence (e.g. in an in vitro transcription/translation system or in a host cell when the vector is introduced into the host cell).
- the vector is introduced into the cell by standard methods including electroporation (e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)), Heat shock, infection by viral vectors, high velocity ballistic penetration by small particles with the nucleic acid either within the matrix of small beads or particles, or on the surface (Klein et al., Nature 327. 70-73 (1987)), and/or the like.
- electroporation e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)
- Heat shock e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)
- infection by viral vectors e.g., as described in From et al., Proc. Natl. Acad. Sci. USA 82, 5824 (1985)
- Heat shock
- promoter refers to a group of transcriptional control modules that are clustered around the initiation site for an RNA polymerase i.e., RNA polymerase II. Promoters are composed of discrete functional modules, each consisting of approximately 7-20 bp of DNA, and containing one or more recognition sites for transcriptional activator or repressor proteins.
- nucleic acid sequences are transcribed by RNA polymerase II (RNAP II and Pol II).
- RNAP II is an enzyme found in eukaryotic cells. It catalyzes the transcription of DNA to synthesize precursors of mRNA and most snRNA and microRNA.
- mammalian expression vectors include, but are not limited to, pcDNA3, pcDNA3.1 ( ⁇ ), pGL3, pZeoSV2( ⁇ ), pSecTag2, pDisplay, pEF/myc/cyto, pCMV/myc/cyto, pCR3.1, pSinRep5, DH26S, DHBB, pNMTl, pNMT41, pNMT81, which are available from Invitrogen, pCI which is available from Promega, pMbac, pPbac, pBK- RSV and pBK-CMV which are available from Strategene, pTRES which is available from Clontech, and their derivatives.
- expression vectors containing regulatory elements from eukaryotic viruses such as retroviruses are used by the present invention.
- SV40 vectors include pSVT7 and pMT2.
- vectors derived from bovine papilloma virus include pBV-lMTHA, and vectors derived from Epstein Bar virus include pHEBO, and p2O5.
- exemplary vectors include pMSG, pAV009/A+, pMTO10/A+, pMAMneo- 5, baculovirus pDSVE, and any other vector allowing expression of proteins under the direction of the SV-40 early promoter, SV-40 later promoter, metallo thionein promoter, murine mammary tumor virus promoter, Rous sarcoma virus promoter, polyhedrin promoter, or other promoters shown effective for expression in eukaryotic cells.
- recombinant viral vectors which offer advantages such as lateral infection and targeting specificity, are used for in vivo expression.
- lateral infection is inherent in the life cycle of, for example, retrovirus and is the process by which a single infected cell produces many progeny virions that bud off and infect neighboring cells.
- the result is that a large area becomes rapidly infected, most of which was not initially infected by the original viral particles.
- viral vectors are produced that are unable to spread laterally. In one embodiment, this characteristic can be useful if the desired purpose is to introduce a specified gene into only a localized number of targeted cells.
- plant expression vectors are used.
- the expression of a polypeptide coding sequence is driven by a number of promoters.
- viral promoters such as the 35S RNA and 19S RNA promoters of CaMV [Brisson et al., Nature 310:511-514 (1984)], or the coat protein promoter to TMV [Takamatsu et al., EMBO J. 3:17-311 (1987)] are used.
- plant promoters are used such as, for example, the small subunit of RUBISCO [Coruzzi et al., EMBO J.
- constructs are introduced into plant cells using Ti plasmid, Ri plasmid, plant viral vectors, direct DNA transformation, microinjection, electroporation and other techniques well known to the skilled artisan. See, for example, Weissbach & Weissbach [Methods for Plant Molecular Biology, Academic Press, NY, Section VIII, pp 421-463 (1988)].
- Other expression systems such as insects and mammalian host cell systems, which are well known in the art, can also be used by the present invention.
- the expression construct of the present invention can also include sequences engineered to optimize stability, production, purification, yield or activity of the expressed polypeptide.
- nucleic acid is well known in the art.
- a “nucleic acid” as used herein will generally refer to a molecule (i.e., a strand) of DNA, RNA or a derivative or analog thereof, comprising a nucleobase.
- a nucleobase includes, for example, a naturally occurring purine or pyrimidine base found in DNA (e.g., an adenine "A,” a guanine "G,” a thymine “T” or a cytosine “C”) or RNA (e.g., an A, a G, an uracil "U” or a C).
- polynucleotide polynucleotide sequence
- nucleic acid sequence nucleic acid molecule
- a polynucleotide may be a polymer of RNA or DNA that is single- or double-stranded, that optionally contains synthetic, non-natural or altered nucleotide bases.
- the term “primer” includes an oligonucleotide, either natural or synthetic, that is capable, upon forming a duplex with a polynucleotide template, of acting as a point of initiation of nucleic acid synthesis and being extended from its 3' end along the template so that an extended duplex is formed. Primers within the scope of the present invention bind adjacent to a target sequence.
- a “primer” may be considered a short polynucleotide, generally with a free 3'-OH group that binds to a target or template potentially present in a sample of interest by hybridizing with the target, and thereafter promoting polymerization of a polynucleotide complementary to the target.
- interaction strength refers to hybridization free energy between a nucleic acid molecule and a ribosomal RNA. Lower and more negative free energy is related to stronger hybridization and stronger interaction strength. Hybridization free energy can be computed based on the Vienna package RNAcoFold, which computes a common secondary structure of two RNA molecules. According to some embodiments, the interaction strength can be defined by a scale of strong, intermediate and weak.
- hybridization or “hybridizes” as used herein refers to the formation of a duplex between nucleotide sequences which are sufficiently complementary to form duplexes via Watson-Crick base pairing. Two nucleotide sequences are “complementary” to one another when those molecules share base pair organization homology. “Complementary” nucleotide sequences will combine with specificity to form a stable duplex under appropriate hybridization conditions.
- RNA sequences are complementary when a section of a first sequence can bind to a section of a second sequence in an anti-parallel sense wherein the 3 '-end of each sequence binds to the 5 '-end of the other sequence and each A, T (U), G and C of one sequence is then aligned with a T (U), A, C and G, respectively, of the other sequence.
- free energy refers to the Gibbs free energy (AG), reflecting the thermodynamic potential that measures the hybridization reaction between a given oligonucleotide and its DNA or RNA complement, a given DNA sequence, such as of a promoter, e.g., of RNA II, and one or more proteins, such as sigma factor, RNA polymerase, or both.
- AG Gibbs free energy
- a computer program product comprising a non-transitory computer-readable storage medium having program code embodied thereon.
- the program code being executable by at least one hardware processor to: (a) receive a sequence of a plasmid; (b) receive a list of interactions associated with the origin of replication (orz) of the plasmid; (c) calculate mutations within at least one sequence associated with an ori associated interaction that: (i) increases the affinity, reduces the free energy, increases the abundance, or any combination thereof, of ori associated interactions supporting high plasmid copy number of the plasmid; (ii) decreases the affinity, increases the free energy, reduces the abundance, or any combination thereof, of ori associated interactions suppressing high plasmid copy number of the plasmid; or both (i) and (ii); and (d) provide an output (sequence of a) plasmid characterized by a maximized copy number, an optimized copy number, or a predetermined copy number, comprising at least one sequence associated with an ori associated interaction comprising at least one calculated mutation.
- the ori associated interaction is selected from: RNA I - RNA II complexation, RNA II - DNA complexation, RNA I - RNA II - repressor of primer (rop) complexation, RNA I - tRNA complexation, RNA II - tRNA complexation, RNA II - RNase H complexation, RNA II - sigma factor complexation, RNA II - sigma factor - RNA polymerase, or any combination thereof.
- the computer program product is for maximizing plasmid copy number. In some embodiments, the computer program product is for optimizing a copy number of a plasmid to host cell. In some embodiments, the computer program product is for obtaining a pre-determined copy number of a plasmid in a host cell for a user in need thereof. [0148] In some embodiments, the computer program product calculates mutation within at least one sequence associated with ori associated interactions in the plasmid.
- the computer program product provides an output (sequence of a) plasmid characterized by a maximized copy number, an optimized copy number, or a predetermined copy number, comprising at least one sequence associated with an ori associated interaction comprising at least one calculated mutation.
- the computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device.
- the computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing.
- a non- exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing.
- RAM random access memory
- ROM read-only memory
- EPROM or Flash memory erasable programmable read-only memory
- SRAM static random access memory
- CD-ROM compact disc read-only memory
- DVD digital versatile disk
- memory stick a floppy disk
- mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon
- a computer readable storage medium is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
- Computer readable program instructions described herein can be downloaded to respective computing/processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and/or a wireless network.
- the network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and/or edge servers.
- a network adapter card or network interface in each computing/processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing/processing device.
- Computer readable program instructions for carrying out operations of the present invention may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the “C” programming language or similar programming languages.
- the computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server.
- the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider).
- electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present invention.
- These computer readable program instructions may be provided to a processor of a general-purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions/acts specified in the flowchart and/or block diagram block or blocks.
- These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and/or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function/act specified in the flowchart and/or block diagram block or blocks.
- Embodiments may comprise a computer program that embodies the functions described and illustrated herein, wherein the computer program is implemented in a computer system that comprises instructions stored in a machine -readable medium and a processor that executes the instructions.
- the embodiments should not be construed as limited to any one set of computer program instructions.
- a skilled programmer would be able to write such a computer program to implement one or more of the disclosed embodiments described herein. Therefore, disclosure of a particular set of program code instructions is not considered necessary for an adequate understanding of how to make and use embodiments.
- the term "about" when combined with a value refers to plus and minus 10% of the reference value.
- a length of about 1,000 nanometers (nm) refers to a length of 1,000 nm+100 nm.
- RNA II and RNA I Promoter sequence affects plasmid copy number
- Rouches et al. generated a library of pUC19 plasmids, each containing a different promoter sequence for either RNA II or RNA I. Notably, each plasmid only contains variations in the promoter regions of either RNA II or RNA I.
- the library was subsequently transformed into E. coli to measure the copy number of each plasmid.
- the inventors created a position- specific scoring matrix (PSSM) for the top and bottom 20% of high and low copy number plasmids.
- PSSM position- specific scoring matrix
- the inventors are building a copy number predictor based on a plasmid sequence.
- the predictor is based on the XGBoost model.
- the model gets sequence-related features for RNAI, RNAII promoters, including gaps, motifs, promotes strength, MIC, pssm, one hot encoding, and more.
- the inventors aim at extended the model also to RNAI and RNAII sequences and model their interactions.
- the inventors aim to model the metabolic burden of the plasmids on the host cell and the effect on gene expression and bioproduction. This can be achieved by building a whole- cell simulation with a pool of ribosomes under different conditions that translates plasmid genes.
- the inventors aim to model plasmid loss over generations and the variability of plasmid copy number in different cell populations.
- the inventors aim to gain new insights to optimize the system, including the effect of different abiotic factors on plasmid copy number.
- different abiotic factors can affect plasmid supercoiling, which is known to affect plasmid copy number.
- the current goal is to develop computational models that can simulate multi-plasmid SynBio systems, where each plasmid in use has a different copy number. By incorporating the different copy numbers of each plasmid, models will enable further investigation on how the expression of genes on each plasmid is affected and how this impacts the overall behavior of the system.
- the data from [Rouches, M.V., Xu, Y., Cortes, L.B.G., et al. (2022)] contains 366 inhibitory RNA promoter variants and 833 priming RNA promoter variants with their corresponding copy number.
- the variants were edited at positions (-33,-30), (-11,-8), and +1 for priming RNA and (-33,-30), (-10,-7), and +1 for inhibitory RNA.
- the original sequence for priming RNA is: TTGAGATCCTTTTTTTCTGCGCGTAATCTGCTGCTT (SEQ ID NO: 1)
- inhibitory RNA is: TTGAAGTGGTGGCCTAACTACGGCTACACTAGAAGA (SEQ ID NO: 2).
- the data contains additional features such as growth rate and predicted promoter strength (KbT).
- RNA polymerase recognizes DNA sequences through the promoter.
- the promoter identification is mediated by the subunit of RNA polymerase G factor in bacteria.
- the housekeeping G70 factor of Escherichia coli recognizes two DNA sequence elements ‘-10’ hexamer (5'-TATAAT-3') located at positions -7 to-12 and ‘-35’ hexamer (5'-TTGACA-3') located at positions -30 to -35.
- the housekeeping G70 factor of Escherichia coli recognizes two DNA sequence elements ‘-10’ hexamer (5'-TATAAT-3') located at positions -7 to-12 and ‘-35’ hexamer (5'-TTGACA-3') located at positions -30 to -35.
- the RNAp and RNAi promoters were mutated at 9 positions. It is assumed that these positions were changed because they correspond to - 10 and -35 hexamers.
- PSSM position-specific scoring matrix
- PSSM score calculated with the 20% high copy number only, was added as a feature called ‘pssm score’.
- Motifs files were downloaded from (1) DPINTERACT database which is a comprehensive library of binding sites for E. coli DNA-binding proteins containing a set of 68 motifs, between 10 and 50 in width (average width 24.5), and (2) SwissRegulon E.coli database which is a database containing genome-wide annotations of regulatory motifs, promoters and TF binding sites (TFBSs) in promoter regions across model organisms, contains a set of 97 motifs, between 6 and 30 in width (average width 20.1).
- the matrix covers base pairs [-4E-1] where 0 denotes the transcription start site. Each row corresponds to a given position; each column corresponds to a value for that base pair. [0180] The promoter strength was calculated for the entire sequence and for the two edited zones of the priming and inhibitory RNAs.
- MIC maximal information coefficient
- Lasso is a linear regression model that performs feature selection and regularization, making it a powerful tool for predicting outcomes in high-dimensional datasets.
- the lasso model uses LI regularization, which adds a penalty term equal to the absolute value of the coefficients of the regression parameters in order to shrink the coefficients of less important features to zero. This results in a more parsimonious model that only includes the most relevant features in the final prediction while also improving the model's performance by reducing overfitting.
- the optimization objective for Lasso is: where the regression coefficients are listed as beta_i.
- the default lambda regulation term is 1
- Lasso regression can be used for both classification and regression tasks, and is particularly useful in cases where the number of features is much larger than the number of observations. It has been widely used in applications such as gene expression analysis, image processing, and financial modeling, among others.
- XGBoost Extreme Gradient Boosting. This algorithm is based on the gradient boosting trees algorithm but improves its speed and has been proven to work well even on sparse data.
- the loss function used is: mean squared error.
- the algorithm starts by guessing 0.5 as an answer, then it calculates the difference between the real training data results and the guess (from now on called residuals).
- the following step will be to build a tree that tries to minimize the residuals.
- Each tree starts from a single leaf, the leaves containing all of the residuals.
- the inventors start splitting the tree based on one feature at a time, and for each feature, the training data is split into quantiles, and the leaf is split for each option, the inventors choose the best-split feature and value that results in the highest Gain.
- a tree has a depth limit, and when one reaches this limit, a new tree that again tries to minimize the residuals calculated by the last tree is stared, in which the output of the matching leaf is: (sum of the residuals)/(num of residuals + regularization value). Then, the inventors add the results of each tree multiplied by a learning rate. The inventors keep building trees until the residuals are small enough or the maximum number of trees is reached.
- XGBoost [0196] A converging random search approach was used in order to find the best hyperparameters for both RNAi and RNAp. The number of trees, maximum depth of each tree, learning rate, gamma, and colsample bytree were optimized.
- the estimated R-squared is defined as:
- chromoprotein is added to the plasmid.
- the inventors use 3 different chromoproteins throughout the current experiments, all of which are official parts from the iGEM registry (BBa_K1692032, BBa_K1033917, BBa_K1033926). Using restriction enzymes and DNA ligase, one chromoprotein is added to the plasmid.
- Quantitative PCR is used to measure the copy number of a given plasmid. This is achieved by measuring the ratio between two single-copy genes (one found in the chromosome and the other on the plasmid itself). This way plasmid copy number can be calculated using the following equation:
- [plasmid copy number] [copy of plasmid gene] / [copy of chromosomal gene].
- LB medium Luria-Bertani (LB) medium
- 20 g of LB powder was added to 1 liter of water, and the mixture was autoclaved to ensure sterility. Once the solution cooled down to 60 °C, antibiotics were added in a 1:1,000 ratio to create a selective medium.
- LB agar plate preparation 20 g of LB powder and 20 g of agar were added to 1 liter of water. The mixture was thoroughly mixed using a stirrer and then autoclaved. After the liquid reached a temperature of 60 °C, antibiotics were added in a 1:1,000 ratio. A flame was lit, and the LB agar solution was carefully poured into sterile Petri dishes to solidify.
- a sterile toothpick was used to carefully lift off a bacterial colony from a culture plate. The colony was then transferred onto the surface of a fresh agar plate. The plate was subsequently incubated overnight at 37 °C to allow for colony growth and expansion.
- a sterile toothpick was used to lift off a bacterial colony from a culture plate of E. coli DH5-alpha strain.
- the colony was inserted into a starter culture containing 3 to 5 ml of growth media.
- the starter culture was then incubated at 37 °C with shaking at 250-300 rpm overnight to promote bacterial growth and achieve a sufficient cell density for subsequent experiments.
- the bacterial culture was centrifuged at maximum speed for 2 minutes using a pull-down mode, resulting in cell pellet formation. The supernatant was carefully removed, and the pellet was resuspended by adding 1 ml of DDW and pipetting thoroughly. Finally, the transformed bacteria were plated onto an LB agar plate supplemented with an appropriate antibiotic. The plate was incubated overnight at 37 °C to enable the formation of colonies, indicative of successful transformation.
- the inventors analyzed the pUC19 plasmid library by Rouches et al., consisting of distinct promoter variants for RNAp or RNAi. The mutations are specifically confined to the promoter regions of RNAp or RNAi.
- the data contains 366 RNAi promoter variants and 833 RNAp promoter variants with their corresponding copy number. The variants were edited at positions (-33,-30), (-11,-8), and +1 for RNAp and (-33,-30), (-10,-7) , and +1 for RNAi (Fig. 25F).
- the inventors divided the data into 3 sets, 15% test set, 15% validation set, and 70% train set. During this process, the inventors ensured that the copy number distribution in each set remained consistent with the original dataset.
- the promoter’s sequences can’t be used directly as input for the models, so, about 6,000 features were engineered to represent the sequence properties that correspond to transcriptions in prokaryotes.
- Motifs files were downloaded from the DPINTERACT database and the SwissRegulon E.coli database. The score and P-value of each motif were calculated with the FIMO module, and sequences with no matching motif were replaced with a 0 score, and p-value of 1. (330 features)
- Promoter strength The promoter strength was calculated as described in Brewster et al. (3 features)
- nucleotide counts features nucleotide counts features from an article by Muhammod et al. (5945 features)
- Entropy The promoter sequence entropy was calculated as follows: is the probability of each nucleotide to appear in each index i. (1 feature)
- RNAi promoters using the above features, produced less successful results than the model based on RNAp promoters.
- the inventors added features describing the RNAp secondary structure, considering the complementarity between RNAi sequences and RNAp promoters. Mutations in RNAi's promoter can impact RNAp's structure, leading to alterations in its secondary structure.
- the inventors investigated the replication mechanism of the ColEl plasmid family, specifically pUC19, where RNAp's persistence on the DNA template is vital for initiating replication. A key interaction involves a G-rich loop in the RNAp transcript binding to a C-rich stretch on the DNA, affecting RNAp conformations and the formation of RNA-DNA hybrids necessary for DNA synthesis.
- RNAp contains three binding sites: alpha, beta, and gamma.
- the RNAp transcript undergoes structural changes during transcription, forming three stem-loop domains and later an extended stem- loop. This folding is influenced by the pairing of alpha-beta or beta- gamma sequences, which determines the transcript's ability to form the hybrid for DNA synthesis.
- RNAi's structure with three stem loops and an unpaired region, binds to RNAp, influencing its folding and favoring beta-gamma pairing. This interaction affects primer formation, essential for replication.
- the current study used RNAi and RNAp sequences from the pUC19 plasmid, analyzing RNAi promoter mutations.
- the inventors focused on nine specific mutation positions in the RNAi promoter, corresponding to key points in the RNAp sequence.
- the inventors employed the ViennaRNA package, particularly RNAfold and RNAeval, to calculate RNAp's secondary structure, considering the impact of RNAi promoter mutations.
- This approach allowed the inventors to generate features describing RNAp's secondary structure at different transcription stages, enhancing the current RNAi- based model's accuracy.
- the features (additional -25,000 features to the features detailed above) are calculated for a range of lengths of RNAp in order to describe its structure during transcription.
- the inventors explored various models, including Neural Networks, Linear Regression models, Random Forest, and Ensemble Learning models, while tuning their hyperparameters with Optuna (more details below).
- the inventors assessed the Spearman and Pearson correlations on both the training and validation sets.
- XGBoost, CatBoost, Light GBM, and Random Forest demonstrated superior performance. Consequently, the inventors advanced to the subsequent stages, feature selection and hyperparameter optimization, with these selected models.
- Optuna automates the search for optimal hyperparameter configurations using Bayesian optimization, where each trial employs a different set of hyperparameter values and evaluates an evaluation score on the validation set.
- the feature selection was performed in two steps: (1) Mutual information-based feature selection: the inventors calculated the mutual information (MI) between each feature and the copy number, selecting those with MI values greater than the mean MI as potential candidates. Subsequently, the inventors assessed the MI between the remaining features, identifying pairs of features exhibiting strong correlations within the top percentile. For each pair of correlated features, the inventors chose the one displaying a higher MI with the copy number, ensuring the retention of the most informative attributes.
- MI mutual information
- Boruta-SHAP is a feature selection algorithm that combines the Boruta algorithm with SHAP values to select important features:
- Boruta is a tree-based feature selection algorithm that works by creating random shadow features (noise) and comparing the importance of the original features to the importance of the shadow features. Features that are more important than the shadow features are kept, while features that are not more important than the shadow features are removed,
- SHAP values are a way to measure the individual impact of each feature on the prediction of a model. SHAP values are calculated by considering all possible combinations of features and their impact on the prediction.
- Boruta-SHAP works by using SHAP values to rank the importance of the features and the shallow features.
- the CatBoost model was employed for predicting the final outcomes of the test set in the case of RNAi.
- a voting model was adopted, wherein the final prediction is determined by averaging the predictions from the ensemble learning model, CatBoost, and XGBoost. This voting model approach not only enhances performance but also mitigates noise and overfitting.
- the model performance was evaluated using Pearson and Spearman correlations, R-squared score, and mean absolute error (MAE). The model development workflow is demonstrated in Fig. 29C.
- XGBoost employs a distinct method for assessing feature importance, known as "gain". This method calculates the contribution of each feature to the model by considering the improvement in accuracy brought about by splitting on that feature, averaged over all trees.
- the gain is a measure of the average contribution of a feature to the model's prediction performance: the higher the gain, the more important the feature.
- the CatBoost model uses a combination of the SHAP values and internal algorithms to rank features, providing a measure of the impact each feature has on the variation in the predicted outcome. This model- agnostic approach is beneficial as it offers a consistent way to interpret feature importance across different types of models.
- the inventors visualized the features' distribution by their corresponding copy number using violin plots for discrete features and scatter histograms for continuous features, in order to identify patterns, trends, and outliers.
- SHAP Siliconley Additive exPlanations
- an interpretable Al approach leveraging Shapley values is a model- agnostic technique designed to elucidate the impact of individual features on model predictions. This method assigns a Shapley value to each feature, providing a quantitative measure of its contribution to the model's output for a specific instance.
- Features with high absolute Shapley values are indicative of having a substantial influence on predictions.
- the inventors employed SHAP to examine the selected features and assess their influence on the decision-making process of the models. This approach allows the inventors to identify the critical features and understand how specific values of each feature contribute to the prediction of a high copy number. This interpretability framework was particularly valuable as the inventors sought to enhance the comprehension of the colEl replication system's mechanisms. By delving into the model structure and leveraging SHAP, the inventors aimed to unravel the intricate dynamics underlying the replication system and glean insights into its functional intricacies.
- the modified origin of replications were synthesized based on the tool’s output.
- the original pUC19 plasmid was amplified using PCR to introduce new flanking regions around the ORI, then a cut site was targeted with a restriction enzyme, to cut the original ORI.
- the modified ORI was inserted using homologues joining with NEBuilder® HiFi DNA Assembly kit to the pUC19 backbone.
- the altered plasmids were introduced into E. coli DH5 alpha through a chemical transformation and verified using sanger sequencing. The process is demonstrated in Fig. 29.
- a colony was picked and mixed in DDW into a dilution of 10“ 6 .
- a reaction mix was prepared with: 10 pl fast SYBR Green Master Mix, 2 pl of each primer and 5 pl DDW. Nineteen (19) pl of the reaction mix was transferred into qPCR plate, and 1 pL of the colony sample was added. The PCR plate went through spin down for 1 minute at 800 rpm prior to its loading in the instrument.
- the inventors Starting from the relative copy number, the inventors improved the mathematical fitting of the data to the measured copy number. For this purpose, the inventors added 30% more data points to the sequences used for the fitting while saving aside an equal number of data points for validation. Analyzing the data, the inventors observed a power log behavior connecting the measured copy number to the relative copy number; therefore, the inventors decided to perform a linear regression of the logarithmic values to transform the relative copy number to copy number estimations.
- a model able to predict the copy number from the ORI sequence will have practical implementations, allowing researchers to design new plasmids with a specific copy number tailored to their individual needs. Creating it as an explainable model will also allow the inventors to gain new insights into the biology of Col-El -like plasmids replication, insights that can be capitalized upon in future research.
- the inventors used a database of tagged ORI sequences (meaning different ORI variants each with a corresponding copy number) that was recently published, wherein two libraries of ORI variants, one with mutations in the RNAp promoter region of the ORI and the second with changes in the RNAi promoter region were created. The inventors therefore chose to create two separate models, for each of the datasets.
- the inventors first set to find the best fitting modal for the task. To test different models, the inventors split the data into training, test, and validation sets (70%, 15%, and 15% respectively) and verified that all sets feature a copy number distribution similar to that of the entire dataset.
- RNAp and RNAi were trained 8 different models for RNAp and RNAi each, and tested their performance on the test set results are shown in Fig. 30.
- CatBoost and XGBoost both performed best, with similar results (see more details about these models in the supplementary).
- the inventors utilized a voting model using both of them, averaging their predictions (model performance shown in Fig. 25).
- CatBoost outperformed the rest of the models (results shown in Figs. 25A-25B).
- both models exhibited a very high correlation between the model predictions and the test set, demonstrating the ability of the model to predict the copy number of plasmids that were not used for its training, based on the RNAp or RNAi promoter sequences alone.
- RNAp The most important feature for RNAp was dG_total for CatBoost and rpoD16_score for XGBoost.
- dG_total is the difference in free energy between an unbound promoter and a promoter bound to a sigma factor + RNA polymerase.
- rpoD16_score is a numerical value representing the strength or likelihood of the presence of the rpoD16 motif, a specific bacterial sigma factor recognition site. Higher scores suggest a higher probability of successful biosynthesis.
- RNAFold can provide the probability of each nucleotide pairing with a nearby nucleotide.
- seq_end_400_375_385 means the inventors used the first 400 bases of the RNA sequences and checked the probability of the nucleotide in index 375 to pair with the nucleotide in index 385.
- the sequence TGCGCGTTTAAT (SEQ ID NO: 7) is a de-novo motif that appears with higher rates in groups of promoters that have high copy number, against a reference group of promoters with a low copy number.
- Figs. 26A-26B display the selected features along with their respective copy numbers and the distribution patterns for these features. It displays the common features between the XGBoost and CatBoost models used in the RNAp model, as well as the two most significant features of the CatBoost model used in the RNAi model.
- RNAp related features distribution graph (Fig. 26A)
- rpoD16_score demonstrates a pronounced peak, signifying a strong effect at a specific score on the copy number.
- the plot for the feature dG_total indicates a spread of values with most data points concentrated around -3.4 to -2.7, implying a region of stability that could be critical for high copy numbers.
- the scatter plot for seq_end_400_375_385 shows a dense accumulation of points at the lower end of the scale, which may indicate a significant correlation with increased copy numbers.
- the feature TGGCGGTTTAAT (SEQ ID NO: 8)_denovo_HIGH displays a spread of data points, with notable bars extending toward middle values, hinting at the possibility of distinct conditions or states influencing the copy number within this feature.
- SHAP is a tool used to interpret machine learning models by assigning each feature an importance value for a particular prediction. It is based on Shapley values from game theory, which allocates fair payoffs to players depending on their contribution to the total payout. SHAP helps to understand which features most strongly influence the model's predictions, revealing the impact of each feature on the predicted outcome. This insight allows the inventors to identify the features that are most significant in determining the copy number, thereby providing a deeper understanding of the underlying plasmid replication system. The SHAP analysis yielded a quantitative measure of each feature's influence on the model's predictions, spotlighting the features with high absolute Shapley values across both RNAp and RNAi models. Figs.
- 27A-27B show the SHAP value for each feature chosen to train the models during the feature selection process.
- dG_total surfaced as the most influential feature, leading to significant positive impacts on predicted copy numbers.
- the strong positive influence of dG_total was particularly striking against the backdrop of more moderate effects from other features, positioning dG_total as a potentially critical determinant in the RNAp model's decision-making process.
- TGCGGTTTAAT SEQ ID NO: 9
- denovo_HIGH was identified as having the highest mean absolute SHAP values, signifying its substantial role in the model's output.
- dG_total appeared to be significant in the current analysis, the inventors continued to analyze it using a SHAP force plot.
- the dG_total feature represents the overall change in free energy (Gibbs free energy) between an unbound promoter and a promoter bound to a sigma factor + RNA polymerase which is associated with the plasmid replication process.
- a lower dG_total value indicates a more thermodynamically favorable biosynthesis reaction.
- This optimal range located between -3.4 to -2.7, corresponds to a more thermodynamically favorable reaction for plasmid replication, thereby indicating a threshold effect where certain energy levels are conducive to higher copy numbers.
- the presence of clustering within this range suggests a stable and efficient replication process, with dG_total values outside this range potentially leading to less favorable replication outcomes.
- This range represents the reaction's thermodynamic 'sweet spot' - a state of maximum favorability characterized by the stable interaction of molecular complexes, efficient regulatory mechanisms, environmental adaptation, and optimal resource utilization.
- understanding this optimal dG_total range is crucial for enhancing the efficiency and predictability of plasmid replication processes, which is of significant interest in genetic engineering and biotechnology applications.
- the inventors were successfully able to replace the origins of replications of apUC19 plasmid with new modified sequences taken from the literature or generated by the current model using a high throughput method.
- This innovative approach allowed the inventors to easily construct a library of plasmid sequences which played a crucial role in the verification and improvement of the current model.
- This procedure involves synthesizing modified origins of replication based on tool outputs, introducing them into the pUC19 plasmid via PCR amplification and homologous recombination, and then verifying the modified plasmids after transformation into E. coli using Sanger sequencing, confirming the presence of the intended mutations within the RNAp and RNAi promoters (see more details in Fig.
- the inventors proceeded to calculate the copy numbers of the newly constructed plasmids.
- the inventors adopted a colony qPCR approach to measure the plasmid copy numbers for the various constructs within our library. The results of this copy number analysis provided valuable insights into the replication dynamics of these engineered plasmids. Based on these findings, as anticipated, it is evident that substituting the promoter sequences at the origin of replication has an impact on the plasmid copy number.
- the inventors used Rouches et al. ddPCR measurements, in addition to in-house qPCR measurements to improve the current model. As done in Rouches et al., the inventors also performed a mathematical transformation to the copy number data used to train the current model for improving the model performance (see more details in the methods section). The inventors achieved a high correlation on both the test set and the current biological results, as can be seen in Fig. 28B. This further verifies the reliability of the current model and the success of introducing a machine-learning model for copy number engineering and prediction.
- the inventors used the machine learning model results and developed a software tool (available at: pen- radient, vercel.app) that enables researchers to design and predict plasmid copy numbers simply and efficiently.
- the model s high performance allows the current tool to accurately predict the copy number of any plasmid, based on its sequence and the host strain in which it will be used.
- the tool offers several key features to enhance plasmid sequence manipulation and customization. As demonstrated in Fig. 28, users can input a plasmid sequence (FASTA or GenBank files) along with the desired copy number, and the tool generates a modified sequence with ORI adjustments to achieve the specified copy number.
- the tool efficiently generates FASTA files, ensuring precision in achieving the desired copy number.
- users can request multiple copy numbers simultaneously, receiving dedicated sequences for each.
- the tool's output includes clear and concise plasmid maps, seamlessly integrated with annotations within the sequence, providing comprehensive information. Accessible on both computers and mobile devices, the tool caters to diverse user needs. Furthermore, developers can seamlessly integrate the tool into their projects using the provided API, enhancing its adaptability and accessibility.
- the inventors suggest for the first time a model that can predict plasmid copy number based on the ORI sequence.
- a library of plasmid sequences was constructed, and a colony qPCR method was used to measure copy numbers, validating the model's predictions and allowing for improvements.
- the current model can be used both for understanding the regulation of plasmid copy number and for engineering it.
- the inventors developed two separate machine learning models to predict copy number from RNAp and RNAi promoter sequences using a database of tagged ORI sequences.
- the models demonstrated high correlation in predicting plasmid copy numbers not used for training the model.
- the inventors found that the sequence features related to transcription regulation of the RNA genes encoded in the ORI (e.g., dG_total and rpoD16_score) are the most important features in the current models. These features represent aspects related to transcription efficiency such as the free energy of promoter binding and the strength of sigma factor recognition sites.
- RNAp and RNAi within the ColEl plasmid family play a pivotal role in the regulation of plasmid copy number.
- the inclusion of secondary structure features into the current computational model has significantly enhanced its predictive accuracy.
- the inventors were able to quantify how mutations in the RNAi promoter impact the RNAp secondary structure, which is critical for designing strategies to optimize plasmid stability and efficiency in synthetic biology applications.
- the inventors developed a web-based software tool which offers a user-friendly platform for researchers to efficiently design, predict, and modify plasmid sequences, thereby streamlining research workflows and enhancing the precision of genetic manipulations.
Landscapes
- Genetics & Genomics (AREA)
- Health & Medical Sciences (AREA)
- Engineering & Computer Science (AREA)
- Life Sciences & Earth Sciences (AREA)
- Chemical & Material Sciences (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Organic Chemistry (AREA)
- Biotechnology (AREA)
- General Engineering & Computer Science (AREA)
- Zoology (AREA)
- Wood Science & Technology (AREA)
- Biomedical Technology (AREA)
- Microbiology (AREA)
- Plant Pathology (AREA)
- Molecular Biology (AREA)
- Physics & Mathematics (AREA)
- Biochemistry (AREA)
- General Health & Medical Sciences (AREA)
- Biophysics (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
Abstract
Description
Claims
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| EP24818917.7A EP4720308A1 (en) | 2023-06-05 | 2024-06-05 | Computational models for effecting copy number of plasmids |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US202363471014P | 2023-06-05 | 2023-06-05 | |
| US63/471,014 | 2023-06-05 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2024252392A1 true WO2024252392A1 (en) | 2024-12-12 |
Family
ID=93795275
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/IL2024/050552 Ceased WO2024252392A1 (en) | 2023-06-05 | 2024-06-05 | Computational models for effecting copy number of plasmids |
Country Status (2)
| Country | Link |
|---|---|
| EP (1) | EP4720308A1 (en) |
| WO (1) | WO2024252392A1 (en) |
-
2024
- 2024-06-05 EP EP24818917.7A patent/EP4720308A1/en active Pending
- 2024-06-05 WO PCT/IL2024/050552 patent/WO2024252392A1/en not_active Ceased
Non-Patent Citations (2)
| Title |
|---|
| FANG DIANNE: "Attempts to Use PCR Site-Directed Mutagenesis to Create a Nonfunctional rop Gene in the Plasmid pBR322", JOURNAL OF EXPERIMENTAL MICROBIOLOGY AND IMMUNOLOGY, vol. 6, 31 December 2004 (2004-12-31), pages 45 - 51, XP093247586 * |
| ROUCHES MILES V., XU YASU, CORTES LOUIS BRIAN GEORGES, LAMBERT GUILLAUME: "A plasmid system with tunable copy number", NATURE COMMUNICATIONS, vol. 13, no. 1, UK, XP093247583, ISSN: 2041-1723, DOI: 10.1038/s41467-022-31422-0 * |
Also Published As
| Publication number | Publication date |
|---|---|
| EP4720308A1 (en) | 2026-04-08 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Chen et al. | Engineering a precise adenine base editor with minimal bystander editing | |
| Leppek et al. | Combinatorial optimization of mRNA structure, stability, and translation for RNA-based therapeutics | |
| Arbab et al. | Determinants of base editing outcomes from target library analysis and machine learning | |
| HamediRad et al. | Towards a fully automated algorithm driven platform for biosystems design | |
| Kelsic et al. | RNA structural determinants of optimal codons revealed by MAGE-Seq | |
| Rabani et al. | A massively parallel reporter assay of 3′ UTR sequences identifies in vivo rules for mRNA degradation | |
| Du et al. | Regulatory transposable elements in the encyclopedia of DNA elements | |
| Liang et al. | Transcriptome-scale RNA-targeting CRISPR screens reveal essential lncRNAs in human cells | |
| Zhao et al. | Massively parallel functional annotation of 3′ untranslated regions | |
| Engreitz et al. | Local regulation of gene expression by lncRNA promoters, transcription and splicing | |
| Xie et al. | sgRNAcas9: a software package for designing CRISPR sgRNA and evaluating potential off-target cleavage sites | |
| Fei et al. | Advancing protein evolution with inverse folding models integrating structural and evolutionary constraints | |
| Kleinstiver et al. | High-fidelity CRISPR–Cas9 nucleases with no detectable genome-wide off-target effects | |
| Patwardhan et al. | Massively parallel functional dissection of mammalian enhancers in vivo | |
| Hoffmann et al. | Accurate mapping of tRNA reads | |
| o’Brien et al. | Unlocking HDR-mediated nucleotide editing by identifying high-efficiency target sites using machine learning | |
| Lin et al. | Expression dynamics, relationships, and transcriptional regulations of diverse transcripts in mouse spermatogenic cells | |
| van Schendel et al. | SIQ: easy quantitative measurement of mutation profiles in sequencing data | |
| Cowan et al. | Development of multiplexed orthogonal base editor (MOBE) systems | |
| Zou et al. | Cas9-PE: a robust multiplex gene editing tool for simultaneous precise editing and site-specific random mutation in rice | |
| Rangan et al. | De novo 3D models of SARS-CoV-2 RNA elements and small-molecule-binding RNAs to aid drug discovery | |
| CN119274660A (en) | A method for analyzing the effect of 5' UTR on mRNA translation efficiency based on machine learning model | |
| Wang et al. | Constructing eRNA-mediated gene regulatory networks to explore the genetic basis of muscle and fat-relevant traits in pigs | |
| Wang et al. | De novo design of insulated cis-regulatory elements based on deep learning-predicted fitness landscape | |
| Dent et al. | A basic framework to explain splice-site choice in eukaryotes |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24818917 Country of ref document: EP Kind code of ref document: A1 |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 2024818917 Country of ref document: EP |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| ENP | Entry into the national phase |
Ref document number: 2024818917 Country of ref document: EP Effective date: 20260105 |
|
| ENP | Entry into the national phase |
Ref document number: 2024818917 Country of ref document: EP Effective date: 20260105 |
|
| ENP | Entry into the national phase |
Ref document number: 2024818917 Country of ref document: EP Effective date: 20260105 |
|
| WWP | Wipo information: published in national office |
Ref document number: 2024818917 Country of ref document: EP |











